Short-focus projection dynamic distortion compensation method based on depth vision modeling

By constructing 3D modeling results and nonlinear mapping tensors using an improved Instant-NGP model, the problem of dynamic distortion compensation in complex environments for short-throw projection technology is solved, achieving high-precision, fast distortion repair and visual consistency.

CN121883322AInactive Publication Date: 2026-04-17GUANGDONG HANYING OPTOELECTRONICS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG HANYING OPTOELECTRONICS TECHNOLOGY CO LTD
Filing Date
2026-01-08
Publication Date
2026-04-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing short-throw projection technologies suffer from problems such as lag in dynamic distortion compensation, response delay, and insufficient compensation accuracy in complex backgrounds and irregular curved surfaces. In particular, the compensation accuracy is insufficient in edge and occluded areas, which affects visual consistency and user immersion experience.

Method used

An improved Instant-NGP model is adopted, and multimodal visual information is combined to construct a 3D modeling result, generate a nonlinear mapping tensor, generate a distortion compensation image through pixel retargeting, and dynamically update the mapping tensor based on the residual image to form a dynamic compensation closed loop.

Benefits of technology

It achieves high-precision distortion compensation under dynamic viewpoint changes and spatial geometric perturbations, improves compensation response speed and robustness, and enhances spatial consistency and visual continuity of images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883322A_ABST
    Figure CN121883322A_ABST
Patent Text Reader

Abstract

The invention discloses a short-focus projection dynamic distortion compensation method based on depth vision modeling, and the method comprises the following steps: S1, collecting the multi-modal vision information of a projection region, and forming a vision data set; s2, inputting into an improved Instant-NGP model, and generating a three-dimensional modeling result; s3, constructing a spatial attitude matrix; s4, generating a nonlinear mapping tensor of the projection image space and the surface space; s5, performing pixel redirection processing on the original projection image to generate a distortion compensation image; s6, constructing a residual image, calculating a structure deviation value and an edge offset, and updating a nonlinear mapping tensor; and S7, dynamically adjusting the reconstruction frequency of the improved Instant-NGP model, the updating frequency of the nonlinear mapping tensor and the output frequency of the distortion compensation image to form a dynamic compensation closed loop. According to the invention, pixel-level dynamic distortion compensation in a short-focus projection scene is realized, and the method has the technical advantages of high modeling precision, fast compensation response and strong space consistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision and image processing technology, and in particular to a method for compensating for dynamic distortion in short-focal-length projection based on depth vision modeling. Background Technology

[0002] With the widespread application of short-throw and ultra-short-throw projection technologies in education, exhibitions, and immersive interactive scenarios, dynamic distortion compensation methods for complex backgrounds and irregular curved surfaces have become an important research direction in the fields of visual enhancement and image reconstruction. Existing image distortion correction schemes mostly rely on static calibration or regular plane modeling for pre-compensation, primarily employing two-dimensional geometric mapping or local stretching interpolation methods to resample the projected image. However, in real-time dynamically changing scenarios, the following problems commonly exist:

[0003] Due to the large projection angle and wide projection range of short-throw projection, it is prone to large-scale nonlinear distortion in complex backgrounds such as curved walls, rough objects, and irregular structures. Traditional static mapping relationships cannot adapt to the nonlinear distortion caused by rapid posture changes. Most existing modeling methods are based on preset patterns or single-frame reprojection sampling, lacking the ability to reconstruct three-dimensional structures in continuous time sequences. This results in lagging compensation strategies and significant response delays, failing to support the need for seamless compensation under high-speed dynamic changes. Most solutions do not fully utilize the causal relationship between depth information and projection distortion, failing to accurately capture minute projection misalignments caused by surface height differences. In particular, the compensation accuracy is insufficient in edge and occluded areas, resulting in blurring, abrupt changes, or ghosting in the overall displayed image, seriously affecting visual consistency and user immersion experience.

[0004] In addition, although some solutions introduce depth cameras to assist in modeling, their reconstruction frequency is fixed and the image distortion output frequency is limited by computing resources. They lack dynamic self-adjustment capabilities and it is difficult to improve the real-time compensation and modeling accuracy while preserving image stability. This restricts the intelligent adaptability and compensation robustness of short-throw projection in complex environments.

[0005] Therefore, how to provide a method for dynamic distortion compensation of short-throw projection based on depth vision modeling is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a dynamic distortion compensation method for short-focal-length projection based on depth vision modeling. This invention introduces an improved Instant-NGP model, combines multimodal visual information of the projection area to construct a 3D model, constructs a spatial attitude matrix based on spatial position data and attitude angle data, further generates a nonlinear mapping tensor between the projection image space and the surface space, performs pixel redirection processing to generate a distortion-compensated image, and dynamically updates the nonlinear mapping tensor based on the structural deviation value and edge offset of the residual image. Finally, a dynamic compensation closed loop is formed by adjusting the modeling and reconstruction frequency, the mapping update frequency, and the compensation output frequency. This invention has the advantages of high modeling accuracy, stable distortion repair, fast compensation response, and strong adaptability.

[0007] A method for compensating for dynamic distortion in short-focal-length projection based on depth vision modeling according to an embodiment of the present invention includes the following steps:

[0008] S1. Collect multimodal visual information of the projection area to form a visual dataset; the multimodal visual information includes color images, depth images, spatial position data of the projection device, attitude angle data and ambient lighting parameters;

[0009] S2. Input the visual dataset into the improved Instant-NGP model to generate a 3D modeling result; the 3D modeling result includes point cloud data, surface normal information, pixel confidence map and color adjustment map; the improved Instant-NGP model includes a pose-aware control structure, a depth confidence guidance structure, a dynamic residual compensation structure and an optimized spiral output structure.

[0010] S3. Construct a spatial attitude matrix based on the spatial position data and attitude angle data;

[0011] S4. Based on the three-dimensional modeling results and the spatial attitude matrix, generate a nonlinear mapping tensor between the projected image space and the surface space;

[0012] S5. Perform pixel redirection processing on the original projected image according to the nonlinear mapping tensor to generate a distortion-compensated image;

[0013] S6. Acquire the display image of the distortion compensation image on the projection surface, construct the residual image, calculate the structural deviation value and edge offset between the residual image and the distortion image, and update the nonlinear mapping tensor.

[0014] S7. Based on the changing trend of the 3D modeling results, the rate of change of the spatial pose matrix, and the deviation index of the residual image, dynamically adjust the reconstruction frequency, the update frequency of the nonlinear mapping tensor, and the output frequency of the distortion compensation image of the improved Instant-NGP model to form a dynamic compensation closed loop.

[0015] Preferably, S1 specifically comprises:

[0016] The system acquires color images, depth images, spatial position data of the projection device, attitude angle data, and ambient lighting parameters of the projection area. The color images are acquired through an imaging component that covers the entire projection area and outputs continuous frames of color image data. The depth images are acquired through a depth sampling component that works synchronously with the imaging component, outputting depth image data with the same time frame as the color images. The spatial position data is acquired through a positioning component that outputs three-dimensional coordinate values. The attitude angle data is acquired through an attitude sensing component that outputs attitude data including pitch, yaw, and roll angles. The ambient lighting parameters are acquired through a lighting detection component that outputs light intensity and color temperature values. The color images, depth images, spatial position data, attitude angle data, and ambient lighting parameters are then synchronously processed according to a unified time reference to form a visual dataset.

[0017] Preferably, the improved Instant-NGP model includes an attitude-aware control structure, a depth confidence guidance structure, a dynamic residual compensation structure, and an optimized spiral output structure, specifically:

[0018] The attitude perception and control structure receives spatial position data and attitude angle data from the projection device, performs attitude vector construction operation, and generates attitude embedding vector; the attitude embedding vector is concatenated with the image coordinate encoding vector to output a fused attitude feature vector.

[0019] The depth confidence-guided structure receives a depth image and a pixel confidence map, constructs a confidence weight map, and then performs a fusion process with the image feature vector to generate a weighted depth feature map.

[0020] The dynamic residual compensation structure receives consecutive frames of color images, extracts pixel difference maps and optical flow estimation maps between adjacent image frames, constructs temporal residual feature vectors, and fuses temporal residual feature vectors with weighted depth feature maps to generate dynamic feature response maps.

[0021] The optimized spiral output structure receives and fuses attitude feature vectors, weighted depth feature maps, and dynamic feature response maps, constructs volume rendering paths and color prediction paths, and outputs 3D modeling results.

[0022] Preferably, the 3D modeling results include point cloud data, surface normal information, pixel confidence maps, and color adjustment maps, specifically:

[0023] The point cloud data is obtained by jointly encoding the image coordinates and depth values ​​of each pixel in the color image and depth image, and combining the spatial position data and attitude angle data of the projection device to calculate the corresponding three-dimensional spatial coordinate values, thus forming a three-dimensional coordinate set.

[0024] The surface normal information is based on the neighborhood structure of each point in the point cloud data. The normal vector direction of each three-dimensional point is determined according to the spatial gradient calculation rules, and a normal information image corresponding to the size of the color image is output.

[0025] The pixel confidence map is generated by weighted analysis based on the pixel stability, occlusion discrimination and temporal consistency features of the depth image, resulting in a confidence image with the same resolution as the color image.

[0026] The color adjustment map performs channel correction processing on the color image based on the light intensity and color temperature values ​​in the ambient light parameters, combined with the attitude angle data, and outputs a color adjustment map after brightness compensation and color equalization.

[0027] Preferably, S3 specifically includes:

[0028] Based on the spatial location data, extract the three-dimensional coordinate values ​​and construct the translation vector;

[0029] Based on the attitude angle data, pitch angle, yaw angle and roll angle are extracted to construct a set of rotation angles; the set of rotation angles is converted into a rotation matrix, and a three-dimensional orientation mapping relationship is generated using Euler angle transformation rules.

[0030] The rotation matrix and translation vector are combined in a matrix concatenation manner to form a spatial attitude matrix; the spatial attitude matrix is ​​a four-dimensional homogeneous transformation matrix, which characterizes the attitude change state and spatial displacement relationship of the projection device in the three-dimensional coordinate system.

[0031] Preferably, S4 specifically comprises:

[0032] Extract point cloud data and surface normal information from the 3D modeling results, combine them with the spatial attitude matrix to perform 3D spatial coordinate transformation, and generate a set of surface points in the target space.

[0033] Perform projection mapping calculations on the set of surface points to construct a pixel mapping relationship from the original image space to the target projection surface, and output the corresponding coordinate set;

[0034] Based on pixel mapping relationships, surface normal information, and pixel confidence maps, the deformation direction vector and occlusion status of each pixel are calculated.

[0035] By combining the color adjustment map, a pixel-level color response offset matrix is ​​established;

[0036] A nonlinear mapping tensor is constructed by merging pixel mapping relationships, deformation direction vectors, occlusion marker states, and color response offset matrices; the nonlinear mapping tensor includes pixel mapping coordinates between the projected image space and the surface space, geometric deformation vectors, occlusion masks, and color correction factors.

[0037] Preferably, S5 specifically includes:

[0038] Based on the pixel mapping coordinates recorded in the nonlinear mapping tensor, a coordinate relocation operation is performed on each pixel of the original projected image to generate a pixel relocation index table.

[0039] Based on the geometric deformation vector in the nonlinear mapping tensor, the pixel positions in the redirection index table are geometrically corrected to adjust the spatial distribution of pixels.

[0040] Based on the occlusion mask in the nonlinear mapping tensor, interpolation compensation is performed on the pixels in the occluded region to generate a continuous pixel distribution;

[0041] Based on the color correction factor in the nonlinear mapping tensor, color channel compensation and brightness balancing are performed on the redirected pixels to generate a set of pixels with consistent color.

[0042] The geometrically corrected pixel distribution is fused with the color compensation result to output a distortion-compensated image.

[0043] Preferably, the step of acquiring the display image of the distortion-compensated image on the projection surface and constructing the residual image specifically involves:

[0044] Acquire display image data on the projection surface, receive continuous frame color images output by the imaging component, and record the frame sequence corresponding to the distortion compensation image;

[0045] Spatial registration processing is performed on the acquired display image data to align the pixel coordinates of the display image with the pixel coordinates of the distortion compensation image, generating a sequence of registered images;

[0046] Perform difference calculation on each pixel of the registered image sequence, extract the pixel difference matrix between the distortion-compensated image and the display image, and generate the initial residual image;

[0047] High-pass filtering and edge enhancement are performed on the initial residual map to extract structural difference features and output a structural deviation map.

[0048] Calculate the brightness difference matrix and color difference matrix based on the structural deviation map, and construct a pixel-level residual intensity map;

[0049] The pixel difference matrix, structural deviation map, brightness difference matrix, and color difference matrix are stitched together according to the channel dimension to generate a residual image; the residual image includes the pixel difference distribution, structural deviation features, brightness difference distribution, and color difference distribution between the projected surface display image and the distortion compensation image.

[0050] Preferably, the step of calculating the structural deviation value and edge offset between the residual image and the distorted image, and updating the nonlinear mapping tensor, specifically involves:

[0051] Extract the brightness and color channels of the residual and distorted images, perform normalization on each channel, and generate standardized image pairs;

[0052] Perform gradient calculations on normalized image pairs to extract edge intensity maps and orientation gradient maps; calculate a set of pixel-level edge orientation vectors based on the orientation gradient maps.

[0053] Based on the edge intensity map and the set of direction vectors, perform edge correspondence matching to determine corresponding edge point pairs, calculate the spatial offset distance between corresponding edge points, and generate an edge offset matrix.

[0054] Correlation analysis is performed on the structural features of the residual image and the distorted image to construct a structural similarity matrix; the average structural deviation value is calculated based on the structural similarity matrix.

[0055] A deviation weight table is constructed based on the edge offset matrix and the average structural deviation value, and a global structural deviation mapping map is generated.

[0056] Perform a weight update operation on the geometric deformation vector and color correction factor in the nonlinear mapping tensor, adjust the pixel mapping coordinates and color response coefficients according to the deviation weight table, and output the updated nonlinear mapping tensor.

[0057] Preferably, S7 specifically includes:

[0058] Extract continuous 3D modeling results and perform difference calculations on point cloud data, surface normal information, and color adjustment maps;

[0059] Extract the attitude matrix in continuous space and perform difference calculation on the translation vector and rotation matrix;

[0060] Extract continuous residual images and perform variation analysis on structural deviation values ​​and edge offsets;

[0061] Based on the differences in 3D modeling results, changes in spatial pose matrix, and deviations in residual images, adjustment factors are generated.

[0062] The reconstruction frequency, the update frequency of the nonlinear mapping tensor, and the output frequency of the distortion-compensated image are dynamically adjusted based on the adjustment factor.

[0063] The beneficial effects of this invention are:

[0064] This invention introduces an improved Instant-NGP model, jointly inputting short-throw projection images, spatial depth information, and attitude angle data to construct a nonlinear mapping tensor between the projection image space and surface space. Pixel redirection processing is then performed to generate distortion-compensated images, solving the dynamic distortion problem caused by viewing angle changes, object deformation, and inconsistencies in depth in existing short-throw projection devices. In the modeling stage, by setting the spatial modeling voxel structure and activation density factor, a 3D reconstruction result is constructed using multiple frames of images and corresponding spatial poses. A construction error constraint function is introduced, and a spatial reconstruction tensor for distortion inversion is output. In the mapping phase, a mapping alignment index is established based on the spatial coordinates of the reconstructed tensor and the projected image. A nonlinear mapping tensor is generated through multi-scale interpolation. In the compensation phase, a compensation image construction operation is performed. Based on the residual image between the compensation image and the real image, structural deviation values ​​and edge offsets are extracted to construct a compensation update vector. The mapping tensor parameters are dynamically updated to achieve real-time dynamic compensation. In the control phase, a frequency control vector is constructed by setting the modeling and reconstruction frequency, the mapping update frequency, and the compensation output frequency. This dynamically adjusts the execution cycles of the mapping update process and the modeling and reconstruction process, forming a closed-loop control link for modeling and compensation. Ultimately, pixel-level distortion compensation output for short-focal-length projection images under dynamic viewing angle changes and spatial geometric disturbances is achieved, improving the spatial consistency and visual continuity of distortion repair. This approach offers technical advantages such as high modeling accuracy, fast compensation speed, and strong robustness. Attached Figure Description

[0065] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0066] Figure 1 This is a flowchart of a short-focal-length projection dynamic distortion compensation method based on depth vision modeling proposed in this invention.

[0067] Figure 2 This is a schematic diagram of the structure of the improved Instant-NGP model proposed in this invention;

[0068] Figure 3 This is a data flow diagram of a short-focal-length projection dynamic distortion compensation method based on depth vision modeling proposed in this invention. Detailed Implementation

[0069] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0070] refer to Figure 1-3 A method for compensating dynamic distortion in short-throw projection based on depth vision modeling includes the following steps:

[0071] S1. Collect multimodal visual information of the projection area to form a visual dataset; the multimodal visual information includes color images, depth images, spatial position data of the projection device, attitude angle data and ambient lighting parameters;

[0072] S2. Input the visual dataset into the improved Instant-NGP model to generate a 3D modeling result; the 3D modeling result includes point cloud data, surface normal information, pixel confidence map and color adjustment map; the improved Instant-NGP model includes a pose-aware control structure, a depth confidence guidance structure, a dynamic residual compensation structure and an optimized spiral output structure.

[0073] S3. Construct a spatial attitude matrix based on the spatial position data and attitude angle data;

[0074] S4. Based on the three-dimensional modeling results and the spatial attitude matrix, generate a nonlinear mapping tensor between the projected image space and the surface space;

[0075] S5. Perform pixel redirection processing on the original projected image according to the nonlinear mapping tensor to generate a distortion-compensated image;

[0076] S6. Acquire the display image of the distortion compensation image on the projection surface, construct the residual image, calculate the structural deviation value and edge offset between the residual image and the distortion image, and update the nonlinear mapping tensor.

[0077] S7. Based on the changing trend of the 3D modeling results, the rate of change of the spatial pose matrix, and the deviation index of the residual image, dynamically adjust the reconstruction frequency, the update frequency of the nonlinear mapping tensor, and the output frequency of the distortion compensation image of the improved Instant-NGP model to form a dynamic compensation closed loop.

[0078] This implementation method collects color images and depth images of the projection area, spatial position data and attitude angle data of the projection device, and ambient lighting parameters to construct a visual dataset containing multimodal information, ensuring the completeness of perception for complex scenes. By inputting the visual dataset into a 3D modeling structure that includes attitude perception control, depth confidence guidance, dynamic residual compensation, and rotational output optimization mechanisms, a 3D modeling result containing point cloud data, surface normal information, pixel confidence maps, and color adjustment maps is generated, achieving high-precision reconstruction of the geometric features and lighting characteristics of the projection surface. Furthermore, a spatial attitude matrix is ​​constructed based on the spatial position data and attitude angle data to clarify the spatial transformation relationship of the projection device between different frames, improving the temporal consistency of subsequent mapping calculations. Based on the 3D modeling result and the spatial attitude matrix, a spatial representation of the projection image is generated. The nonlinear mapping tensor between surface spaces effectively establishes the image redirection path. Pixel redirection processing is performed on the original projected image based on the nonlinear mapping tensor to generate a distortion-compensated image, improving the geometric consistency of the image on uneven surfaces. Furthermore, by acquiring the display results of the distortion-compensated image on the real projection surface, a residual image is constructed, and the structural deviation value and edge offset between it and the distorted image are calculated. The nonlinear mapping tensor is dynamically updated to enhance compensation accuracy. Simultaneously, based on the changing trend of the 3D modeling results, the rate of change of the spatial pose matrix, and the residual image deviation index, the 3D reconstruction frequency, mapping tensor update frequency, and compensation image output frequency are dynamically adjusted to construct a highly responsive and stable dynamic distortion compensation closed-loop system, significantly improving the real-time compensation capability and display consistency of the short-throw projection system in complex scenes.

[0079] In this embodiment, S1 specifically refers to:

[0080] The system acquires color images, depth images, spatial position data of the projection device, attitude angle data, and ambient lighting parameters of the projection area. The color images are acquired through an imaging component that covers the entire projection area and outputs continuous frames of color image data. The depth images are acquired through a depth sampling component that works synchronously with the imaging component, outputting depth image data with the same time frame as the color images. The spatial position data is acquired through a positioning component that outputs three-dimensional coordinate values. The attitude angle data is acquired through an attitude sensing component that outputs attitude data including pitch, yaw, and roll angles. The ambient lighting parameters are acquired through a lighting detection component that outputs light intensity and color temperature values. The color images, depth images, spatial position data, attitude angle data, and ambient lighting parameters are then synchronously processed according to a unified time reference to form a visual dataset.

[0081] In this embodiment, the improved Instant-NGP model includes an attitude-aware control structure, a depth confidence guidance structure, a dynamic residual compensation structure, and an optimized spiral output structure, specifically:

[0082] In the attitude perception and control structure, spatial position data and attitude angle data of the projection device are collected. Based on the Euler angle transformation rule, the attitude angle data is converted into a direction vector. The attitude embedding vector is constructed by splicing it with the spatial position data of the projection device through a three-dimensional space transformation formula. The attitude embedding vector is then fused with the image coordinate encoding vector, which is expanded based on a multi-scale trigonometric sine position function to form a spatial perception feature code. The fused vector is then embedded into the encoding tensor space through a linear mapping transformation to output a fused attitude feature vector. The fused attitude feature vector contains information on the spatial orientation and projection direction change trend corresponding to the current viewpoint, which can improve the robustness to attitude changes during the three-dimensional modeling process.

[0083] In the depth confidence-guided structure, the depth image and pixel confidence map are spatially aligned to generate a pixel confidence matrix at a uniform resolution. A confidence weight map is constructed, which is obtained by linearly normalizing each pixel value in the pixel confidence map with a set confidence interval. The confidence weight map is fused with the image feature vector, which is extracted from the color image through shallow convolution. A confidence weighting mechanism is used to enhance the feature response value of high-confidence regions and suppress the blur contribution of occluded or edge regions, outputting a weighted depth feature map. The weighted depth feature map can enhance the depth continuity and boundary sharpness in 3D reconstruction.

[0084] In the dynamic residual compensation structure, a sequence of color image frames at consecutive time points is input. A pixel displacement map between adjacent frames is calculated based on an optical flow estimation method, and then weighted pixel-wise fused with the pixel difference map to construct a temporal residual feature vector. The pixel displacement map is based on the PWC-Net structure to achieve optical flow field estimation, and the pixel difference map is obtained by calculating the absolute difference between two frames through each channel. The temporal residual feature vector is fused with the weighted depth feature map, and a channel attention mechanism is used to highlight the feature response of changing regions, generating a dynamic feature response map. The dynamic feature response map can effectively compensate for the loss of detail caused by scene movement and lighting changes during modeling, enhancing the temporal consistency of the modeling output.

[0085] In optimizing the spiral output structure, a fused attitude feature vector, a weighted depth feature map, and a dynamic feature response map are received to construct a volume rendering path and a color prediction path. The volume rendering path uses sparse hash coding and a multilayer perceptron for position density fitting, while the color prediction path performs color value regression in the RGB space using fused feature vectors. The volume rendering path outputs point cloud data and surface normal information, calculating the normal direction of each point on the surface using a Gaussian spherical fitting method. The color prediction path outputs a pixel confidence map and a color adjustment map, which are optimized and trained using mean-variance normalization and color residual minimization functions. The color adjustment map reflects the color shift trend under global illumination changes.

[0086] Furthermore, the sparse hash coding structure improves the sampling density of the model in the edge region by introducing a multi-scale sampling strategy, optimizes the computational efficiency and expressive power by dynamically adjusting the number of coding layers and the number of hidden dimensions, and the volume density function and color prediction function are jointly trained by minimizing the weighted total loss function of reconstruction loss, regularization loss and residual loss, and the weight coefficients of the loss function are obtained by fitting a multi-objective optimal function.

[0087] Furthermore, during execution, the fusion of pose feature vectors and weighted depth feature maps can enhance the model's adaptability to pose shifts and depth blurs, the dynamic feature response map can enhance the model's ability to track image temporal changes, and the collaborative optimization of volume rendering path and color prediction path can improve the geometric accuracy and color fidelity of 3D modeling results, achieving a more stable and higher fidelity 3D reconstruction output effect.

[0088] In this embodiment, the 3D modeling result includes point cloud data, surface normal information, pixel confidence map, and color adjustment map, specifically:

[0089] The point cloud data is obtained by jointly encoding the image coordinates and depth values ​​of each pixel in the color image and depth image, and combining the spatial position data and attitude angle data of the projection device to calculate the corresponding three-dimensional spatial coordinate values, forming a three-dimensional coordinate set. Specifically, it includes: performing pixel alignment processing on the color image and depth image to establish a one-to-one pixel mapping table; calculating the initial three-dimensional coordinate values ​​of the pixels based on the pinhole imaging model according to the depth value and corresponding image coordinates of each pixel in the depth image; performing coordinate translation transformation on the initial three-dimensional coordinate values ​​and the spatial position data of the projection device, and performing attitude rotation transformation in combination with the rotation matrices of pitch angle, yaw angle and roll angle in the attitude angle data; completing the three-dimensional coordinate integrated transformation by using a homogeneous coordinate matrix stitching method to generate a global three-dimensional coordinate set; performing a neighborhood weighted smoothing operation on the local spatial point cloud in the global three-dimensional coordinate set, constructing a smoothing weight matrix using spatial distance and depth gradient, and calculating the local continuous surface morphology of the spatial point set by least squares fitting method; the generated point cloud data can completely represent the surface geometry and spatial deformation relationship of the projection area, improving the geometric accuracy and morphological stability of three-dimensional modeling under complex attitude conditions.

[0090] The surface normal information is based on the neighborhood structure of each point in the point cloud data. The normal vector direction of each 3D point is determined according to the spatial gradient calculation rules, and a normal information image corresponding to the size of the color image is output. Specifically, this includes: extracting the neighborhood point set of each 3D point from the point cloud data; calculating the spatial distance between each point and its neighbors based on a set neighborhood radius; using a minimum eigenvalue constrained covariance matrix decomposition method to obtain the normal vector direction by fitting the principal direction vector; performing consistency correction on the normal vector of each point in the neighborhood set to construct a local normal smoothing tensor; remapping the corrected normal vector to the 2D image plane according to the pixel correspondence of the color image, and outputting a normal information image with the same size as the color image; the normal information image reflects the surface tilt direction and local curvature characteristics of each pixel in space, achieving accurate reconstruction and detail preservation of the micromorphology of the projected surface.

[0091] The pixel confidence map is generated by weighted analysis of pixel stability, occlusion discrimination, and temporal consistency features of the depth image, producing a confidence image with the same resolution as the color image. Specifically, this includes: calculating inter-frame depth difference maps for consecutive frame depth images; establishing a pixel stability matrix based on the pixel depth change rate; constructing an occlusion discrimination matrix based on the neighborhood depth variance and occlusion boundary gradient; calculating temporal consistency features and using a sliding time window to calculate the temporal mean squared error of pixel depth; weighting and fusing the pixel stability matrix, occlusion discrimination matrix, and temporal consistency matrix according to weight coefficients, and generating a pixel confidence tensor through normalized fitting; and interpolating and extending the pixel confidence tensor to match the resolution, generating a pixel confidence map with the same resolution as the color image.

[0092] The color adjustment map performs channel correction processing on the color image based on the light intensity and color temperature values ​​in the ambient lighting parameters, combined with attitude angle data, and outputs a color adjustment map after brightness compensation and color equalization. Specifically, it includes: performing a brightness response function transformation on the pixel values ​​of each channel in the color image to construct a light intensity correction matrix; normalizing the light intensity and color temperature values ​​in the ambient lighting parameters and calculating the channel response weights; performing linear correction operations on the RGB channels using the channel weight matrix and obtaining the color equalization factor through a color difference fitting method; adjusting the light direction influence component based on the attitude angle data, performing a tensor multiplication operation between the attitude angle vector and the light intensity matrix to obtain the attitude lighting mapping coefficients; fusing the channel weight matrix, the color equalization factor, and the attitude lighting mapping coefficients to generate a color adjustment matrix; applying the color adjustment matrix to perform pixel-level channel correction and brightness compensation on the color image and outputting the color adjustment map. The color adjustment map can maintain color consistency and brightness balance under conditions of light changes and projection angle shifts, improving the visual fidelity and realism of the projected image.

[0093] In this embodiment, S3 specifically refers to:

[0094] Based on the spatial location data, extract the three-dimensional coordinate values ​​and construct the translation vector;

[0095] Based on the attitude angle data, pitch, yaw, and roll angles are extracted to construct a set of rotation angles. The set of rotation angles is then input into an Euler angle transformation rule to generate corresponding rotation matrices. The Euler angle transformation rule performs a three-axis angle combination transformation according to the spatial rotation order, mapping the pitch, yaw, and roll angles to their respective rotation axes, and calculating the basic rotation matrices around the X, Y, and Z axes. The basic rotation matrices are then multiplied according to a set combination order to form a three-dimensional direction mapping matrix.

[0096] The rotation matrix and translation vector are combined in a matrix concatenation manner to construct a four-dimensional homogeneous transformation matrix; the first three columns of the four-dimensional homogeneous transformation matrix represent the rotation transformation part, the fourth column represents the translation transformation part, and the last row is the homogeneous coordinate supplementary item; the four-dimensional homogeneous transformation matrix is ​​defined as a spatial attitude matrix, which is used to represent the attitude state and spatial displacement relationship of the projection device in three-dimensional space.

[0097] An affine transformation operation is performed on the spatial pose matrix and the image coordinate encoding vector to generate a pose mapping vector; the pose mapping vector represents the pose-related positional relationship of the current image pixel in three-dimensional space.

[0098] The attitude mapping vector is input into the subsequent attitude sensing and control structure and the optimized spiral output structure to improve the adaptability of the modeling results to the pose changes of the projection device, thereby enhancing the geometric consistency and spatial continuity of the modeling process under attitude changes.

[0099] In this embodiment, S4 specifically refers to:

[0100] Extract point cloud data and surface normal information from the 3D modeling results, and obtain the spatial coordinates and corresponding unit normal vector of each 3D point;

[0101] The spatial attitude matrix is ​​applied to the point cloud data to perform a homogeneous transformation operation, which completes the transformation of the three-dimensional coordinates in the target space and obtains the set of surface points in the target space. The homogeneous transformation operation includes performing rotation matrix transformation and translation vector weighting on the point cloud coordinates to maintain the consistency of the spatial geometric structure.

[0102] Projection mapping calculation is performed on the set of surface points. A projection transformation function is constructed based on the camera intrinsic parameter matrix and the image projection model to map the three-dimensional spatial coordinates to two-dimensional image pixel coordinates, generating a pixel mapping relationship from the original image space to the target projection surface; the corresponding coordinate set is output, including the original position and target position coordinate pairs of each pixel.

[0103] Based on the pixel mapping relationship, surface normal information, and pixel confidence map, the angle between the projection vector direction and the viewing angle of each pixel is calculated. It is determined whether the angle exceeds the occlusion angle threshold. If it does, it is marked as an occlusion state, and an occlusion identification state set is constructed. At the same time, the difference vector between the surface normal direction and the viewing angle direction is used to construct a deformation direction vector set to reflect the geometric offset trend of the pixel during the mapping process.

[0104] A color adjustment image is acquired, and color difference calculation and response fitting operation are performed on each pixel in the original image to construct a color response offset matrix. The color response offset matrix is ​​generated by color channel difference fitting, including brightness increment factor and contrast adjustment factor for the RGB three channels respectively, to express the color response change of each pixel in the projection transformation.

[0105] The pixel mapping relationship, deformation direction vector, occlusion status and color response offset matrix are fused at the channel level to construct a four-channel nonlinear mapping tensor; the first channel of the nonlinear mapping tensor is the set of pixel position mapping coordinates, the second channel is the set of deformation direction vectors, the third channel is the set of occlusion masks, and the fourth channel is the set of color correction factors.

[0106] By inputting the nonlinear mapping tensor into the projection image generation path, pixel-level geometric deformation, occlusion compensation, and color adaptive processing are completed, achieving precise alignment and realistic reproduction of the projection image on the target surface, and improving image fitting effect and visual consistency.

[0107] In this embodiment, S5 specifically refers to:

[0108] Based on the pixel mapping coordinates recorded in the nonlinear mapping tensor, a coordinate relocation operation is performed on each pixel of the original projected image to construct a pixel mapping path from the source image space to the target surface space, and a pixel relocation index table containing the original pixel position and the target mapping position is generated. The pixel relocation index table is recorded using a two-dimensional hash mapping structure to ensure that the index access efficiency is maintained in high-density pixel mapping scenarios.

[0109] Based on the geometric deformation vector in the nonlinear mapping tensor, the pixel positions in the redirection index table are subjected to coordinate offset processing, and vector superposition operation is performed to adjust the spatial distribution coordinates of each pixel; the coordinate offset is calculated by the deformation direction vector obtained in the 3D modeling stage, and the projection function constructed based on the normal difference value and the view vector is fitted to obtain the spatial micro-variation coefficient of each pixel point, so as to maintain the consistency of the geometric structure.

[0110] Based on the occlusion mask information in the nonlinear mapping tensor, the set of pixels marked as occluded is extracted, and bilateral interpolation is performed in the local pixel neighborhood. Similarity weighting is performed in both color space and spatial location to construct a continuous and abrupt pixel compensation region. The interpolation compensation process adopts a combination of color gradient guidance and spatial weight fitting to improve the structural consistency and edge clarity of the restored occluded region.

[0111] Based on the color correction factor in the nonlinear mapping tensor, gain adjustment and brightness balancing operations are performed on each pixel channel after retargeting. The color channel compensation is fitted based on a preset response normalization model, and channel correction gain coefficients and brightness offset factors are set for the RGB channels respectively to ensure that pixels in the same surface area maintain color consistency under changes in illumination.

[0112] The pixel distribution results after geometric correction are fused with the pixel set after color compensation at the channel level to construct a unified image structure and generate a continuous two-dimensional image matrix in row and column order. The spatial coordinates, color values ​​and occlusion states of each pixel in the output distortion-compensated image are controlled by the mapping tensor to achieve accurate compensation for image distortion caused by changes in viewing angle, surface curvature and illumination, thereby improving image mapping accuracy and visual fusion effect.

[0113] In this embodiment, the step of acquiring the display image of the distortion-compensated image on the projection surface and constructing the residual image specifically involves:

[0114] Acquire display image data on the projection surface, receive continuous frame color images output by the imaging component, and record the frame sequence corresponding to the distortion compensation image;

[0115] Spatial registration processing is performed on the acquired display image data to align the pixel coordinates of the display image with the pixel coordinates of the distortion compensation image, generating a sequence of registered images;

[0116] Perform difference calculation on each pixel of the registered image sequence, extract the pixel difference matrix between the distortion-compensated image and the display image, and generate the initial residual image;

[0117] High-pass filtering and edge enhancement are performed on the initial residual map to extract structural difference features and output a structural deviation map.

[0118] Calculate the brightness difference matrix and color difference matrix based on the structural deviation map, and construct a pixel-level residual intensity map;

[0119] The pixel difference matrix, structural deviation map, brightness difference matrix, and color difference matrix are stitched together according to the channel dimension to generate a residual image; the residual image includes the pixel difference distribution, structural deviation features, brightness difference distribution, and color difference distribution between the projected surface display image and the distortion compensation image.

[0120] In this embodiment, the calculation of the structural deviation value and edge offset between the residual image and the distorted image, and the updating of the nonlinear mapping tensor, specifically involves:

[0121] The luminance and color channels of the residual and distorted images are extracted. Linear normalization is performed on the luminance channel, and chromaticity normalization is performed on the color channel to generate normalized image pairs. The linear normalization is obtained by performing a minimum-maximum stretch transformation on the pixel grayscale range, and the chromaticity normalization is calculated by normalizing the channel mean and standard deviation to obtain the normalization coefficient.

[0122] Gradient calculation is performed on the standardized image pairs, and the horizontal and vertical gradient components are calculated based on the pixel intensity changes. The gradient components are used to construct an orientation gradient map and an edge intensity map. The orientation gradient map records the gradient direction angle distribution of each pixel, and the edge intensity map records the magnitude of the pixel intensity change rate. The pixel-level edge direction vector set is calculated based on the orientation gradient map, and a statistical model of the direction distribution is generated by fitting the vector angle distribution.

[0123] An edge matching operation is performed based on the edge intensity map and the set of direction vectors. Edge point pairs with similar gradient directions and close positions are searched in the residual image and the distorted image. The spatial offset distance is calculated based on the coordinate difference of the corresponding edge point pairs to generate an edge offset matrix. The spatial offset distance is calculated using the Euclidean distance function. The edge offset matrix records the offset magnitude and direction vector of each matching point pair.

[0124] Correlation analysis is performed on the structural features of the residual image and the distorted image to extract the texture feature matrix within the local window. Cross-correlation is then performed on the texture feature matrix to construct a structural similarity matrix. The element values ​​of the structural similarity matrix are calculated using a similarity fitting function, with the inputs being the mean brightness, standard deviation of contrast, and covariance index of the corresponding window. The average structural deviation value is calculated based on the structural similarity matrix to describe the degree of global structural consistency.

[0125] A deviation weight table is constructed based on the edge offset matrix and the average structural deviation value. The edge offset value and the structural deviation value are normalized and weighted to generate a global structural deviation mapping map. The deviation weight table is calculated by a weight fitting function, which is a non-linear weighted combination. The coefficients are determined by the ratio of the offset magnitude to the similarity gradient.

[0126] Perform weight update operations on the geometric deformation vector and color correction factor in the nonlinear mapping tensor, and adjust the pixel mapping coordinates and color response coefficients according to the corresponding index positions in the deviation weight table; the geometric deformation vector update is performed by weighted vector superposition, and the color correction factor update is performed by response gain compensation mode; output the updated nonlinear mapping tensor.

[0127] In this embodiment, S7 specifically refers to:

[0128] Multiple continuous 3D modeling results are extracted, and spatial coordinate fitting is performed on the point cloud data in each frame. Euclidean distance calculation is performed based on the point cloud fitting results of adjacent frames to generate a point cloud difference matrix. Angle calculation is performed on the surface normal information in adjacent frames using vector angle measurement to generate a set of normal offset angles. Difference calculation is performed on the channel value of each pixel in the color adjustment map to obtain a color change amplitude map.

[0129] Multiple continuous spatial attitude matrices are extracted, and rotation matrices and translation vectors are separated from them. A rotation transformation vector sequence is constructed based on the Euler angle difference in the rotation matrix. A three-dimensional spatial difference calculation is performed based on the translation vector to generate a translation transformation vector sequence. The rotation transformation vector and the translation transformation vector are then dimensionally aligned and concatenated to construct a set of spatial attitude changes.

[0130] Multiple consecutive residual images are extracted, and a weighted moving average is calculated for the structural deviation values ​​in each frame to construct a structural deviation change trend map. Euclidean distance difference is calculated for the corresponding edge points in the edge offset matrix to construct an edge offset change map. Based on the structural deviation change trend map and the edge offset change map, the change gradient value is calculated to construct a set of residual fluctuation indices.

[0131] The point cloud difference matrix, normal offset angle set, color change amplitude map, spatial attitude change set, and residual fluctuation index set are normalized respectively, and weighted fusion is performed using a multi-factor aggregation function to generate a fusion response score map. The multi-factor aggregation function obtains a set of weighted coefficients by fitting the function parameters through the minimum mean square error. The weight parameters correspond to point cloud difference, normal offset, color change, spatial attitude, and residual fluctuation index respectively, and a fusion adjustment factor is constructed.

[0132] Threshold judgment is performed based on the change range of the fusion adjustment factor. If the change range exceeds the preset dynamic adjustment threshold, the output frequency of the 3D reconstruction module, the parameter update frequency of the nonlinear mapping tensor, and the output frame rate of the distortion compensation image are adjusted according to the set mapping relationship.

[0133] This implementation method constructs a dynamic adjustment factor by integrating multi-source variation indicators, and dynamically adjusts the processing frequency of key modules based on the adjustment factor, thereby achieving adaptive allocation of system resources and dynamic maintenance of image quality, effectively improving the robustness and real-time performance of the system in complex scenarios.

[0134] Example 1:

[0135] To verify the feasibility of this invention in practice, it was applied to a high-precision component assembly and inspection scenario in an aerospace parts factory. This scenario requires real-time capture of the three-dimensional morphology of the component surface from multiple angles to detect assembly errors, pose deviations, and component edge misalignments. Traditional inspection methods often rely on fixed vision systems or manual review, which are not only inefficient but also difficult to accurately identify minute errors in complex spatial structures or dynamic assembly processes, leading to high rework rates, frequent misjudgments, and delayed system feedback.

[0136] In this scenario, the self-calibration method for shield assembly error described in this invention is applied. Image structure encoding is performed based on the improved NRI model, and an adjustment factor is generated by the difference between the continuous three-dimensional modeling results, the spatial attitude matrix, and the residual image. This adjustment factor is used to dynamically adjust the three-dimensional reconstruction frequency, the nonlinear mapping update frequency, and the distortion compensation image output frequency, thereby constructing a system-level feedback mechanism, enhancing the sensitivity to structural deviations and attitude changes, and improving the automatic calibration effect and image restoration accuracy.

[0137] This solution deploys a 3D structured light scanning device and a multi-angle camera module on the assembly line, combined with a GPU-accelerated NRI feature encoding network, to acquire continuous images at a rate of 30 frames per second. The system performs difference extraction and weighted normalization calculations on point cloud data, normal information, color adjustment maps, spatial pose matrices, and residual images, outputting a fusion adjustment factor to dynamically adjust the operating frequency of the processing module. The entire deployment cycle was three months, and the operational status is as follows:

[0138] In the first stage, the system uses traditional fixed-frequency image processing as a baseline comparison group without employing a fusion adjustment mechanism. In the second stage, the dynamic frequency adjustment mechanism of this invention is introduced to adjust the frequency of the core module of the model in real time. In the third stage, the range of adjustment factors is refined and an error feedback module is added to improve the system's rapid response capability and adaptability.

[0139] To verify the effectiveness of the method, statistical analysis was performed on typical time periods extracted from three months of test data. The following are the core comparative data regarding 3D reconstruction accuracy, edge error, pose deviation, and system response latency:

[0140] Table 1. Comparison of image reconstruction accuracy and error control between self-calibration systems and traditional methods.

[0141] Comparison items Traditional methods Self-calibration system Self-calibration system + error feedback Average error of 3D reconstruction (mm) 1.28 0.74 0.53 Average edge offset (mm) 0.96 0.51 0.38 Posture recognition deviation angle (°) 3.12 1.65 1.21 Image processing response latency (ms) 142 98 76

[0142] As shown in Table 1, traditional methods exhibit significant deviations in 3D modeling error control, particularly in pose recognition and edge reconstruction, where substantial delays and distortions occur. Since the introduction of the adjustment mechanism described in this invention, the reconstruction error has decreased by 42.2%, edge offset by 46.9%, and system response latency by approximately 31%, indicating that the adjustment factor's control over the model output rhythm effectively improves the accuracy of image structure analysis. Further integration of an error feedback module reduces the 3D reconstruction error to 0.53 mm, and the pose deviation is controlled within 1.2 degrees, demonstrating a significant improvement in overall error control capability. This indicates that the dynamic adjustment mechanism not only optimizes image decoding accuracy but also improves the anomaly feedback response speed, providing a more stable image foundation for assembly quality inspection.

[0143] Table 2. Comparison of System False Positive Rate and Number of Repeat Calibrations within Three Months

[0144] Comparison items Traditional methods Self-calibration system Self-calibration system + error feedback False positive rate (%) 9.3 4.7 2.1 Daily average number of recalibrations 6.2 2.7 1.3 Adjustment factor trigger frequency (times / day) 0 18.5 24.1 Frequency of manual user intervention (times / week) 13 5 2

[0145] As shown in Table 2, without the method of this invention, the system exhibits a high false detection rate and calibration repetition rate in complex assembly structures, and heavily relies on manual intervention. After adopting the method of this invention, the false detection rate decreased by 49.5%, and the average daily number of recalibrations decreased by more than 56%, indicating that the adjustment factor generated by fusing point cloud differences, attitude matrix changes, and residual image fluctuations effectively avoids misjudgments and hysteresis feedback problems at static frequencies. After introducing an error feedback path, the adjustment factor trigger frequency increased to 24.1 times / day, the system became more sensitive to abnormal structures, and the frequency of manual intervention decreased to only twice per week, demonstrating a stable enhancement in self-calibration capabilities.

[0146] This embodiment introduces a fusion adjustment mechanism to transform various changing factors in complex 3D scenes into dynamic control signals, adjusting the image reconstruction rhythm and structural analysis accuracy. This effectively solves the problems of slow feedback, error accumulation, and high misjudgment in dynamic assembly scenarios that exist in traditional methods, significantly improving image processing quality, feedback timeliness, and system intelligence, and verifying the practicality and advancement of this invention in industrial scenarios.

[0147] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A short-focus projection dynamic distortion compensation method based on depth vision modeling, characterized in that, Includes the following steps: S1. Collect multimodal visual information of the projection area to form a visual dataset; the multimodal visual information includes color images, depth images, spatial position data of the projection device, attitude angle data and ambient lighting parameters; S2. Input the visual dataset into the improved Instant-NGP model to generate a 3D modeling result; the 3D modeling result includes point cloud data, surface normal information, pixel confidence map and color adjustment map; the improved Instant-NGP model includes a pose-aware control structure, a depth confidence guidance structure, a dynamic residual compensation structure and an optimized spiral output structure. S3. Construct a spatial attitude matrix based on the spatial position data and attitude angle data; S4. Based on the three-dimensional modeling results and the spatial attitude matrix, generate a nonlinear mapping tensor between the projected image space and the surface space; S5. Perform pixel redirection processing on the original projected image according to the nonlinear mapping tensor to generate a distortion-compensated image; S6. Acquire the display image of the distortion compensation image on the projection surface, construct the residual image, calculate the structural deviation value and edge offset between the residual image and the distortion image, and update the nonlinear mapping tensor. S7. Based on the changing trend of the 3D modeling results, the rate of change of the spatial pose matrix, and the deviation index of the residual image, dynamically adjust the reconstruction frequency, the update frequency of the nonlinear mapping tensor, and the output frequency of the distortion compensation image of the improved Instant-NGP model to form a dynamic compensation closed loop.

2. The short-focus projection dynamic distortion compensation method based on depth vision modeling according to claim 1, characterized in that, Specifically, S1 is: The system acquires color images, depth images, spatial position data of the projection device, attitude angle data, and ambient lighting parameters of the projection area. The color images are acquired through an imaging component that covers the entire projection area and outputs continuous frames of color image data. The depth images are acquired through a depth sampling component that works synchronously with the imaging component, outputting depth image data with the same time frame as the color images. The spatial position data is acquired through a positioning component that outputs three-dimensional coordinate values. The attitude angle data is acquired through an attitude sensing component that outputs attitude data including pitch, yaw, and roll angles. The ambient lighting parameters are acquired through a lighting detection component that outputs light intensity and color temperature values. The color images, depth images, spatial position data, attitude angle data, and ambient lighting parameters are then synchronously processed according to a unified time reference to form a visual dataset.

3. The short-focus projection dynamic distortion compensation method based on depth vision modeling according to claim 1, characterized in that, The improved Instant-NGP model includes an attitude-aware control structure, a depth confidence guidance structure, a dynamic residual compensation structure, and an optimized spiral output structure, specifically: The attitude perception and control structure receives spatial position data and attitude angle data from the projection device, performs attitude vector construction operation, and generates attitude embedding vector; the attitude embedding vector is concatenated with the image coordinate encoding vector to output a fused attitude feature vector. The depth confidence-guided structure receives a depth image and a pixel confidence map, constructs a confidence weight map, and then performs a fusion process with the image feature vector to generate a weighted depth feature map. The dynamic residual compensation structure receives consecutive frames of color images, extracts pixel difference maps and optical flow estimation maps between adjacent image frames, constructs temporal residual feature vectors, and fuses temporal residual feature vectors with weighted depth feature maps to generate dynamic feature response maps. The optimized spiral output structure receives and fuses attitude feature vectors, weighted depth feature maps, and dynamic feature response maps, constructs volume rendering paths and color prediction paths, and outputs 3D modeling results.

4. The short-focus projection dynamic distortion compensation method based on depth vision modeling according to claim 3, characterized in that, The 3D modeling results include point cloud data, surface normal information, pixel confidence maps, and color adjustment maps, specifically: The point cloud data is obtained by jointly encoding the image coordinates and depth values ​​of each pixel in the color image and depth image, and combining the spatial position data and attitude angle data of the projection device to calculate the corresponding three-dimensional spatial coordinate values, thus forming a three-dimensional coordinate set. The surface normal information is based on the neighborhood structure of each point in the point cloud data. The normal vector direction of each three-dimensional point is determined according to the spatial gradient calculation rules, and a normal information image corresponding to the size of the color image is output. The pixel confidence map is generated by weighted analysis based on the pixel stability, occlusion discrimination and temporal consistency features of the depth image, resulting in a confidence image with the same resolution as the color image. The color adjustment map performs channel correction processing on the color image based on the light intensity and color temperature values ​​in the ambient light parameters, combined with the attitude angle data, and outputs a color adjustment map after brightness compensation and color equalization.

5. The short-focus projection dynamic distortion compensation method based on depth vision modeling according to claim 1, characterized in that, Specifically, S3 is: Based on the spatial location data, extract the three-dimensional coordinate values ​​and construct the translation vector; Based on the attitude angle data, pitch angle, yaw angle and roll angle are extracted to construct a set of rotation angles; the set of rotation angles is converted into a rotation matrix, and a three-dimensional orientation mapping relationship is generated using Euler angle transformation rules. The rotation matrix and translation vector are combined in a matrix concatenation manner to form a spatial attitude matrix; the spatial attitude matrix is ​​a four-dimensional homogeneous transformation matrix, which characterizes the attitude change state and spatial displacement relationship of the projection device in the three-dimensional coordinate system.

6. The short-focus projection dynamic distortion compensation method based on depth vision modeling according to claim 1, characterized in that, Specifically, S4 is: Extract point cloud data and surface normal information from the 3D modeling results, combine them with the spatial attitude matrix to perform 3D spatial coordinate transformation, and generate a set of surface points in the target space. Perform projection mapping calculations on the set of surface points to construct a pixel mapping relationship from the original image space to the target projection surface, and output the corresponding coordinate set; Based on pixel mapping relationships, surface normal information, and pixel confidence maps, the deformation direction vector and occlusion status of each pixel are calculated. By combining the color adjustment map, a pixel-level color response offset matrix is ​​established; A nonlinear mapping tensor is constructed by merging pixel mapping relationships, deformation direction vectors, occlusion marker states, and color response offset matrices; the nonlinear mapping tensor includes pixel mapping coordinates between the projected image space and the surface space, geometric deformation vectors, occlusion masks, and color correction factors.

7. The short-focus projection dynamic distortion compensation method based on depth vision modeling according to claim 1, characterized in that, Specifically, S5 is: Based on the pixel mapping coordinates recorded in the nonlinear mapping tensor, a coordinate relocation operation is performed on each pixel of the original projected image to generate a pixel relocation index table. Based on the geometric deformation vector in the nonlinear mapping tensor, the pixel positions in the redirection index table are geometrically corrected to adjust the spatial distribution of pixels. Based on the occlusion mask in the nonlinear mapping tensor, interpolation compensation is performed on the pixels in the occluded region to generate a continuous pixel distribution; Based on the color correction factor in the nonlinear mapping tensor, color channel compensation and brightness balancing are performed on the redirected pixels to generate a set of pixels with consistent color. The geometrically corrected pixel distribution is fused with the color compensation result to output a distortion-compensated image.

8. The short-focus projection dynamic distortion compensation method based on depth vision modeling according to claim 1, characterized in that, The process of acquiring the distortion-compensated image and displaying it on the projection surface to construct the residual image specifically involves: Acquire display image data on the projection surface, receive continuous frame color images output by the imaging component, and record the frame sequence corresponding to the distortion compensation image; Spatial registration processing is performed on the acquired display image data to align the pixel coordinates of the display image with the pixel coordinates of the distortion compensation image, generating a sequence of registered images; Perform difference calculation on each pixel of the registered image sequence, extract the pixel difference matrix between the distortion-compensated image and the display image, and generate the initial residual image; High-pass filtering and edge enhancement are performed on the initial residual map to extract structural difference features and output a structural deviation map. Calculate the brightness difference matrix and color difference matrix based on the structural deviation map, and construct a pixel-level residual intensity map; The pixel difference matrix, structural deviation map, brightness difference matrix, and color difference matrix are stitched together according to the channel dimension to generate a residual image; the residual image includes the pixel difference distribution, structural deviation features, brightness difference distribution, and color difference distribution between the projected surface display image and the distortion compensation image.

9. The short-focus projection dynamic distortion compensation method based on depth vision modeling according to claim 1, characterized in that, The calculation of the structural deviation and edge offset between the residual image and the distorted image, and the updating of the nonlinear mapping tensor, specifically involves: Extract the brightness and color channels of the residual and distorted images, perform normalization on each channel, and generate standardized image pairs; Perform gradient calculations on normalized image pairs to extract edge intensity maps and orientation gradient maps; calculate a set of pixel-level edge orientation vectors based on the orientation gradient maps. Based on the edge intensity map and the set of direction vectors, perform edge correspondence matching to determine corresponding edge point pairs, calculate the spatial offset distance between corresponding edge points, and generate an edge offset matrix. Correlation analysis was performed on the structural features of the residual image and the distorted image to construct a structural similarity matrix; Calculate the average structural deviation value based on the structural similarity matrix; A deviation weight table is constructed based on the edge offset matrix and the average structural deviation value, and a global structural deviation mapping map is generated. Perform a weight update operation on the geometric deformation vector and color correction factor in the nonlinear mapping tensor, adjust the pixel mapping coordinates and color response coefficients according to the deviation weight table, and output the updated nonlinear mapping tensor.

10. The short-focus projection dynamic distortion compensation method based on depth vision modeling according to claim 1, characterized in that, Specifically, S7 is: Extract continuous 3D modeling results and perform difference calculations on point cloud data, surface normal information, and color adjustment maps; Extract the attitude matrix in continuous space and perform difference calculation on the translation vector and rotation matrix; Extract continuous residual images and perform variation analysis on structural deviation values ​​and edge offsets; Based on the differences in 3D modeling results, changes in spatial pose matrix, and deviations in residual images, adjustment factors are generated. The reconstruction frequency, the update frequency of the nonlinear mapping tensor, and the output frequency of the distortion-compensated image are dynamically adjusted based on the adjustment factor.