Real scene graph data processing method and device based on multi-dimensional data dynamic correction

By collecting, evaluating, and fusing multidimensional data, a coordinate transformation mechanism is constructed. Combined with feature matching and contour analysis, efficient processing of real-scene image data is achieved, solving the data acquisition and image fusion problems in existing technologies and improving the accuracy and visual effect of the processing.

CN121582095BActive Publication Date: 2026-05-19BEIJING KEANKE INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING KEANKE INTELLIGENT TECH CO LTD
Filing Date
2026-01-26
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing methods for processing real-scene image data have shortcomings in data acquisition, coordinate transformation, and image fusion, failing to effectively integrate multi-source data and affecting processing accuracy and visual effects.

Method used

Point cloud and real-scene image data are collected using a handheld image acquisition device. Data quality score vectors are calculated, scene classification and data correction modules are constructed, coordinate system transformation parameters are calculated, feature matching matrices are generated, and multi-level image stitching algorithms are used to fuse image data to achieve effective data management and accurate registration.

Benefits of technology

It improves the accuracy and visual effect of real-scene image processing, solves the shortcomings of traditional technologies in data acquisition, coordinate transformation and image fusion, and provides technical support for real-scene image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121582095B_ABST
    Figure CN121582095B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of based on multi-dimensional data dynamic correction real scene graph data processing method and device, through the innovative design data acquisition system, through quality evaluation and time mark, realize the effective management of data.Coordinate conversion mechanism is constructed, combined with feature matching and contour analysis, establish reliable registration model.Smoother transition is introduced, through multi-level splicing and depth fusion, ensure the accuracy of processing.The method effectively solves the deficiency of traditional technology in data acquisition, coordinate conversion and image fusion, etc., provides technical support for real scene graph processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of real-scene images, specifically to a method and apparatus for processing real-scene image data based on dynamic correction of multi-dimensional data. Background Technology

[0002] Existing methods for processing real-world image data have significant shortcomings. Traditional systems perform poorly in data acquisition and quality assessment, failing to effectively integrate multi-source data and impacting processing accuracy.

[0003] Furthermore, existing technologies suffer from bottlenecks in coordinate transformation and feature matching. Most systems lack robust projection transformation mechanisms and correspondence calculation strategies, resulting in suboptimal registration outcomes.

[0004] Existing systems have technical shortcomings in image fusion. They lack in-depth analysis of data changes, making it difficult to achieve smooth transitions and natural image stitching, thus affecting visual quality. Solving these problems is crucial for improving the quality of real-scene image processing. Summary of the Invention

[0005] To address the problems in the existing technology, this application provides a method and apparatus for processing real-scene image data based on dynamic correction of multi-dimensional data, which can effectively solve the shortcomings of traditional technologies in data acquisition, coordinate transformation and image fusion, and provide technical support for real-scene image processing.

[0006] To solve at least one of the above problems, this application provides the following technical solution:

[0007] In a first aspect, this application provides a method for processing real-scene image data based on dynamic correction of multi-dimensional data, including:

[0008] Real-world data is acquired using a handheld image acquisition device, receiving 3D point cloud image data generated by a radar device and 2D real-world image data acquired by a camera. The density distribution and noise level of the point cloud image data are calculated, the clarity and exposure of the real-world image data are evaluated, a data quality score vector is generated, and a unified timestamp identifier is assigned to the point cloud image data and the real-world image data.

[0009] A scene classification module is constructed to perform semantic segmentation on image data and identify scene types and dynamic objects. A data correction module is constructed to calculate the projection transformation parameters between the three-dimensional spatial coordinate system and the two-dimensional plane coordinate system, generate a coordinate system transformation matrix, extract the contour feature points of the point cloud image and the real scene image, calculate the correspondence of the contour feature points based on the iterative nearest point algorithm, and generate a feature matching matrix.

[0010] Analyze the changing characteristics of data at adjacent time points, establish data smooth transition constraints, map the coordinate system transformation matrix to point cloud image data, generate a two-dimensional plane coordinate mapping table, establish the correspondence between feature points based on the feature matching matrix, and use a multi-level image stitching algorithm to fuse the registered image data to generate a real-world image containing depth information.

[0011] Furthermore, it also includes: reading the data stream output by the handheld image acquisition device, constructing a data acquisition module, receiving point cloud image data generated by the radar device scanning, receiving real-scene image data generated by the camera, synchronizing the point cloud image data with the real-scene image data in time, constructing an acquisition record table containing the acquisition time, device identifier, and data type, and writing the acquisition record table into the acquisition database;

[0012] The system reads image data from the acquisition database, constructs a quality assessment module, calculates the point density per unit volume of the point cloud image data, calculates the signal-to-noise ratio parameter of the point cloud image data, generates a point cloud quality score vector, and writes the score vector into the assessment database.

[0013] Furthermore, it also includes: reading real-scene image data, constructing an image evaluation module, calculating the brightness distribution histogram of the image, extracting the gradient feature matrix of the image, calculating the image sharpness index based on the histogram and gradient matrix, statistically analyzing the proportion of bright and dark areas in the image, generating an image quality score vector, and writing the score vector into the evaluation database.

[0014] The scoring vector in the evaluation database is read, a timestamp allocation module is constructed, the clock signal of the hardware device is extracted, a unified time reference sequence is generated, the point cloud image data and the real scene image data are marked with timestamp numbers, an acquisition record table containing data identifiers, time numbers and scoring values ​​is constructed, and the acquisition record table is written into the acquisition database.

[0015] Furthermore, it also includes: reading image data from the acquisition database, constructing a semantic segmentation module, performing regional block processing on the image data, extracting texture and color features of the image, generating scene feature vectors, identifying scene types and dynamic objects based on the feature vectors, constructing a classification record table containing scene labels, object identifiers, and feature attributes, and writing the classification record table into the scene database;

[0016] The classification record table in the scene database is read, a spatial coordinate system transformation module is constructed, a three-dimensional spatial coordinate system is established from point cloud image data, a two-dimensional plane coordinate system is established from real scene image data, projection transformation parameters are selected based on the scene label, a coordinate system transformation matrix is ​​generated, and the transformation matrix is ​​written into the transformation database.

[0017] Furthermore, it also includes: reading point cloud image data and real scene image data, constructing a contour extraction module, calculating the density gradient of the point cloud data, marking the abrupt change position of the density gradient, extracting the edge contour point set, performing edge detection processing on the real scene image, extracting the image contour feature point set, constructing a contour record table containing point set identifiers, spatial positions, and feature attributes, and writing the contour record table into the feature database.

[0018] Read the contour record table in the feature database, construct a feature matching module, calculate the correspondence between two sets of contour points based on the iterative nearest point algorithm, generate a feature matching matrix, combine the transformation matrix and the matching matrix into a correction parameter table, verify the integrity of the correction parameter table, and write the correction parameter table into the correction database.

[0019] Furthermore, it also includes: reading matrix data from the calibration database, constructing a continuity test module, calculating the difference vector of adjacent time data, extracting the change characteristics of the difference vector, generating time series continuity constraints, constructing a continuity record table containing time number, difference value, and constraint parameters, and writing the continuity record table into the test database.

[0020] Read the continuous record table in the inspection database, construct a coordinate mapping module, map the coordinate system transformation matrix to the point cloud image data, calculate the two-dimensional plane projection coordinates, generate a coordinate mapping table, apply smooth transition constraints to the coordinate mapping table, and write the coordinate mapping table into the mapping database.

[0021] Furthermore, it also includes: reading the correspondence table and smooth transition constraints, constructing an image registration module, calculating image transformation parameters based on the correspondence table, applying the transformation parameters to point cloud image data and real scene image data, constructing a registration record table containing registration point pairs, transformation matrix, and registration error, and writing the registration record table into the registration database;

[0022] The registration record table in the registration database is read, an image fusion module is constructed, and a multi-level image stitching algorithm is used to fuse the registered image data to generate a real-scene image containing depth information. The real-scene image is then subjected to integrity verification, and a result record table containing image number, depth data, and fusion parameters is constructed. The result record table is then written into the result database.

[0023] Secondly, this application provides a method for processing real-scene image data based on dynamic correction of multi-dimensional data, including:

[0024] The point cloud image correction module is used to collect real-scene data through a handheld image acquisition device, receive three-dimensional point cloud image data generated by radar scanning, receive two-dimensional real-scene image data generated by a camera, calculate the density distribution and noise level of the point cloud image data, evaluate the clarity and exposure of the real-scene image data, generate a data quality score vector, and assign a unified timestamp identifier to the point cloud image data and the real-scene image data.

[0025] The image feature matching module is used to construct a scene classification module, perform semantic segmentation on image data, identify scene types and dynamic objects, construct a data correction module, calculate the projection transformation parameters between the three-dimensional spatial coordinate system and the two-dimensional plane coordinate system, generate a coordinate system transformation matrix, extract contour feature points from point cloud images and real scene images, calculate the correspondence between the contour feature points based on the iterative nearest point algorithm, and generate a feature matching matrix.

[0026] The real-scene image construction module is used to analyze the changing characteristics of data at adjacent time points, establish data smooth transition constraints, map the coordinate system transformation matrix to point cloud image data, generate a two-dimensional plane coordinate mapping table, establish the correspondence of feature points based on the feature matching matrix, and use a multi-level image stitching algorithm to fuse the registered image data to generate a real-scene image containing depth information.

[0027] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the real-scene image data processing method based on multi-dimensional data dynamic correction.

[0028] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the real-scene image data processing method based on multi-dimensional data dynamic correction.

[0029] Fifthly, this application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the real-scene image data processing method based on multi-dimensional data dynamic correction.

[0030] As can be seen from the above technical solution, this application provides a method and apparatus for processing real-scene image data based on multi-dimensional data dynamic correction. Through innovative data acquisition system design, and by using quality assessment and time stamping, effective data management is achieved. A coordinate transformation mechanism is constructed, and a reliable registration model is established by combining feature matching and contour analysis. A smooth transition is introduced, and multi-level stitching and deep fusion are used to ensure processing accuracy. This method effectively solves the shortcomings of traditional technologies in data acquisition, coordinate transformation, and image fusion, providing technical support for real-scene image processing. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a flowchart illustrating the real-scene image data processing method based on dynamic correction of multi-dimensional data in the embodiments of this application;

[0033] Figure 2 This is a structural diagram of the real-scene image data processing method based on dynamic correction of multi-dimensional data in the embodiments of this application;

[0034] Figure 3 This is a schematic diagram of the structure of the electronic device in the embodiments of this application.

[0035] Figure label:

[0036] Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver storage unit 9144, antenna 9111, speaker 9131, microphone 9132. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0038] The acquisition, storage, use, and processing of data in this application all comply with relevant laws and regulations.

[0039] In view of the problems existing in the prior art, this application provides a method and apparatus for real-scene image data processing based on multi-dimensional data dynamic correction. Through an innovative data acquisition system design, and by using quality assessment and time stamping, effective data management is achieved. A coordinate transformation mechanism is constructed, combined with feature matching and contour analysis, to establish a reliable registration model. A smooth transition is introduced, and through multi-level stitching and deep fusion, the accuracy of the processing is ensured. This method effectively solves the shortcomings of traditional technologies in data acquisition, coordinate transformation, and image fusion, providing technical support for real-scene image processing.

[0040] To effectively address the shortcomings of traditional technologies in data acquisition, coordinate transformation, and image fusion, and to provide technical support for real-scene image processing, this application provides an embodiment of a real-scene image data processing method based on multi-dimensional data dynamic correction. See [link to embodiment]. Figure 1 The real-scene image data processing method based on multi-dimensional data dynamic correction specifically includes the following:

[0041] Step S101: Collect real-scene data through a handheld image acquisition device, receive three-dimensional point cloud image data generated by radar scanning, receive two-dimensional real-scene image data generated by camera acquisition, calculate the density distribution and noise level of the point cloud image data, evaluate the clarity and exposure of the real-scene image data, generate a data quality score vector, and assign a unified timestamp identifier to the point cloud image data and the real-scene image data.

[0042] This embodiment executes S101 for an outdoor pedestrian inspection application scenario. The operator holds a handheld image acquisition device and moves at a constant speed along the park road. The device integrates a lightweight LiDAR and a binocular camera. The LiDAR provides three-dimensional point cloud frames, and the camera outputs two-dimensional real-scene frames.

[0043] To avoid clock drift between different sub-devices, this embodiment first acquires hardware pulses and GNSS / IMU time signals from each device to establish a unified time reference, and then maps the two data streams to the same time axis. Input data is carried in a circular buffer queue. Point cloud frames contain the three-dimensional coordinates of the points and reflection intensity, while image frames contain the original pixel matrix and exposure parameter metadata. To prevent missing frames from causing difficulties in subsequent registration, the buffer adopts a slight lag strategy, prioritizing the pairing of point clouds / images with the closest timestamps into the quality evaluation channel.

[0044] This embodiment calculates the density distribution and noise level of point cloud image data. The process first divides the space into units according to voxel grids, counts the number of points in each unit to obtain the point density per unit volume, and then normalizes the sparsity differences caused by different distances based on the scanning geometry to avoid misjudging structural deviations such as high density at close range and low density at far range as quality problems.

[0045] The noise level is estimated using a joint index of neighborhood plane fitting residual and intensity variance: a fixed radius neighborhood is searched for each point, a local plane is fitted by least squares, and the quantile value of the residual distribution is used to measure geometric jitter, which is combined with the stability of reflection intensity as a material-related noise reference; if the inspection path passes through glass facades or water surfaces, the geometric residual is high and the intensity fluctuation is abnormal, this embodiment marks it as a high-reflection area, and it is deducted in the subsequent weights through scene labels rather than simply judged as low quality.

[0046] The density and noise statistics mentioned above are aggregated in a sliding manner within a time window. The output point cloud quality score vector includes components such as density uniformity, local geometric consistency, intensity stability, and temporal stability. The physical meaning of each component of the vector corresponds to "sufficient spatial coverage", "shape recoverability", "material consistency", and "inter-frame continuity".

[0047] This embodiment evaluates the sharpness and exposure of two-dimensional real-world images. First, a brightness histogram is calculated to observe the accumulation of dark and bright ends and the dynamic range utilization. Second, edge intensity and high-frequency energy are extracted from the gradient feature matrix. A combination of Laplacian energy and Sobel gradient magnitude is used to measure sharpness, avoiding the influence of texture directionality on a single operator. For backlit scenes, the histogram exhibits a bimodal pattern with overflow at the bright end. If the lens metadata indicates a short exposure time, the system determines it as excessive scene contrast rather than focus failure, and the sharpness evaluation does not impose excessive penalties. Motion blur occurring during movement is cross-validated using the consistency of inter-frame optical flow and gradient direction dispersion. If the optical flow direction is consistent with the angular velocity of the device's IMU and its degree is interpretable, the sharpness deduction is reduced, and usable frames are retained to maintain temporal continuity. Finally, an image quality score vector is generated, including exposure balance, detail sharpness, motion interpretability, and noise control. Each component has a direct correspondence with whether the image can be subsequently extracted and registered, conforming to the natural laws of image formation and kinematics.

[0048] This embodiment aligns the point cloud and image scoring vectors at different scales to form a data quality scoring vector. Considering the varying importance of each dimension under different operating conditions, exposure balance is given a higher weight when there is strong outdoor light and shadow; under low-light conditions at night, point cloud intensity stability better reflects usability. The weights are not fixed constants but rather a strategy table that records sensor model, scene label, and weather conditions. To avoid scoring interference from short-term fluctuations, time smoothing is used to gently suppress abnormal frames without excessively delaying the response. The scoring results are written back to the evaluation database, forming a complete record along with the acquisition location, camera pose, and environmental labels, which can be used for subsequent filtering or weighting.

[0049] This embodiment uses a unified timestamp identifier allocation. The device's internal hardware clock establishes a phase-locked relationship with the GNSS PPS to generate a monotonically increasing time reference sequence; point cloud frames and image frames are corrected according to the difference between arrival time and acquisition time to eliminate transmission delay.

[0050] Considering that the integration time of a radar scan frame is not synchronized with the camera exposure, this embodiment adopts a nearest neighbor interpolation strategy: for each image frame, a point cloud scan segment covering its exposure range is found, and the timestamp of this segment is used as the main number. The image and point cloud each retain their original sub-sequence numbers to express an "aligned but not confused" relationship. If packet loss or PPS interruption occurs, the timestamp module triggers a backoff mechanism, using the IMU integration time to fill in the gaps and adding a low-confidence flag to ensure that subsequent modules can still work, but with reduced impact at the weight level.

[0051] Point cloud density and noise are related to the reliability of ICP corresponding points. If no weight is applied at the source, the iteration process will be pulled by outliers, leading to convergence to incorrect local extrema. Image clarity and exposure directly affect the stability of semantic segmentation and edge extraction. Early removal of low-quality frames can prevent the spread of subsequent scene classification errors. For actual inspections, when encountering construction vehicles crossing or crowds gathering, there may be sudden occlusion in the data stream. This embodiment introduces a "scene interpretability" component into the scoring, combining motion interpretability to determine whether the frame can be retained and subsequently removed through dynamic object recognition, rather than rudely discarding the entire time period, ensuring that the temporal continuity constraint still has sufficient effective support.

[0052] This embodiment provides a formula for the composite score expression:

[0053] Q = w1·Dp + w2·Np + w3·Ex + w4·Cl + w5·Mo,

[0054] Where Q is the overall data quality score, Dp represents the point cloud density uniformity, Np represents the point cloud noise controllability, Ex represents the exposure uniformity, Cl represents the image sharpness, Mo represents the motion interpretability, and w1~w5 are the weight coefficients under the current scene strategy. Each term in this expression can be traced back to its source through the physical and statistical measures mentioned above, without involving an empirical black box. Frame pairs with Q exceeding the threshold are marked as priority registration objects, while frame pairs below the lower limit but in continuous segments are labeled "order-preserving but weighted down" to ensure smoothing in subsequent continuity constraints.

[0055] This embodiment considers two typical applications. First, in the narrow alleyways of historical districts, dense metal signs and glass windows result in high noise levels and large geometric residuals in the point cloud reflection intensity. This embodiment relies on high-reflectivity area labeling and temporal smoothing to retain interpretable frames while reducing their weighting. Registration can still be completed in stable areas such as alley entrances and walls. Second, in nighttime park roads, streetlights create bright spots, resulting in a bimodal image histogram with high noise. The system selects frames with higher ISO but acceptable gradient energy based on a comprehensive judgment of exposure balance and sharpness to fill the gaps in the long-distance sparse point cloud. Through the data quality score vector and unified timestamp identifier output by S101, the subsequent semantic segmentation, coordinate transformation, and registration in S102 to S104 have stable and interpretable input prerequisites, reducing the probability of repeated recalculations and failed convergence, and improving the reproducibility of the entire pipeline.

[0056] Step S102: Construct a scene classification module to perform semantic segmentation on image data, identify scene types and dynamic objects, construct a data correction module to calculate the projection transformation parameters between the three-dimensional spatial coordinate system and the two-dimensional plane coordinate system, generate a coordinate system transformation matrix, extract contour feature points from point cloud images and real scene images, calculate the correspondence between the contour feature points based on the iterative nearest point algorithm, and generate a feature matching matrix.

[0057] This embodiment uses mobile surveying and indoor AR navigation as application scenarios, and executes step S102. The prerequisite is that time synchronization and quality scoring of multi-sensor data have been completed in S101. The input includes point cloud frames and real-scene image frames at the same timestamp, as well as initial values ​​of the device's intrinsic and extrinsic parameters. This embodiment first constructs a scene classification module, performing semantic segmentation and dynamic object recognition on the image stream.

[0058] Based on the application's lighting fluctuations and material complexity, a lightweight segmentation network is selected as the backbone. The input consists of the current image frame and the lighting-normalized version of the previous keyframe, and the output is a pixel-level category mask and a dynamic confidence map. Considering that moving vehicles introduce camera self-motion, and that inter-frame differences alone cannot distinguish between "self-moving backgrounds" and "real dynamic objects," the parallax consistency of point cloud sparse projection on the image is incorporated: if a region significantly deviates from the point cloud projection displacement trend in optical flow, it is labeled as a potential dynamic object. The classification results are written into a scene database as "scene labels (roads / factories / corridors, etc.), dynamic object masks, and material and texture summaries," serving as the basis for subsequent projection model selection and feature filtering.

[0059] This embodiment calculates the projection transformation parameters from the 3D spatial coordinate system to the 2D planar coordinate system in the data correction module. The core reasoning is to adaptively select the imaging model of the camera-radar combination based on the scene label: outdoor roads are modeled using a pinhole camera with radial / tangential distortion terms superimposed; indoor corridors are corrected by introducing slight rolling shutter delay. The initial values ​​come from the device calibration and attitude estimation output by S101, but considering the slight changes in intrinsic parameters caused by carrier attitude drift and temperature drift, a secondary solution is required within a small range. To avoid interference from dynamic volumes, this embodiment projects the dynamic regions in the semantic mask onto point cloud coordinates and removes the corresponding points to ensure that the correspondences participating in the solution come from the static background. Robust estimation is used to iteratively correct the projection parameters: with the goal of minimizing the reprojection error of static features, quantile loss is used instead of pure L2 to reduce the impact of a small number of mismatches. The convergence condition is that the error reduction is less than a threshold and the parameter update norm converges.

[0060] This embodiment requires handling the scale difference between sparse point cloud features and dense image features simultaneously at the feature level. To this end, density gradients are first calculated on the point cloud side to identify boundary points with prominent surface normal changes, forming a 3D contour point set. On the image side, a combined Canny and structural tensor approach is used to extract strong edges and sub-pixel corner points, forming a 2D contour feature point set. Since semantic segmentation already provides the material category, image edges of mirror materials and glass regions are prone to introducing spurious features. Therefore, this embodiment reduces the weight or eliminates candidate points with high brightness and unstable gradient directions during image feature filtering. Both ends of the feature set include descriptors: point cloud points carry local curvature, normal, and intensity echo values; image points carry gradient direction and scale. At this point, the projection parameters have initial values, so the 3D contour points can be projected onto the image plane to establish an initial nearest neighbor candidate set, reducing the subsequent matching search space.

[0061] This embodiment employs the Iterative Nearest Point (ICP) approach for the correspondence of contour feature points, but it has been modified to address the "3D-2D cross-modal" characteristics. Standard ICP minimizes the 3D-to-3D distance; here, it is modified to minimize the joint cost of the projected 2D distance and normal consistency: given the current projection matrix, the 3D contour points are projected onto the image plane and matched with the image contour feature points using nearest neighbor matching. During matching, the edge direction difference is required to be less than a threshold, and descriptor similarity is used as the order constraint. After establishing the correspondence, small pose perturbations and micro-intrinsic parameter corrections are applied to reduce reprojection errors. To improve convergence robustness, a layered approach is used: first, sparse features are used to establish approximate alignment at a low-frequency, coarse scale, and then dense edges are used for refinement at high resolution. Dynamic volume regions are masked in each round of matching and do not participate in error accumulation. After convergence, a stable feature matching matrix is ​​obtained, whose elements record the 3D point index, the corresponding 2D point index, the matching confidence, and the local residual.

[0062] Semantic segmentation is not used to provide decorative labels, but rather serves two core subsequent steps: first, selecting an appropriate projection distortion model; and second, cropping dynamic and reflective regions to ensure that parameter estimation revolves around static, structurally strong elements, conforming to the real-world principle that "static geometry dominates coordinate solving." Modifications to ICP ensure that the error term is consistent with the physical quantities observed by the sensor: after the point cloud is projected onto the image, the observation error actually occurs in the image plane. Combined with consistent edge normals, this avoids mismatching irrelevant strong textures to spatial points. Layered matching considers the local extrema problem caused by motion blur and scale aliasing, using a coarse-to-fine approach to avoid getting trapped in local optima.

[0063] In this embodiment, in road scenarios during mobile surveying, ground lane lines and curbs provide stable contours, and the point cloud boundaries correspond clearly to image edges, resulting in rapid convergence of ICP iterations. However, in rainy nights or when there is standing water, specular reflection areas produce false edges on the image. The semantic module uses "water surface / slippery" as a low-confidence material label input to the filter to eliminate most false features. Parameter estimation relies on rigid targets such as curbs and road sign posts to maintain stability. In glass curtain wall corridors for indoor AR navigation, point clouds have weak glass reflection and strong image edges. This embodiment reduces the matching weight of such areas by intersecting the point cloud intensity echo threshold with the image highlight mask. When necessary, combined geometric features such as door frames and baseboards are introduced as anchor points to maintain the recognizability of the projection matrix.

[0064] Step S103: Analyze the changing characteristics of data at adjacent time points, establish data smooth transition constraints, map the coordinate system transformation matrix to point cloud image data, generate a two-dimensional plane coordinate mapping table, establish the correspondence of feature points based on the feature matching matrix, and use a multi-level image stitching algorithm to fuse the registered image data to generate a real-world image containing depth information.

[0065] This embodiment executes S103 for two scenarios: pedestrian inspection and indoor computer room passage. The inputs are the coordinate system transformation matrix, feature matching matrix, and point cloud and image frame pairs with unified timestamps from the previous steps.

[0066] The starting point is the analysis of changes in data at adjacent time points. This embodiment constructs a short-window sequence using frames with consecutive timestamps, calculating two lines: attitude difference and content difference. The attitude difference is given by the pose ΔT_k from IMU integration and radar odometry, while the content difference is extracted from the local plane normal changes of the point cloud and the position vector in the dense optical flow field of the image. If ΔT_k is consistent with the main direction of optical flow, and the point cloud normal changes are concentrated at the foreground boundary, it is determined to be a "change explained by motion". If the two are inconsistent and concentrated in the dynamic class (pedestrian, vehicle, wind-blown leaves) of the semantic segmentation annotation, then the weight of this region is reduced in subsequent constraints to avoid incorrectly including transient objects in the smoothing model. The statistics of change features are output as a sliding window, showing the difference vector and its variance information, as the soft constraint input for smooth transition.

[0067] In this embodiment, when establishing constraints for smooth data transition, a general global regularization is not used. Instead, it is decomposed into two constraints: pose continuity and projection consistency. Pose continuity originates from the motion inertia of the handheld device; pose changes should be smooth within small time steps, denoted as a constraint on ΔT_k. Projection consistency refers to the fact that the position of the same static scene in adjacent frames should not change drastically in the two-dimensional projected coordinate system. To avoid over-smoothing swallowing up the real corners, the constraint weights are adaptively adjusted according to the aforementioned differential variance. When the corner changes abruptly but the optical flow is consistent with the IMU, the pose constraint is relaxed. At locations with high incidence of dynamic occlusion, the robust kernel for projection consistency is increased. In this embodiment, the constraint parameters and time window numbers are written into the continuity record for reading during the mapping and stitching stages, achieving consistency across modules.

[0068] This embodiment maps a coordinate system transformation matrix to point cloud image data, generating a two-dimensional planar coordinate mapping table. Specifically, it reads the 3D-to-2D projection matrix P_k (a perspective or weak perspective model selected by the scene label) from the transformation database, calculates u_k = P_k·X_k for the point cloud X_k at time k using homogeneous coordinates, and obtains the pixel planar coordinates. Considering the time delay in intra-frame scanning by the LiDAR, this embodiment uses linear time interpolation to correct the sampling time of X_k within the frame, aligning the projection with the camera's exposure range; points whose projections fall outside the image are removed, and multiple points with overlapping depths are retained based on the closest depth. The mapping result forms a two-dimensional planar coordinate mapping table, with fields including point ID, pixel coordinates, depth value, and confidence weight (synthesized from the local residual of the point cloud and image quality components), which is subsequently used to guide feature correspondence and stitching occlusion judgment.

[0069] This embodiment establishes the correspondence between feature points based on the feature matching matrix and performs consistency verification with the mapping table. The feature matching matrix comes from the iterative results of ICP on the contour point set and includes cross-frame feature pair indices and matching residuals. To reduce the damage caused by erroneous matches, this embodiment introduces dual verification: first, geometric projection consistency, where the 3D positions of the matching pair in frames k and k+1, after projection by P_k and P_{k+1}, should be close to the pixel position of the feature detection; second, temporal smoothness consistency, where the pixel trajectory of the matching pair in three consecutive frames should be approximately a straight line or a gentle curve, with excessive deviations being weighted lower or eliminated. For regions semantically labeled as dynamic objects, the matching is retained but not involved in the global stitching parameter solution, and is only used for local rendering to avoid "dragging" static structures.

[0070] This embodiment employs a multi-level image stitching algorithm to fuse registered image data. The multi-level approach refers to a dual-channel approach: pyramid resolution and semantic layer. At the pyramid level, residual transformations are estimated progressively from low to high resolution, with lower levels correcting overall misalignment and higher levels correcting detail jitter. At the semantic level, a rigid stitching model is used for rigid areas such as building facades and roads, while a flexible weighted blending model is used for vegetation and water surfaces to reduce texture jitter artifacts. During fusion, occlusion inference is performed using depth values ​​from a mapping table, with foreground elements preferentially covering the background. Seam areas are blended across scales using a Laplacian pyramid to preserve edge sharpness without introducing brightness artifacts. Registration errors are jointly evaluated by feature pair residuals and projection overlap; blocks exceeding a threshold are recalculated at the previous level to maintain overall consistency.

[0071] This embodiment considers the non-constant nature of handheld motion and the impact of lighting changes on color consistency. During the stitching process, local gain compensation is introduced in the color space. The gain parameter is not independently fitted but constrained by the exposure balance and sharpness components derived from S101, preventing underexposed frames from being overstretched. Depth information is generated from the superposition of two paths: one is the geometric depth of point cloud projection, and the other is the parallax depth of binocular or multi-view geometric reconstruction. When dynamic objects and sparse point cloud areas exist, parallax depth is prioritized with point cloud as the regularization term; otherwise, point cloud depth is used as the primary source, and image gradients guide boundary alignment. The two depth sources are fused after confidence weighting, and the output depth map, along with the texture map, constitutes a real-world image containing depth information.

[0072] This embodiment provides a simplified expression:

[0073] E = α·E_motion + β·E_proj + γ·E_photo,

[0074] Where E is the stitching energy function, E_motion represents the cost of pose continuity, derived from the deviation of ΔT_k; E_proj represents the projection consistency cost, measuring the overlap error of feature points in adjacent frames after projection; E_photo represents the photometric consistency cost, measuring the difference in brightness and gradient in the seam region; α, β, and γ are constraint weight coefficients, determined by scene labels and quality scores. The physical meanings of the three terms correspond to the natural laws of kinematics, geometric projection, and imaging processes, respectively. The adaptive selection of weights is derived from the statistics of the preceding modules, avoiding a black box approach.

[0075] In this embodiment, in the narrow environment of the computer room corridor, the point cloud reflects complexly on the surface of the metal cabinet, resulting in low depth confidence in the mapping table. The system automatically reduces γ and relies on more feature overlap and the overall constraints of the low-level pyramid to complete the stitching. In the scene of dappled shade on the park road, dynamic leaf areas are marked with low weight, and parallax depth has a higher weight in these areas. In the final output real-world image, the road edges and building polylines remain continuous, and the leaf texture does not produce obvious stretching at the seams. Through the processing of S103, subsequent downstream tasks for 3D reconstruction or surveying annotation can directly utilize the real-world image with depth, reducing the workload of repeated registration and manual repair, making the process more stable and the backtracking chain clearer.

[0076] As described above, the real-scene image data processing method based on multi-dimensional data dynamic correction provided in this application can achieve effective data management through innovative data acquisition system design, quality assessment, and time stamping. It constructs a coordinate transformation mechanism, combining feature matching and contour analysis to establish a reliable registration model. By introducing smooth transitions and employing multi-level stitching and deep fusion, it ensures processing accuracy. This method effectively addresses the shortcomings of traditional technologies in data acquisition, coordinate transformation, and image fusion, providing technical support for real-scene image processing.

[0077] In one embodiment of the real-scene image data processing method based on multi-dimensional data dynamic correction in this application, it may further include the following:

[0078] Step S201: Read the data stream output by the handheld image acquisition device, construct a data acquisition module, receive point cloud image data generated by the radar device scanning, receive real scene image data generated by the camera, synchronize the point cloud image data and the real scene image data in time, construct an acquisition record table containing acquisition time, device identifier, and data type, and write the acquisition record table into the acquisition database.

[0079] Step S202: Read the image data from the acquisition database, construct a quality assessment module, calculate the point density per unit volume of the point cloud image data, calculate the signal-to-noise ratio parameter of the point cloud image data, generate a point cloud quality score vector, and write the score vector into the assessment database.

[0080] This embodiment uses a typical scenario of walking inspection of a park and mapping of factory corridors, and executes S201 and S202. The data source comes from a handheld image acquisition device, which has a built-in lightweight LiDAR and a binocular camera, and is externally connected to an IMU and a PPS timing unit.

[0081] To ensure the reliability of subsequent registration and quality assessment, this embodiment first constructs a data acquisition module at the acquisition end to uniformly access and cache multi-source data streams. The radar channel writes data to a circular buffer, with scan lines as the smallest unit, containing 3D coordinates, reflection intensity, and scan line ID; the camera channel writes the original pixel array and exposure metadata; the IMU outputs angular velocity and acceleration at a fixed frequency. Different sensors have acquisition and transmission delay differences. This embodiment does not directly use arrival time for alignment, but instead generates a time reference sequence based on the hardware PPS and the device's local clock difference model, appending the acquisition time and clock offset to each frame as input for subsequent time synchronization.

[0082] The time synchronization in this embodiment follows the principle of "inter-frame interpretability and inter-frame traceability." Specifically, the process first estimates the linear drift of each device's local clock based on the PPS (Pressure Point Size) to obtain short-term correction parameters. Then, it aligns the radar's scanning period with the camera's exposure range: when the camera's exposure range is crossed by the radar scan, the radar point cloud is linearly interpolated into a subframe within the frame according to the timestamp, ensuring it aligns with the center time of the image frame. Considering the possibility of PPS loss or IMU saturation during inspection, this embodiment sets a fallback path in the synchronization module, using the drift parameters estimated by the previous valid window for short-term maintenance, and labels the relevant frames with "synchronization confidence" to prevent subsequent modules from treating them with equal weight to normal frames. After alignment, the acquisition record table is constructed in entry form, with fields including acquisition time t, device ID, data type (point cloud / image / IMU), original clock value, corrected time number, and confidence label. This data is written to the acquisition database for cross-module retrieval and auditing.

[0083] In this embodiment, point cloud data and paired time-series segments are read from the acquisition database in S202 to construct a quality assessment module. The goal is to obtain quantitative indicators that reflect spatial coverage and geometric stability. The statistical analysis of point density per unit volume is not a simple count, but rather employs a method of voxel grid overlay and distance compensation: first, the three-dimensional space is divided using an adaptive voxel scale, determined based on the expected resolution overlapping with the camera's field of view, avoiding excessive fineness that could lead to sparsity misjudgments; then, a normalization coefficient of the squared distance is applied to voxels at different ranges to offset the natural sparsity caused by radar geometric divergence. For multi-echo or high-reflectivity surfaces, this embodiment uses peak intensity and number of echoes as additional weighting factors, excluding low-confidence echoes from the effective density to reduce the artificial density overestimation caused by glass, water surfaces, etc.

[0084] This embodiment combines two chains of evidence—geometric and strength—to calculate the signal-to-noise ratio parameter. Considering that the inspection path often contains reflective materials and thin rods, the single strength stability is insufficient to characterize the noise.

[0085] This embodiment performs local plane fitting on the fixed-radius neighborhood of a point, and uses the robust quantile of the residual to measure geometric noise; it also calculates the coefficient of variation of the neighborhood reflection intensity sequence as a quantification of material-related noise. These two methods are not simply summed, but rather assigned different weights based on scene labels: in outdoor shady road sections, strong winds cause leaf vibrations, leading to increased geometric residuals, but their semantics are marked as dynamic, so the weight is reduced to avoid damaging static structures; in densely packed metal cabinet areas of a server room, the intensity flickers but the geometry is flat, so this embodiment increases the weight of the geometric component to ensure that the score is consistent with rebuildability. The evaluation module smooths the above indicators within a short time window, reducing instability caused by occasional vibrations without slowing down the response to sudden changes in the real environment.

[0086] This embodiment maps the unit volume point density and signal-to-noise ratio parameter to a point cloud quality score vector. The vector components include: density coverage (average density after voxel occupancy and distance normalization), geometric consistency (robust statistics of local plane residuals), intensity stability (inverse quantization of the neighborhood intensity variation coefficient), and temporal stability (a measure of consistency between adjacent frame metrics). The physical meaning of the vector directly relates to the usability of subsequent modules: density coverage relates to the sufficiency of feature extraction and projection; geometric consistency corresponds to ICP convergence; intensity stability is related to material-related edge recognizability; and temporal stability relates to the reliability of the continuity term in the splicing energy function. For ease of implementation, a simplified synthetic expression is used:

[0087] Q_pc = a·D + b·G + c·I + d·T,

[0088] Where Q_pc is the overall point cloud quality score, D represents density coverage, G represents geometric consistency, I represents intensity stability, T represents temporal stability, and a, b, c, and d are weighting coefficients under the current scene and device configuration, derived from the aforementioned strategy table. Each parameter has a clear meaning, and the corresponding original statistics are traceable, avoiding a black-box approach.

[0089] This embodiment retains the foreign key association with the acquisition record table when writing to the evaluation database and records meta-parameters such as scoring version, voxel scale, neighborhood radius, and time window length to ensure repeatability across periods. To prevent scoring fluctuations from misleading scheduling, the system sets a threshold band and hysteresis interval: when Q_pc is in the gray area, it is marked as "order preservation and weight reduction," and the decision is made by the subsequent scene classification in S301 and coordinate transformation in S401, rather than being immediately removed. For abnormally high density values, the system triggers reflection anomaly verification, checks whether the proportion of multiple echoes and close-range density are abnormal, and if necessary, reverts to the original point cloud for secondary filtering.

[0090] This embodiment provides two real-world examples for verification. At the glass corridor in the park, the point cloud generates strong echoes and virtual images in front of the glass, nearly doubling the original density. However, the intensity stability and geometric consistency decrease simultaneously, with Q_pc remaining moderate and falling into the gray area. It is labeled for weight reduction and order preservation, and subsequently, during stitching, the pseudo-structure is naturally suppressed through projection depth and occlusion relationships. In the server room corridor, dense server racks form a regular plane with high geometric consistency, but the density coverage is slightly lower in the center of the corridor. The system suggests slowing down at turns to increase voxel occupancy. This strategy is executed via data transmission from the acquisition end, requiring no manual intervention. Through the collaboration of S201 and S202, this embodiment solves three types of problems: asynchronous multi-source data, difficulty in quantifying point cloud quality, and complex scene interference. This ensures that the inputs entering subsequent semantic segmentation, feature matching, and projection stitching have both temporal consistency and quality labels, ultimately reducing unnecessary iterations and failed convergences in the imaging chain, making the process more stable and easier to backtrack.

[0091] In one embodiment of the real-scene image data processing method based on multi-dimensional data dynamic correction in this application, it may further include the following:

[0092] Step S301: Read real-scene image data, construct an image evaluation module, calculate the brightness distribution histogram of the image, extract the gradient feature matrix of the image, calculate the image sharpness index based on the histogram and gradient matrix, count the proportion of bright and dark areas of the image, generate an image quality score vector, and write the score vector into the evaluation database.

[0093] Step S302: Read the scoring vector in the evaluation database, construct a timestamp allocation module, extract the clock signal of the hardware device, generate a unified time reference sequence, mark the point cloud image data and the real scene image data with timestamp numbers, construct an acquisition record table containing data identifiers, time numbers, and scoring values, and write the acquisition record table into the acquisition database.

[0094] This embodiment executes S301 and S302 for two scenarios: pedestrian inspection of the park and indoor corridor mapping. The input is the original real-scene image frames and their acquired metadata already collected in the database in S201. The goal is to form an interpretable image quality score vector and establish a time numbering system that is consistent with the point cloud without relying on manual parameter tuning.

[0095] The initial image evaluation module performs linear gamma removal and illumination normalization on each frame to prevent local saturation caused by automatic camera exposure from obscuring structural details. Then, a brightness distribution histogram is calculated, and binning employs an adaptive equal-frequency strategy to stably reflect the proportion changes of dark, bright, and mid-gray areas. Based on this, a gradient feature matrix is ​​extracted, with channels containing Sobel horizontal / vertical gradients and Laplacian responses. A gradient direction histogram is also stored for later use in identifying motion blur and defocus. Considering the sudden changes in lighting during inspections, this embodiment does not directly use full-frame statistics but instead employs gridded local statistics and summarizes quantile values ​​to avoid a small number of bright signs or reflective ground surfaces dominating the global evaluation.

[0096] This embodiment provides a clear logical deduction for constructing the sharpness index. Defocusing causes high-frequency energy attenuation, while motion blur causes gradient direction concentration. This embodiment uses Laplacian energy to measure detail intensity and gradient direction entropy to characterize whether unidirectional ghosting occurs, supplemented by edge width estimation correction. If the image is in a backlit or high-contrast environment, the brightness histogram will show a double peak with the bright end at the top. In such cases, the deduction in sharpness score needs to be offset by exposure interpretability to avoid misinterpreting the exposure strategy as defocusing. When calculating the proportion of bright and dark areas, two paths are used for cross-validation: the Otsu threshold reflects the overall distribution, while the latter focuses on shadow blocks. If the difference between the two is too large, it is marked as "uneven illumination," prompting an increase in the weight of the luminance consistency constraint in the subsequent stitching stage.

[0097] This embodiment forms an image quality scoring vector and clarifies the correspondence between each component and the downstream physical process. The scoring vector q_img contains four principal components: exposure equalization E (derived from histogram peak shape, dark end overflow, and saturation ratio), detail sharpness C (normalized value of Laplacian energy after noise suppression), motion interpretability M (mapped from gradient direction entropy and IMU angular velocity consistency, reflecting whether blur is consistent with self-motion), and illumination uniformity U (given by the spatial variance of local threshold results). Among them, E corresponds to the contrast availability of subsequent feature extraction, C affects edge and corner stability, M determines whether to use stronger temporal smoothing, and U indicates the photometric correction intensity required during stitching. To reduce black-box color, each component of q_img can be traced back to specific statistics and physical interpretations, and the calculation window, filtering parameters, and camera exposure metadata are included when writing to the library.

[0098] In this embodiment, a robustness check is performed before the scoring vector is written into the evaluation database.

[0099] The checks include: the q_img difference of adjacent images in the same keyframe should not cross the threshold band in a short period of time; if a sudden change occurs, the original metadata needs to be checked for exposure strategy switching or sensitivity transitions. If confirmed, the change is retained and a "switching exposure" label is recorded; if there is no supporting metadata, it is marked as "abnormal fluctuation" and the acquisition end is triggered to reduce its movement speed for a short period of time. In low-light noise scenes, the Laplacian response is inflated by noise. This embodiment introduces a noise control factor (estimated from high-frequency noise power) in the C component calculation to weaken the artificially inflated effect of noise on sharpness and maintain consistency with the actual resolution.

[0100] This embodiment enters S302 to construct the timestamp allocation module. After reading the scoring vector from the evaluation database, the clock signal of the hardware device (PPS pulse and local clock) is extracted and combined with the clock offset model already saved in the acquisition record to generate a unified time base sequence τ. Considering that the radar point cloud is a scan integral and the camera is an instantaneous exposure, τ is not a simple integer frame number, but a strictly monotonic sequence on the physical time axis, with the unit being microseconds. For each image frame and the point cloud subframes that are temporally adjacent to it, a timestamp number n is marked, which is defined as the index on τ that is closest to the acquisition center time of that frame; to preserve alignment accuracy, the relative displacement parameter of the scan line is further preserved within a point cloud frame and written as a sub-time offset δ, while the image records the exposure start and end offsets, forming a two-level time representation of "number n + sub-offset".

[0101] This embodiment binds the score value with the data identifier while marking the timestamp number, constructing a collection record table and writing it into the collection database. The record table fields include: data identifier (unique frame ID and source device ID), time number n, sub-offset δ, score vector q_img or Q_pc (if it is a point cloud entry), synchronization confidence, and scene label summary. The reason for binding the score in S302 is that in the subsequent S103, when establishing smooth transition and photometric consistency constraints, it is necessary to assign weights to different frames according to quality. If time and quality are stored separately, cross-database aggregation will introduce unnecessary delay and consistency risks. For short-term PPS loss or clock jitter, this embodiment extrapolates τ from the drift estimate of the previous steady-state window and marks the time number of that period as low confidence. The subsequent splicing energy function automatically lowers its weight, which conforms to the natural principle of "weak evidence has little impact".

[0102] In this embodiment, during the nighttime road mapping in the park, the histogram shows a clear bimodal distribution, with a low U component, a medium-high E component, and a slightly high C component due to noise, which is maintained at a medium level after noise correction. The system labels this frame segment as "uneven illumination" in the acquisition record table, and S103 increases the weight of the photometric consistency item and uses stronger seam blending. Under strong white light in the computer room corridor, the histogram is concentrated in the highlights and lacks mid-gray. If the exposure metadata indicates a shortened shutter speed and increased gain, and the M and IMU directions are consistent, it is judged as "high gain interpretable," allowing more frames to be retained for stitching to maintain temporal continuity. Through the connection between S301 and S302, this embodiment solves three types of problems: difficulty in quantifying image quality, difficulty in interpreting exposure changes, and inconsistent time numbering among multiple sources. The output score and unified numbering provide a traceable weighting basis for subsequent coordinate mapping and multi-level stitching, reducing recalculation and convergence failures caused by misjudgment.

[0103] In one embodiment of the real-scene image data processing method based on multi-dimensional data dynamic correction in this application, it may further include the following:

[0104] Step S401: Read the image data from the acquisition database, construct a semantic segmentation module, perform regional block processing on the image data, extract the texture and color features of the image, generate scene feature vectors, identify scene types and dynamic objects based on the feature vectors, construct a classification record table containing scene labels, object identifiers, and feature attributes, and write the classification record table into the scene database.

[0105] Step S402: Read the classification record table in the scene database, construct a spatial coordinate system transformation module, establish a three-dimensional spatial coordinate system from point cloud image data, establish a two-dimensional plane coordinate system from real scene image data, select projection transformation parameters based on the scene label, generate a coordinate system transformation matrix, and write the transformation matrix into the transformation database.

[0106] This embodiment executes S401 and S402 for two scenarios: pedestrian inspection in the park and indoor server room corridor. The input consists of raw real-scene images from the acquisition database and the timestamp and quality score written in S302. The goal is to identify scenes and dynamic objects at the image layer and constrain the 3D-to-2D coordinate system transformation. The processing begins with the construction of the semantic segmentation module. Considering the significant changes in lighting and materials during inspections, direct end-to-end segmentation is susceptible to domain drift. This embodiment first divides the image into regions, cutting the entire image into several sub-blocks according to an overlapping grid. Texture and color features are calculated independently for each block. Texture features consist of multi-scale Gabor energy, the ratio of structure tensor eigenvalues, and local binary pattern histograms, corresponding to repetitive textures (brick walls, grilles), anisotropic strong edges (road edges, cabinet edges), and fine textures (grass, fabric), respectively. Color features are calculated using the mean and variance in the CIE Lab space, and a hue peak distribution is established to distinguish stable color groups such as road gray, vegetation green, and warning yellow. To improve robustness under low light and backlight conditions, the above features are all localized for contrast under the guidance of the S301 illuminance uniformity U index to avoid misjudgment caused by feature attenuation in dark areas.

[0107] This embodiment concatenates sub-block-level features into a scene feature vector and introduces a lightweight model for classification, but the decision boundary is not determined solely by the model. The model input is a vector concatenated from sub-block textures and colors, and the output is the scene category (roads, walls, glass, vegetation, server racks, ground markings, etc.) and dynamic volume confidence (pedestrians, vehicles, swaying leaves). Simultaneously, it reads the motion interpretability M from S101 and the dynamic mask from S102 to impose consistency constraints on the model output: if the optical flow is consistent with the IMU and the segmentation is determined to be dynamic, it is downgraded to "interpretable dynamic" and directly masked in subsequent coordinate calculations; if a high-confidence glass label appears and the point cloud intensity stability is low, it is further confirmed as a reflective area to avoid treating specular reflection edges as real structures. The classification results are organized into a classification record table, including scene labels, object identifiers, feature attribute summaries (dominant texture scale, hue peak, anisotropy index), time number, and confidence level, and written to the scene database to provide a basis for projection model selection and static feature screening.

[0108] This embodiment adheres to a two-layer structure of region segmentation and feature-semantic mapping because subsequent coordinate system transformations rely on interpretable geometric priors. Rigid regions such as roads and walls are suitable for pinhole models with small distortions, while glass and water surfaces are prone to introducing reflection pseudopoints, requiring additional occlusion and confidence processing during projection. Although vegetation textures are strong, their irregular geometric shapes can easily skew feature matching. This embodiment includes a "geometric utilization strategy" for each category in the classification record table. For example, walls / cabinet surfaces are set as rigid anchor points, road markings allow sub-pixel edges to participate, and vegetation is only used for photometric fusion and not for geometric calculation. These strategies will be read and mapped into solution weights in S402.

[0109] This embodiment enters the S402 spatial coordinate system transformation module, first constructing a three-dimensional spatial coordinate system from the point cloud, and then constructing a two-dimensional planar coordinate system from the image. On the three-dimensional side, the radar coordinate system is used as the parent system. Based on the initial extrinsic parameters in S201 and the IMU attitude interpolated on time number n, the pose aligned to the time is obtained. The original point cloud is then cleared of points corresponding to dynamic volumes and candidate points of highly reflective mirrors labeled by classification records, retaining geometrically stable structural point clouds. On the two-dimensional side, the image plane coordinates are defined using camera intrinsic parameters and distortion parameters. Photometric normalization is performed based on the exposure parameters in S301 to ensure comparability of subsequent reprojection errors under uneven brightness conditions. The purpose of establishing two sets of coordinate systems is not simple mapping, but rather to provide a stable foundation for selecting appropriate projection transformation parameters for different scene labels.

[0110] This embodiment adaptively selects the projection model and parameter solution path based on scene labels. Road and wall scenes use perspective projection π_p, with intrinsic parameters including focal length and principal point. Distortion is handled by second-order radial projection plus one-order tangential projection. Slight rolling shutter effect in the corridor introduces time-dependent line offsets, which are corrected using line delay terms. Glass / water surface labels are guided into "low-weighted projection," used only for occlusion inference and not contributing to geometric constraints. The parameter solution aims to minimize the reprojection error of static anchor points, using the contour feature correspondence output by S102 as observations. A robust kernel function suppresses the long tail of residuals, avoiding minor mismatches that determine extrinsic parameter corrections. To avoid ill-posedness due to high coupling between intrinsic and extrinsic parameters, this embodiment employs a block-based solution: first, the pose is fixed to estimate a small amount of intrinsic parameter drift; then, the intrinsic parameters are fixed to refine the extrinsic parameters, alternating several times until the solution stops. The convergence criterion is determined by the error reduction rate and the parameter step size.

[0111] This embodiment adds a consistency check before generating the coordinate system transformation matrix.

[0112] Verification 1: Cross-category residual differences. If the residuals of rigid body classes (walls, cabinets) are significantly smaller than those of non-rigid body classes, it indicates that the model selection is reasonable. If the opposite is true, check for timestamp offsets or uncorrected rolling shutter. If necessary, backtrack and adjust the line delay parameters.

[0113] Verification 2: Spatial Coverage. The density coverage D of S202 is read. If the spatial distribution involved in the solution is too concentrated, the extrapolation risk of this matrix is ​​marked, prompting a suggestion in the S103 stitching stage to relax the photometric consistency threshold in areas far from the anchor point. Parameters that pass verification are encapsulated into a coordinate system transformation matrix P_n, along with an applicable scenario label set and confidence level, and written to the transformation database.

[0114] Understandably, the purpose of parallel semantic segmentation and feature statistics is not to pursue more refined pixel classification, but to separate "usable geometry" from "visible texture," which aligns with the inherent differences between imaging and geometric observation. The adaptive selection of projection parameters follows the physical laws of the sensor and environment: the pinhole model is sufficient to describe most rigid body scenes; rolling shutter speeds are geometric deformations caused by line delay, which must be modeled when the device moves with angular velocity; specular reflections lack geometric stability, and reducing their weight avoids dragging virtual images into the coordinate solution. Block-based parameter solving aims to reduce the mutual compensation effect between intrinsic and extrinsic parameters, commonly seen in cases of slight temperature drift and minor focal length changes. Fine-tuning the intrinsic parameters first and then refining the extrinsic parameters is more stable than solving large variables simultaneously at once.

[0115] This embodiment demonstrates a differentiated strategy in two landing scenarios. During the daytime in the park, walls and lane lines provide high-contrast stable edges, which the classification module marks as rigid anchor points. The transformation matrix is ​​dominated by these features. The glass window area participates in occlusion judgment but not in geometry, ultimately resulting in stable convergence of P_n within the window. At night, in the server room corridor, strong white light causes brightness overflow. While the image sharpness index C is acceptable, the illuminance uniformity U is low. The module performs local contrast balancing on the 2D side and relies on the high flatness of the server rack facade on the 3D side. The rolling shutter correction term takes effect during rapid oscillations, and parameter solving prioritizes ensuring consistent reprojection in the depth direction. In both scenarios, the transformation matrix is ​​written to the transformation database with version and evidence indexes. After reading by S103, coordinate mapping and multi-layer stitching can be performed directly without repeated discrimination and parameter tuning.

[0116] The technical effects of this embodiment can be summarized in three points. First, scene semantics and texture-color joint features enable dynamic objects and specular reflections to be identified at the image layer and reflected in the geometric utilization strategy, reducing subsequent parameter offsets. Second, the adaptive and block-based solution of the projection model based on scene labels suppresses the coupling of intrinsic and extrinsic parameters and temporal distortion, resulting in a stable and interpretable coordinate system transformation matrix. Third, both the classification records and the transformation matrix are equipped with confidence and version fingerprints, facilitating cross-time period verification and reconstruction reproduction, providing a clean geometric premise for subsequent deep fusion and stitching.

[0117] In one embodiment of the real-scene image data processing method based on multi-dimensional data dynamic correction in this application, it may further include the following:

[0118] Step S501: Read point cloud image data and real scene image data, construct a contour extraction module, calculate the density gradient of the point cloud data, mark the abrupt change position of the density gradient, extract the edge contour point set, perform edge detection processing on the real scene image, extract the image contour feature point set, construct a contour record table containing point set identifier, spatial location, and feature attributes, and write the contour record table into the feature database.

[0119] Step S502: Read the contour record table in the feature database, construct a feature matching module, calculate the correspondence between two sets of contour points based on the iterative nearest point algorithm, generate a feature matching matrix, combine the transformation matrix and the matching matrix into a correction parameter table, verify the integrity of the correction parameter table, and write the correction parameter table into the correction database.

[0120] This embodiment executes S501 and S502 for two scenarios: pedestrian inspection of the park and mapping of the computer room corridor. The inputs are the original point cloud and real-world image in S201, the scene labels and dynamic mask in S401, the initial value of the coordinate system transformation matrix in S402, and the image quality score in S301. To avoid interference from dynamic objects and reflective areas in feature extraction, this embodiment first removes low-confidence specular surfaces and strongly reflective areas based on the scene labels before proceeding to contour extraction. Pixels marked as dynamic and their projection neighborhoods in the point cloud are weighted less, ensuring that subsequent contour extraction focuses more on static geometry.

[0121] This embodiment constructs the density gradient calculation process on the point cloud side. First, a regularized mesh is established in the point cloud using adaptive voxelization. For each voxel, the number of points *n* and the weighted density *d* (considering distance compensation and echo confidence) are statistically analyzed. A discrete approximation of the 3D gradient ∇d is performed on *d* on the mesh. Locations with large gradient magnitudes often correspond to object boundaries or line-of-sight occlusion edges. To suppress noise spikes, a dual-threshold hysteresis strategy is used to mark locations of abrupt density gradient changes: high thresholds are directly identified as boundary candidates, and voxels between low and high thresholds are included if they are connected to strong boundaries, avoiding breaks in the gradient. For points within candidate voxels, local PCA is used to obtain the normal and curvature. Isolated points with abnormal curvature but discontinuous density gradients are removed, forming a stable 3D edge contour point set P3D. For each point, spatial location, normal, local curvature, and intensity statistics are recorded as characteristic attributes.

[0122] This embodiment performs edge detection and sub-pixel feature extraction on the image side. The image is first locally contrast-stretched based on the exposure equalization and illumination uniformity of S301, then Canny detection is performed to obtain the main edge map. Sub-pixel corner points are extracted at locations with high corner response using the structure tensor. For areas labeled as glass, water, or strongly highlighted, edge points with large gradient direction fluctuations are downweighted or filtered. To improve cross-scale robustness, a pyramid-shaped multi-layer feature set (P2D) is constructed. The lower layers retain stable long edges, while the higher layers retain detailed corner points. Each feature point is accompanied by a histogram descriptor for direction, scale, and local gradient. Considering the characteristics of edge direction concentration and amplitude decrease during motion blur, direction entropy and the motion interpretability of S301 are cross-judged to retain strong structural edges under interpretable blur and reduce sparsity caused by excessive cleaning.

[0123] In this embodiment, P3D and P2D are organized into a contour record table, with fields including point set identifier (3D / 2D), time number, spatial location (XYZ coordinates of 3D points or pixel coordinates of 2D points), feature attributes (normal, curvature, intensity, orientation, scale, descriptor fingerprint), scene label, and confidence weight.

[0124] Before the record table is written into the feature database, it undergoes two consistency checks: first, temporal consistency, the number of features and the average position of the same spatial region across adjacent frames should not change significantly; second, semantic consistency, regions marked as dynamic objects should not have high-confidence static anchor points, if they do, the mask boundary of S401 is checked and corrected.

[0125] This embodiment enters the feature matching module S502, employing an iterative nearest-neighbor approach for 3D-2D cross-modal mapping. First, using the projection matrix P given in S402, the 3D contour points are projected onto the image plane. An initial nearest-neighbor set is established in P2D based on pixel distance and orientation consistency. Orientation consistency is determined by the angle between the normal of the projected edge and the image edge direction; those exceeding a threshold are discarded. The matching cost function jointly considers reprojection distance and normal consistency. Matching of the dynamic mask region is only retained for rendering and does not participate in the solution. To avoid local extrema, this embodiment performs the process from top to bottom at the pyramid scale. The coarse scale estimates large pose deviations, the fine scale refines local edge registration, and robust loss is used to mitigate long-tail errors.

[0126] This embodiment provides a structured description of the solution variables and constraints for ICP. The variables include minute extrinsic perturbations and small intrinsic parameter drifts, while the constraints are matched pairs from the static region.

[0127] Each iteration consists of two steps: first, solving for the incremental extrinsic parameters with fixed intrinsic parameters to reduce the projection residual; then, fixing the extrinsic parameters again to fine-tune any possible minor changes in focal length. Weights are assigned according to a scene-specific strategy during the solution process: rigid anchor points such as walls and server racks have the highest weight, followed by road markings, and vegetation has the lowest weight; areas with strong reflections only participate in occlusion and depth prior calculations. Iteration convergence is determined by both the residual reduction rate and the step size norm to prevent parameter oscillations.

[0128] This embodiment generates a feature matching matrix M after iteration. Matrix entries record 3D point indices, 2D point indices, matching confidence, residuals, scale levels, and scene categories. To ensure the reliability of the results for subsequent registration and stitching, this embodiment combines the transformation matrix of S402 with M to form a correction parameter table, containing the projection model version, solution weight strategy, and convergence index. Before writing to the correction database, integrity verification is performed: first, geometric consistency—the residual distribution of rigid classes should be significantly better than that of non-rigid classes; if this is not met, it reverts to the previous scale for re-estimation; second, temporal consistency—the matching trajectories of three consecutive frames should be smooth; if "skipped matching" occurs, check for timestamp offsets or uncompensated shutter speeds; third, coverage consistency—read the density coverage D of S202; if matching points are mainly concentrated in narrow areas, mark "extrapolation risk" in the parameter table, reminding S103 to increase the weight of the photometric term in non-covered areas.

[0129] This embodiment provides a solution expression:

[0130] J = Σ_k ρ(||π(P, X_k) − u_k||) + α Σ_k (1 − cosθ_k).

[0131] Where J is the matching cost function, X_k is the spatial coordinate of the k-th 3D contour point, u_k is the pixel coordinate of the corresponding 2D contour point, π(P,·) represents the projection function defined by intrinsic and extrinsic parameters, θ_k is the angle between the projection edge normal and the image edge direction, ρ is the robust loss, and α is the orientation consistency weight. The physical meaning of each parameter directly corresponds to sensor observation: the error occurs in the image plane, and the orientation term constrains the edge geometry, conforming to the natural law of the relationship between imaging and geometry.

[0132] In this embodiment, during the daytime scene of the park roads, the lane lines and curbs form stable long sides, the point cloud density gradient changes clearly, and the 3D-2D matching converges rapidly. The glass window area is marked as highly reflective, so its weight is weakened in the matching process to avoid virtual images dragging extrinsic parameter estimation. At night in the server room corridor, strong white light causes the image histogram to be high, the edges are clear but the illumination is uneven. The module relies on the cabinet facade as a rigid anchor point, and hierarchical matching suppresses mismatches caused by local highlights. The final generated matching matrix is ​​stable in the depth direction, meeting the depth and occlusion inference requirements of subsequent S103 stitching.

[0133] The technical problems solved by this embodiment are concentrated in three aspects:

[0134] First, the cross-modal difference between sparse point clouds and dense images is addressed by using density gradient and orientation consistency constraints to accurately extract "registerable structures" from massive textures. Second, interference from dynamic and reflective regions is prevented from being mistakenly included in the geometric solution through semantically driven weighting and filtering. Third, the integrity and traceability of parameters and matching results are ensured by versioning the calibration parameter table and using evidence fingerprints to support reliable reuse and verification in subsequent modules. Finally, the data written into the calibration database includes the linkage status of matching relationships and projection parameters, providing a solid and interpretable input for multi-level stitching and deep fusion.

[0135] In one embodiment of the real-scene image data processing method based on multi-dimensional data dynamic correction in this application, it may further include the following:

[0136] Step S601: Read the matrix data in the calibration database, construct a continuity test module, calculate the difference vector of adjacent time data, extract the change characteristics of the difference vector, generate time series continuity constraints, construct a continuity record table containing time number, difference value, and constraint parameters, and write the continuity record table into the test database.

[0137] Step S602: Read the continuity record table in the inspection database, construct a coordinate mapping module, map the coordinate system transformation matrix to the point cloud image data, calculate the two-dimensional plane projection coordinates, generate a coordinate mapping table, apply smooth transition constraints to the coordinate mapping table, and write the coordinate mapping table into the mapping database.

[0138] In this embodiment, S601 and S602 are executed in two scenarios: walking inspection of the park and mapping of the computer room corridor. The projection / external participation feature matching matrix from the calibration database is input, and the unified time number generated by S302 and the quality label of S202 / S301 are read in combination to avoid equal weighting of low confidence frames in time series analysis.

[0139] The starting point is to establish an interpretable and continuous characterization of data from adjacent time points. In this embodiment, difference vectors are calculated simultaneously from two lines: "motion state and observation content". One is the pose difference, which calculates the relative pose ΔTn based on the extrinsic parameter matrices Tn and Tn+1 of adjacent time points in the calibration database; the other is the observation difference, which reads the feature matching after the previous correction step and statistically analyzes the reprojection offset of the corresponding point on the image plane and the normal change in three dimensions, forming δu_n and δn_n. If the angular velocity direction of ΔTn is consistent with the principal direction of δu_n and the amplitude can be explained by the imaging scale, the change is labeled as "externally explained by motion"; if it deviates and is concentrated in the dynamic mask region, it is recorded as "externally generated change". The two differences are robustly aggregated within a short window to extract change features such as mean, variance, principal component of direction, and long-tail index, which are used for subsequent constraint adaptive weighting.

[0140] This embodiment generates temporal continuity constraints based on differential features. Considering the motion continuity of the handheld device and the stability of the scene geometry, rigid global smoothing is not directly used. Instead, a combination of pose smoothing and projection smoothing constraints is constructed. Pose smoothing prioritizes the stability of ΔTn, limiting translational and rotational jumps between adjacent time steps. Projection smoothing requires the static anchor point to evolve gradually on the two-dimensional image plane, avoiding misalignment at seams caused by jitter. The weights are driven by differential features. When there is a sudden change in the angle but δu_n is consistent with the IMU, the pose constraint strength is reduced. When the proportion of the dynamic region increases, the robustness kernel of the projection constraint is increased. The continuity verification module writes the constraint parameters into a continuity record table at the same time. The fields include time number n, pose difference index, projection difference index, robust kernel type and weight coefficient, and low confidence flag. This data is stored in the verification database for subsequent mapping stages to read one record at a time, maintaining consistency across modules.

[0141] This embodiment maps the coordinate system transformation matrix to the point cloud and image, generates two-dimensional projected coordinates, and performs smoothing constraints. The projection matrix P_n corresponding to time n in the transformation database is read, and homogeneous projection is performed on the point cloud frame X_n to obtain pixel coordinates u_n. Considering the temporal tail of radar frame scanning, this embodiment performs time interpolation on each scan line within X_n according to the sub-offset δ of S302 to ensure that the projection references the same exposure center. For points whose projections fall out of the field of view or compete at multiple depths within the same pixel, selection is made based on the depth proximity priority and semantic occlusion strategy, establishing an initial coordinate mapping table with fields including point ID, pixel coordinates, depth, semantic label, and quality weight. Subsequently, the temporal continuity constraint of S601 is introduced, and the mapping table is smoothed in two stages: at the global level, a small affine correction is performed on u_n according to the pose smoothing constraint to remove slight jitter; at the local level, curve fitting is performed on the pixel trajectory in the neighborhood of the static anchor point to eliminate abnormal jump points, and only weak smoothing is performed on the dynamic region to avoid the expansion of the tail.

[0142] This embodiment avoids misinterpreting real corners or accelerations as jitter during smooth transitions. The judgment criterion originates from the variation characteristics of S601: when the angular component of ΔTn is significant and the principal direction of δu_n is in the same direction, and the image quality scores C and M are within the confidence interval, the width of the local smoothing window automatically shrinks to preserve the real edge displacement at the corner; conversely, if the image illumination uniformity U is low and noise increases, the projection smoothing uses a stronger robust kernel to prevent low-quality frames from exerting too much influence on the trajectory. For the mirror / water surface label area, only the pixel coordinates used for photometric fusion are retained in the mapping table, and the depth value is not output to avoid subsequent depth field contamination by virtual images.

[0143] This embodiment provides a simplified energy representation:

[0144] E = α·||ΔTn − ΔTn−1|| Σ + β·Σ_k ρ(||u_n(k) − u {n−1}(k)||) + γ·Ω.

[0145] In the formula, ΔTn represents the external covariance difference between adjacent time steps, ||·||_Σ represents the weighted norm under the covariance Σ, characterizing the pose continuity; u_n(k) is the projected coordinate of the k-th static anchor point at time n, ρ is the robust loss, characterizing the projection continuity; Ω is the set of constraint regularization terms, including dynamic mask culling and quality weight modulation; α, β, and γ are weight coefficients, jointly determined by the difference features, semantic proportions, and quality scores. The physical meaning of each quantity clearly corresponds to the device kinematics and imaging projection process, and the parameters can be traced back to the statistics of the preceding modules.

[0146] In this embodiment, in the narrow scenario of a computer room corridor, the angular component ΔTn rises sharply at equipment corners. The displacement of the static anchor point (edge ​​of the cabinet facade) in u_n is consistent with that of the IMU. The system shrinks the smoothing window, and the mapping table maintains sharp edge turns to avoid "blurred corners" during splicing. In the shaded sections of the park roads, wind-blown leaves cause local irregular fluctuations in δu_n, but the classification label is dynamic. The projection smoothing adopts weak constraints on this area, enhancing the continuity only for road markings and curbs. The static trajectory of the mapping table is smooth, and the dynamic texture is not rigidly straightened, maintaining a natural feel. For clock jitter caused by short-term PPS loss, the noise of ΔTn estimation increases, and the continuous recording table automatically lowers α to prevent the time error from being translated into geometric smoothness. The weights are restored after the clock recovers.

[0147] In this embodiment, the smoothed coordinate mapping table is written into the mapping database, recording the mapping version, constraint parameters, participating quality weights, and anomaly markers.

[0148] The reason for embedding the continuity constraint parameters into the library, rather than using them only briefly in memory, is to ensure that subsequent multi-layer stitching and depth fusion require auditing the source and weight of each pixel. This allows for backtracking to the specific constraint window and weight decisions in case of seam misalignment or depth jumps. Overall, S601 solves the problem of distinguishing between self-motion and external interference in adjacent frame changes and translates it into executable constraints. S602 specifically implements the constraints onto the temporal trajectory of each projection point, generating stable coordinate mappings that retain transition details. This provides a reliable geometric path for subsequent generation of real-world images containing depth information, reducing the probability of recalculation and convergence failure.

[0149] In one embodiment of the real-scene image data processing method based on multi-dimensional data dynamic correction in this application, it may further include the following:

[0150] Step S701: Read the correspondence table and smooth transition constraints, construct an image registration module, calculate image transformation parameters based on the correspondence table, apply the transformation parameters to point cloud image data and real scene image data, construct a registration record table containing registration point pairs, transformation matrix, and registration error, and write the registration record table into the registration database;

[0151] Step S702: Read the registration record table in the registration database, construct an image fusion module, use a multi-level image stitching algorithm to fuse the registered image data, generate a real-scene image containing depth information, perform integrity verification on the real-scene image, construct a result record table containing image number, depth data, and fusion parameters, and write the result record table into the result database.

[0152] This embodiment executes S701 and S702 around the output stage of the park's pedestrian inspection and computer room corridor mapping. The inputs are the correspondence table and correction parameter table written in S502, and the smooth transition constraints generated in S103. The starting point is the determination of the solution strategy in the image registration module. This embodiment first filters feature point pairs that meet the confidence standard and groups them according to scene labels: rigid anchor points (walls, cabinet surfaces, curbs, and markings) are used for the main constraints, while weakly stable classes (vegetation and glass highlights) are only used for verification and occlusion assistance; matching pairs covered by dynamic masks are assigned zero or very low weights to avoid dragging the global picture. Then, combined with the smooth transition constraints, a cross-frame pose initial value chain is established to avoid single-frame solutions falling into local extrema. On this basis, the image-point cloud transformation parameters are solved, with variables including extrinsic perturbations and small-amplitude intrinsic drifts. The goal is to minimize the reprojection error and orientation cost on static anchor points, and to suppress abrupt changes using a temporal continuity term.

[0153] This embodiment applies the obtained transformation parameters to point cloud and real-world images to generate registration results and error statistics. For the point cloud side, the projection model in the transformation database is read, and the 3D points are mapped to the image plane according to homogeneous coordinates. Foreground and background cropping is performed using the depth and confidence weights of the mapping table to eliminate overlap misjudgments caused by occlusion. For the image side, the residual transformation is updated from bottom to top according to the scale pyramid. The low-resolution layer corrects global bias, and the high-resolution layer fine-tunes local distortion.

[0154] The registration error is not calculated based on a single pixel distance, but rather uses a composite index of local similarity and geometric consistency: edge pairs are evaluated using reprojection distance and normal angle, while texture regions are evaluated using NCC or gradient correlation. The weights are derived from the quality score and scene strategy. Finally, a registration record table is constructed, with fields including registration point pair index, transformation matrix version, residual statistics (mean, quantiles, long tail ratio), time window number, and confidence label, and written to the registration database for retrieval during the fusion stage.

[0155] The reason for introducing a direction term into the registration target is based on the fact that the observability of edge geometry is more stable. Simple pixel differences are easily misled by changes in illumination and texture repetition. Introducing a temporal continuity term follows the natural law of handheld motion inertia and avoids pose jumps caused by single-frame noise in weak texture segments. For glass highlights or water reflections, although the image edges are obvious, stable supports are often lacking in the point cloud. The strategy is to identify false matches by the inconsistency between the occlusion relationship and depth after registration, and to mark the source when writing back the error to prevent error statistics from masking systematic biases.

[0156] In this embodiment, step S702 is used to construct the image fusion module. The input consists of registered multi-frame images and point cloud projections. The fusion algorithm employs multi-level stitching: at the spatial level, a weight map of the overlapping region guides the seams, and the weight map is derived from exposure balance, sharpness, viewing angle difference, and occlusion priority; at the frequency domain level, a Laplacian pyramid is used to mix at different scales, with low frequencies ensuring smooth brightness transitions and high frequencies preserving edge sharpness.

[0157] Depth information is gathered from two sources: first, the geometric depth of the point cloud is projected onto the depth map; second, the dense depth is estimated from multi-view geometry (binocular or adjacent frame disparity). These two sources are then fused using a confidence-weighted approach. The sources of confidence are traceable: point cloud confidence is based on local planar residuals and echo stability, while disparity confidence is based on reprojection consistency and photometric matching residuals. To reduce motion blur caused by dynamic volumes, the most recent single-frame texture is prioritized in dynamic mask regions, and depth is primarily determined by the disparity path with a short-term update window.

[0158] This embodiment inserts local consistency corrections for color and exposure before fusion. Gain and offset parameters are not independently fitted, but are constrained by S301 exposure equalization and illumination uniformity to avoid overfitting that could cause color drift. For cross-frame blocks with abrupt changes in illumination, local histogram matching is used to limit them within a reasonable range, preventing excessive flattening of shadow boundaries and thus disrupting geometric seam cues. During fusion, seam lines are selected using the minimum energy path. The energy term considers brightness differences, gradient breaks, and depth discontinuities, prioritizing areas that pass through low-texture regions with minimal depth changes to reduce "ghost seams."

[0159] This embodiment outputs a real-world image containing depth information and performs an integrity check. The check includes two dimensions: geometric consistency and photometric consistency. Geometrically, the linearity and parallelism deviation of known anchor points (roadsides, cabinet edges) in the fused image are evaluated. Photometrically, brightness and chromaticity gradients are sampled on both sides of the seam to detect abrupt changes. If anomalies are found, this embodiment traces the cause: if they are concentrated in a certain time window and accompanied by low synchronization confidence, the process backtracks to the S302 time segment for weight reduction; if they are concentrated in high-brightness reflection areas, the process backtracks to the S401 label to correct the reflection mask. The results of the checks are organized into a result record table, recording the image number, depth data summary (confidence histogram and coverage), fusion parameter version, and anomaly markers, and written to the result database.

[0160] This embodiment provides a formula for expressing the overall energy, which is used to unify registration and fusion constraints within the same framework:

[0161] F = αE_geo + βE_photo + γE_depth,

[0162] Where F represents the total energy, E_geo is the geometric term derived from reprojection error and anchor point orientation consistency, E_photo is the photometric term, measuring the brightness / gradient difference in the overlapping area, and E_depth is the depth consistency term, measuring the deviation between point cloud depth and multi-parallax depth in the overlapping area. α, β, and γ are weighting coefficients, adaptively set based on scene labels and quality scores. The three terms correspond to the natural laws of motion geometry, imaging, and 3D geometry, respectively. The parameters are not black-box extrapolations but are all supported by upstream statistics.

[0163] For example, during the day, the road edges and markings provide strong geometric anchors in the park roads, and the registration error converges quickly at the low-frequency layer. The stitching seams are guided along the low-gradient area of ​​the paving texture, and the depth is mainly point cloud with parallax trimming. At night, the computer room corridor suffers from uneven illumination and rolling shutter speeds, causing local deformation. The temporal term weight is increased during registration, and the gain is limited during fusion. The depth is mainly parallax with point cloud providing planar regularization. The final output real-world image maintains stable linearity in the long corridor's depth direction, and the equipment label text is not stretched at the seams. In summary, the implementation of S701 and S702 integrates the preceding correspondence, temporal smoothing, and scene semantics into a traceable transformation-fusion link, solving the problem of cross-modal registration being susceptible to illumination and dynamic disturbances. The output real-world results with depth meet the direct usage requirements of surveying, inspection annotation, and AR navigation.

[0164] To effectively address the shortcomings of traditional technologies in data acquisition, coordinate transformation, and image fusion, and to provide technical support for real-scene image processing, this application provides an embodiment of a real-scene image data processing method based on multidimensional data dynamic correction, which implements all or part of the aforementioned method. See [link to embodiment]. Figure 2 The real-scene image data processing method based on multi-dimensional data dynamic correction specifically includes the following:

[0165] The point cloud image correction module 10 is used to collect real-scene data through a handheld image acquisition device, receive three-dimensional point cloud image data generated by radar scanning, receive two-dimensional real-scene image data generated by a camera, calculate the density distribution and noise level of the point cloud image data, evaluate the clarity and exposure of the real-scene image data, generate a data quality score vector, and assign a unified timestamp identifier to the point cloud image data and the real-scene image data.

[0166] Image feature matching module 20 is used to construct a scene classification module, perform semantic segmentation on image data, identify scene types and dynamic objects, construct a data correction module, calculate the projection transformation parameters between the three-dimensional spatial coordinate system and the two-dimensional plane coordinate system, generate a coordinate system transformation matrix, extract contour feature points from point cloud images and real scene images, calculate the correspondence between the contour feature points based on the iterative nearest point algorithm, and generate a feature matching matrix.

[0167] The real-scene image construction module 30 is used to analyze the changing characteristics of data at adjacent time points, establish data smooth transition constraints, map the coordinate system transformation matrix to point cloud image data, generate a two-dimensional plane coordinate mapping table, establish the correspondence of feature points based on the feature matching matrix, and use a multi-level image stitching algorithm to fuse the registered image data to generate a real-scene image containing depth information.

[0168] As described above, the real-scene image data processing method based on multi-dimensional data dynamic correction provided in this application can achieve effective data management through innovative data acquisition system design, quality assessment, and time stamping. It constructs a coordinate transformation mechanism, combining feature matching and contour analysis to establish a reliable registration model. By introducing smooth transitions and employing multi-level stitching and deep fusion, it ensures processing accuracy. This method effectively addresses the shortcomings of traditional technologies in data acquisition, coordinate transformation, and image fusion, providing technical support for real-scene image processing.

[0169] From a hardware perspective, in order to effectively address the shortcomings of traditional technologies in data acquisition, coordinate transformation, and image fusion, and to provide technical support for real-scene image processing, this application provides an embodiment of an electronic device for implementing all or part of the aforementioned real-scene image data processing method based on multi-dimensional data dynamic correction. The electronic device specifically includes the following components:

[0170] The system comprises a processor, memory, a communications interface, and a bus; wherein the processor, memory, and communications interface communicate with each other via the bus; the communications interface is used to realize information transmission between the real-scene image data processing method based on multi-dimensional data dynamic correction and core business systems, user terminals, and related databases and other related devices; the logic controller can be a desktop computer, tablet computer, or mobile terminal, etc., and this embodiment is not limited to these. In this embodiment, the logic controller can be implemented with reference to the embodiments of the real-scene image data processing method based on multi-dimensional data dynamic correction in the previous embodiments, and the contents of these embodiments are incorporated herein, and repeated details will not be described again.

[0171] It is understood that the user terminal may include smartphones, tablet computers, network set-top boxes, portable computers, desktop computers, personal digital assistants (PDAs), in-vehicle devices, smart wearable devices, etc. Among these, the smart wearable devices may include smart glasses, smartwatches, smart bracelets, etc.

[0172] In practical applications, some parts of the real-scene image data processing method based on multi-dimensional data dynamic correction can be executed on the electronic device side as described above, or all operations can be completed in the client device. The specific choice depends on the processing power of the client device and the limitations of the user's usage scenario. This application does not impose any limitations on this. If all operations are completed in the client device, the client device may further include a processor.

[0173] The aforementioned client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side; in other implementation scenarios, it may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a distributed server structure.

[0174] Figure 3This is a schematic block diagram illustrating the system configuration of the electronic device 9600 according to an embodiment of this application. Figure 3 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that... Figure 3 This is an example; other types of structures can also be used to supplement or replace this structure to achieve telecommunications functions or other functions.

[0175] In one embodiment, the real-scene image data processing method based on multi-dimensional data dynamic correction can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following control:

[0176] Step S101: Collect real-scene data through a handheld image acquisition device, receive three-dimensional point cloud image data generated by radar scanning, receive two-dimensional real-scene image data generated by camera acquisition, calculate the density distribution and noise level of the point cloud image data, evaluate the clarity and exposure of the real-scene image data, generate a data quality score vector, and assign a unified timestamp identifier to the point cloud image data and the real-scene image data.

[0177] Step S102: Construct a scene classification module to perform semantic segmentation on image data, identify scene types and dynamic objects, construct a data correction module to calculate the projection transformation parameters between the three-dimensional spatial coordinate system and the two-dimensional plane coordinate system, generate a coordinate system transformation matrix, extract contour feature points from point cloud images and real scene images, calculate the correspondence between the contour feature points based on the iterative nearest point algorithm, and generate a feature matching matrix.

[0178] Step S103: Analyze the changing characteristics of data at adjacent time points, establish data smooth transition constraints, map the coordinate system transformation matrix to point cloud image data, generate a two-dimensional plane coordinate mapping table, establish the correspondence of feature points based on the feature matching matrix, and use a multi-level image stitching algorithm to fuse the registered image data to generate a real-world image containing depth information.

[0179] As described above, the electronic device provided in this application, through an innovatively designed data acquisition system, achieves effective data management via quality assessment and time stamping. It constructs a coordinate transformation mechanism, combining feature matching and contour analysis to establish a reliable registration model. A smooth transition is introduced, and multi-level stitching and deep fusion ensure processing accuracy. This method effectively addresses the shortcomings of traditional technologies in data acquisition, coordinate transformation, and image fusion, providing technical support for real-scene image processing.

[0180] In another embodiment, the real-scene image data processing method based on multi-dimensional data dynamic correction can be configured separately from the central processing unit 9100. For example, the real-scene image data processing method based on multi-dimensional data dynamic correction can be configured as a chip connected to the central processing unit 9100, and the function of the real-scene image data processing method based on multi-dimensional data dynamic correction can be realized through the control of the central processing unit.

[0181] like Figure 3 As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily need to include these components. Figure 3 All components shown; in addition, the electronic device 9600 may also include Figure 3 For components not shown, please refer to existing technologies.

[0182] like Figure 3 As shown, the central processing unit 9100, sometimes also referred to as a controller or operating control, may include a microprocessor or other processor device and / or logic device, which receives inputs and controls the operation of various components of the electronic device 9600.

[0183] The memory 9140 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It may store the aforementioned failure-related information, and also store a program for executing that information. The central processing unit 9100 may execute the program stored in the memory 9140 to perform information storage or processing, etc.

[0184] Input unit 9120 provides input to central processing unit 9100. Input unit 9120 may be, for example, a keypad or touch input device. Power supply 9170 provides power to electronic device 9600. Display 9160 displays images and text. Display may be, for example, an LCD display, but is not limited thereto.

[0185] The memory 9140 can be a solid-state memory, such as a read-only memory (ROM), random access memory (RAM), a SIM card, etc. It can also be a memory that retains information even when power is off, can be selectively erased, and contains more data; examples of this type of memory are sometimes referred to as EPROMs. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs or processes for executing the operation of the electronic device 9600 via the central processing unit 9100.

[0186] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers for the electronic device for communication functions and / or for performing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0187] The communication module 9110 is a transmitter / receiver that sends and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processing unit 9100 to provide input signals and receive output signals, which is the same as in a conventional mobile communication terminal.

[0188] Based on different communication technologies, multiple communication modules 9110 can be configured in the same electronic device, such as cellular network modules, Bluetooth modules, and / or wireless LAN modules. The communication module 9110 (transmitter / receiver) is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide audio output via the speaker 9131 and receive audio input from the microphone 9132, thereby realizing typical telecommunications functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Additionally, the audio processor 9130 is coupled to a central processing unit 9100, enabling on-device recording via the microphone 9132 and on-device playback of stored audio via the speaker 9131.

[0189] Embodiments of this application also provide a computer-readable storage medium capable of implementing all steps of the real-scene image data processing method based on multi-dimensional data dynamic correction, where the execution subject is a server or client, as described in the above embodiments. The computer-readable storage medium stores a computer program that, when executed by a processor, implements all steps of the real-scene image data processing method based on multi-dimensional data dynamic correction, where the execution subject is a server or client, as described in the above embodiments. For example, when the processor executes the computer program, it implements the following steps:

[0190] Step S101: Collect real-scene data through a handheld image acquisition device, receive three-dimensional point cloud image data generated by radar scanning, receive two-dimensional real-scene image data generated by camera acquisition, calculate the density distribution and noise level of the point cloud image data, evaluate the clarity and exposure of the real-scene image data, generate a data quality score vector, and assign a unified timestamp identifier to the point cloud image data and the real-scene image data.

[0191] Step S102: Construct a scene classification module to perform semantic segmentation on image data, identify scene types and dynamic objects, construct a data correction module to calculate the projection transformation parameters between the three-dimensional spatial coordinate system and the two-dimensional plane coordinate system, generate a coordinate system transformation matrix, extract contour feature points from point cloud images and real scene images, calculate the correspondence between the contour feature points based on the iterative nearest point algorithm, and generate a feature matching matrix.

[0192] Step S103: Analyze the changing characteristics of data at adjacent time points, establish data smooth transition constraints, map the coordinate system transformation matrix to point cloud image data, generate a two-dimensional plane coordinate mapping table, establish the correspondence of feature points based on the feature matching matrix, and use a multi-level image stitching algorithm to fuse the registered image data to generate a real-world image containing depth information.

[0193] As described above, the computer-readable storage medium provided in this application embodiment achieves effective data management through an innovatively designed data acquisition system, quality assessment, and time stamping. It constructs a coordinate transformation mechanism, combining feature matching and contour analysis to establish a reliable registration model. A smooth transition is introduced, and multi-level stitching and deep fusion ensure processing accuracy. This method effectively addresses the shortcomings of traditional technologies in data acquisition, coordinate transformation, and image fusion, providing technical support for real-scene image processing.

[0194] Embodiments of this application also provide a computer program product capable of implementing all steps of the real-scene image data processing method based on multi-dimensional data dynamic correction, in which the execution subject is a server or client as described in the above embodiments. When executed by a processor, this computer program / instruction implements the steps of the real-scene image data processing method based on multi-dimensional data dynamic correction. For example, the computer program / instruction implements the following steps:

[0195] Step S101: Collect real-scene data through a handheld image acquisition device, receive three-dimensional point cloud image data generated by radar scanning, receive two-dimensional real-scene image data generated by camera acquisition, calculate the density distribution and noise level of the point cloud image data, evaluate the clarity and exposure of the real-scene image data, generate a data quality score vector, and assign a unified timestamp identifier to the point cloud image data and the real-scene image data.

[0196] Step S102: Construct a scene classification module to perform semantic segmentation on image data, identify scene types and dynamic objects, construct a data correction module to calculate the projection transformation parameters between the three-dimensional spatial coordinate system and the two-dimensional plane coordinate system, generate a coordinate system transformation matrix, extract contour feature points from point cloud images and real scene images, calculate the correspondence between the contour feature points based on the iterative nearest point algorithm, and generate a feature matching matrix.

[0197] Step S103: Analyze the changing characteristics of data at adjacent time points, establish data smooth transition constraints, map the coordinate system transformation matrix to point cloud image data, generate a two-dimensional plane coordinate mapping table, establish the correspondence of feature points based on the feature matching matrix, and use a multi-level image stitching algorithm to fuse the registered image data to generate a real-world image containing depth information.

[0198] As described above, the computer program product provided in this application, through an innovatively designed data acquisition system, achieves effective data management through quality assessment and time stamping. It constructs a coordinate transformation mechanism, combining feature matching and contour analysis to establish a reliable registration model. By introducing smooth transitions and employing multi-level stitching and deep fusion, it ensures processing accuracy. This method effectively addresses the shortcomings of traditional technologies in data acquisition, coordinate transformation, and image fusion, providing technical support for real-scene image processing.

[0199] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0200] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (devices), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0201] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0202] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0203] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A method for processing real-scene image data based on dynamic correction of multi-dimensional data, characterized in that, The method includes: Real-world data is acquired using a handheld image acquisition device, receiving 3D point cloud image data generated by a radar device and 2D real-world image data acquired by a camera. The density distribution and noise level of the point cloud image data are calculated, the clarity and exposure of the real-world image data are evaluated, a data quality score vector is generated, and a unified timestamp identifier is assigned to the point cloud image data and the real-world image data. A scene classification module is constructed to perform semantic segmentation on image data and identify scene types and dynamic objects. A data correction module is constructed to establish a three-dimensional spatial coordinate system from point cloud image data and a two-dimensional planar coordinate system from real scene image data. Based on the identified scene types and dynamic objects, projection transformation parameters are selected to generate a coordinate system transformation matrix. Contour feature points of point cloud images and real scene images are extracted. The correspondence between the contour feature points is calculated based on the iterative nearest point algorithm to generate a feature matching matrix. Analyze the changing characteristics of data at adjacent time points, establish data smooth transition constraints, map the coordinate system transformation matrix to point cloud image data, generate a two-dimensional plane coordinate mapping table, establish the correspondence between feature points based on the feature matching matrix, and use a multi-level image stitching algorithm to fuse the registered image data to generate a real-world image containing depth information. The multi-level image stitching algorithm for fusing the registered image data includes: according to the scene type identified by the semantic segmentation, using a rigid stitching model for regions identified as rigid structures and a flexible weighted hybrid model for regions identified as non-rigid structures.

2. The method for processing real-scene image data based on dynamic correction of multi-dimensional data according to claim 1, characterized in that, The process of acquiring real-scene data via a handheld image acquisition device, receiving 3D point cloud image data generated by a radar device, receiving 2D real-scene image data acquired by a camera, and calculating the density distribution and noise level of the point cloud image data includes: Read the data stream output by the handheld image acquisition device, construct a data acquisition module, receive point cloud image data generated by radar scanning, receive real scene image data generated by camera acquisition, synchronize the point cloud image data with the real scene image data in time, construct an acquisition record table containing acquisition time, device identifier, and data type, and write the acquisition record table into the acquisition database. The system reads image data from the acquisition database, constructs a quality assessment module, calculates the point density per unit volume of the point cloud image data, calculates the signal-to-noise ratio parameter of the point cloud image data, generates a point cloud quality score vector, and writes the score vector into the assessment database.

3. The method for processing real-scene image data based on dynamic correction of multi-dimensional data according to claim 1, characterized in that, The process of evaluating the sharpness and exposure of the real-scene image data, generating a data quality score vector, and assigning a unified timestamp identifier to the point cloud image data and the real-scene image data includes: Read real-scene image data, construct an image evaluation module, calculate the brightness distribution histogram of the image, extract the gradient feature matrix of the image, calculate the image sharpness index based on the histogram and gradient matrix, count the proportion of bright and dark areas of the image, generate an image quality score vector, and write the score vector into the evaluation database. The scoring vector in the evaluation database is read, a timestamp allocation module is constructed, the clock signal of the hardware device is extracted, a unified time reference sequence is generated, the point cloud image data and the real scene image data are marked with timestamp numbers, an acquisition record table containing data identifiers, time numbers and scoring values ​​is constructed, and the acquisition record table is written into the acquisition database.

4. The method for processing real-scene image data based on dynamic correction of multi-dimensional data according to claim 1, characterized in that, The scene classification module performs semantic segmentation on image data to identify scene types and dynamic objects. The data correction module calculates the projection transformation parameters between the three-dimensional spatial coordinate system and the two-dimensional planar coordinate system, generating a coordinate system transformation matrix. Read image data from the acquisition database, construct a semantic segmentation module, perform regional block processing on the image data, extract texture and color features of the image, generate scene feature vectors, identify scene types and dynamic objects based on the feature vectors, construct a classification record table containing scene labels, object identifiers, and feature attributes, and write the classification record table into the scene database. The classification record table in the scene database is read, a spatial coordinate system transformation module is constructed, a three-dimensional spatial coordinate system is established from point cloud image data, a two-dimensional plane coordinate system is established from real scene image data, projection transformation parameters are selected based on the scene label, a coordinate system transformation matrix is ​​generated, and the transformation matrix is ​​written into the transformation database.

5. The method for processing real-scene image data based on dynamic correction of multi-dimensional data according to claim 1, characterized in that, The step of extracting contour feature points from the point cloud image and the real-world image, calculating the correspondence between the contour feature points based on the iterative nearest point algorithm, and generating a feature matching matrix includes: Read point cloud image data and real scene image data, construct a contour extraction module, calculate the density gradient of the point cloud data, mark the abrupt change position of the density gradient, extract the edge contour point set, perform edge detection processing on the real scene image, extract the image contour feature point set, construct a contour record table containing point set identifier, spatial position, and feature attributes, and write the contour record table into the feature database. Read the contour record table in the feature database, construct a feature matching module, calculate the correspondence between two sets of contour points based on the iterative nearest point algorithm, generate a feature matching matrix, combine the transformation matrix and the matching matrix into a correction parameter table, verify the integrity of the correction parameter table, and write the correction parameter table into the correction database.

6. The method for processing real-scene image data based on dynamic correction of multi-dimensional data according to claim 1, characterized in that, The process involves analyzing the changing characteristics of data at adjacent time points, establishing constraints for smooth data transition, mapping the coordinate system transformation matrix to point cloud image data, and generating a two-dimensional plane coordinate mapping table, including: Read the matrix data in the calibration database, construct a continuity test module, calculate the difference vector of adjacent time data, extract the change characteristics of the difference vector, generate time series continuity constraints, construct a continuity record table containing time number, difference value, and constraint parameters, and write the continuity record table into the test database. Read the continuous record table in the inspection database, construct a coordinate mapping module, map the coordinate system transformation matrix to the point cloud image data, calculate the two-dimensional plane projection coordinates, generate a coordinate mapping table, apply smooth transition constraints to the coordinate mapping table, and write the coordinate mapping table into the mapping database.

7. The method for processing real-scene image data based on dynamic correction of multi-dimensional data according to claim 1, characterized in that, The process of establishing feature point correspondences based on the feature matching matrix, fusing the registered image data using a multi-level image stitching algorithm, and generating a real-world image containing depth information includes: Read the correspondence table and smooth transition constraints to construct an image registration module. Calculate image transformation parameters based on the correspondence table. Apply the transformation parameters to point cloud image data and real-world image data. Construct a registration record table containing registration point pairs, transformation matrices, and registration errors. Write the registration record table into the registration database. The registration record table in the registration database is read, an image fusion module is constructed, and a multi-level image stitching algorithm is used to fuse the registered image data to generate a real-scene image containing depth information. The real-scene image is then subjected to integrity verification, and a result record table containing image number, depth data, and fusion parameters is constructed. The result record table is then written into the result database.

8. A real-scene image data processing device based on multi-dimensional data dynamic correction, characterized in that, The device includes: The point cloud image correction module is used to collect real-scene data through a handheld image acquisition device, receive three-dimensional point cloud image data generated by radar scanning, receive two-dimensional real-scene image data generated by a camera, calculate the density distribution and noise level of the point cloud image data, evaluate the clarity and exposure of the real-scene image data, generate a data quality score vector, and assign a unified timestamp identifier to the point cloud image data and the real-scene image data. The image feature matching module is used to construct a scene classification module, perform semantic segmentation on image data, identify scene types and dynamic objects, construct a data correction module, establish a three-dimensional spatial coordinate system from point cloud image data, establish a two-dimensional planar coordinate system from real scene image data, select projection transformation parameters based on the identified scene type and dynamic objects, generate a coordinate system transformation matrix, extract contour feature points from point cloud images and real scene images, calculate the correspondence of the contour feature points based on the iterative nearest point algorithm, and generate a feature matching matrix. The real-scene image construction module is used to analyze the changing characteristics of data at adjacent time points, establish data smooth transition constraints, map the coordinate system transformation matrix to point cloud image data, generate a two-dimensional plane coordinate mapping table, establish the correspondence of feature points based on the feature matching matrix, and use a multi-level image stitching algorithm to fuse the registered image data to generate a real-scene image containing depth information. The multi-level image stitching algorithm is used to fuse the registered image data, which includes: according to the scene type identified by the semantic segmentation, a rigid stitching model is used for regions identified as rigid structures, and a flexible weight mixing model is used for regions identified as non-rigid structures.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the real-scene image data processing method based on multi-dimensional data dynamic correction as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the real-scene image data processing method based on multi-dimensional data dynamic correction as described in any one of claims 1 to 7.