Visual positioning method and system based on manifold optimization and two-dimensional code
By using a visual positioning method based on manifold optimization and QR codes, efficient real-time positioning that adapts to environmental changes is achieved. This solves the problems of low computational efficiency and time delay deviation in multi-source fusion positioning, and improves the positioning accuracy and operation and maintenance efficiency of mobile device clusters.
Patent Information
- Application Number
- CN202511781620.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-29
- Publication Date
- 2026-03-17
AI Technical Summary
Existing multi-source fusion positioning methods cannot adapt to environmental changes, have low computational efficiency, and suffer from time delay deviations in digital twin synchronization, affecting the high-precision positioning and real-time performance of mobile device clusters.
A visual positioning method based on manifold optimization and QR codes is adopted. By collecting multi-source positioning data for spatiotemporal alignment, anomaly detection and quality assessment are performed. An incremental optimization framework is constructed, adaptive weight fusion is carried out, and a deep learning model is combined to predict the positioning quality trend, perform latency compensation and trajectory prediction, and achieve high-fidelity synchronization and intelligent visualization.
It improves the robustness and accuracy of the positioning system, reduces computational complexity, supports parallel optimization processing of large-scale device clusters, and provides high-fidelity real-time synchronization and intelligent operation and maintenance monitoring capabilities.
Smart Images

Figure CN121677679A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of visual positioning, more particularly, it relates to a visual positioning method and system based on manifold optimization and two-dimensional code. BACKGROUND
[0002] With the rapid development of intelligent manufacturing and logistics automation, mobile robots, especially automatic guided vehicles (AGV), are widely used in port, warehouse, factory and other scenarios. Accurate real-time positioning is the basis for autonomous navigation of mobile robots, directly affecting the efficiency and safety of operation. Visual positioning based on two-dimensional code is widely used in indoor and outdoor scenarios due to its low cost and flexible deployment, but single visual positioning system has problems such as unstable positioning accuracy, being easily affected by light and occlusion in complex environments.
[0003] In the prior art, multi-source sensor fusion positioning methods improve positioning accuracy and robustness by combining visual, inertial navigation unit (IMU), odometry and other sensor data. However, traditional fusion methods mostly use fixed weight strategy, which cannot dynamically adjust the reliability weight of each sensor according to environmental changes, resulting in poor fusion effect in harsh environments. At the same time, batch optimization methods have high computational complexity, making it difficult to meet the real-time requirements of large-scale device clusters. In addition, the time delay of data transmission and processing causes synchronization deviation between the digital twin model and the physical device, affecting the accuracy of operation and monitoring.
[0004] To solve the above problems, a multi-source fusion positioning method is needed that can adapt to environmental changes, efficiently and real-time solve, and compensate for system time delay, to achieve high-precision positioning of mobile device clusters and real-time digital twin synchronization, and improve the efficiency and quality of intelligent operation and supervision. SUMMARY
[0005] The present application provides a visual positioning method and system based on manifold optimization and two-dimensional code, which solves the technical problems of multi-source fusion positioning in related technologies that cannot adapt to environmental changes, low computational efficiency, and time delay deviation in digital twin synchronization.
[0006] The present application provides a visual positioning method based on manifold optimization and two-dimensional code, comprising the following steps:
[0007] S1, collecting multi-source positioning data and performing space-time alignment to obtain a synchronized observation data sequence;
[0008] S2, receiving the observation data sequence, performing anomaly detection and quality evaluation on the observation data sequence to obtain a filtered observation data set and a quality label;
[0009] S3, obtaining the filtered observation data set and the quality label, constructing an incremental optimization framework of manifold space, and obtaining an optimized pose sequence;
[0010] S4, obtaining the optimized pose sequence, adaptively fusing the optimized pose sequence of the manifold based on environment perception to obtain a fused pose;
[0011] S5, obtaining the fused pose, predicting a positioning quality trend by using a deep learning model to obtain a positioning quality change trend;
[0012] S6, obtaining the fused pose, establishing a motion model through the fused pose to perform time delay compensation and trajectory prediction to obtain real-time synchronous pose estimation;
[0013] S7, driving a digital twin model according to the real-time synchronous pose estimation and the positioning quality change trend to realize high-fidelity synchronization and intelligent visualization to obtain a real-time three-dimensional monitoring scene.
[0014] In a preferred embodiment, the collecting multi-source positioning data and performing space-time alignment comprises:
[0015] Collecting original sensor data of a visual positioning module, an inertial navigation unit (IMU) and a wheel odometer;
[0016] Implementing clock synchronization between a device end and a server end by using a network time protocol (NTP), calculating network delay and clock deviation through exchanging timestamp information, and calibrating all data timestamps to a server time reference according to the clock deviation;
[0017] Setting a uniform time grid interval, generating intermediate time data by using an interpolation method for sensor data with a sampling frequency lower than a grid frequency, and selecting data points closest to grid time by using a downsampling method for sensor data with a sampling frequency higher than the grid frequency.
[0018] In a preferred embodiment, the performing anomaly detection and quality assessment on the observation data sequence comprises:
[0019] Setting a confidence threshold, marking as to-be-verified data and reducing the weight of the data in subsequent fusion optimization when the actual recognition confidence is lower than the threshold;
[0020] Calculating reasonable position and angle change thresholds according to the maximum linear speed and maximum angular speed of the device, marking as an abnormal value if the position difference or angle difference exceeds the reasonable upper limit, and pre-constructing a two-dimensional code map of the work area, comparing the map position with the calculated position difference when the visual positioning module recognizes the two-dimensional code, and marking the observation data corresponding to the two-dimensional code as an abnormal value if the position difference exceeds the preset threshold.
[0021] In a preferred embodiment, the incremental optimization framework of the manifold space comprises:
[0022] The pose of the mobile device on the plane is represented as an element of the special Euclidean group SE(2), a Lie algebra element is mapped to a Lie group element through an exponential mapping, and a Lie group element is mapped back to a Lie algebra element through a logarithmic mapping;
[0023] The pose optimization problem is represented by a factor graph structure, the factor nodes include visual observation factors, IMU motion factors, odometry motion factors and prior factors, and the optimization objective is to minimize the weighted sum of squared errors of all factors;
[0024] The incremental smoothing and mapping iSAM algorithm is used to solve the pose optimization problem, and the incremental update is realized by maintaining the QR decomposition of the information matrix, and when new observation data arrives, the QR decomposition is updated by the Givens rotation numerical method.
[0025] In a preferred embodiment, the adaptive weight fusion of the pose sequence after manifold optimization based on environmental perception comprises:
[0026] The image brightness mean and variance, image sharpness, two-dimensional code region edge strength, and similarity feature between consecutive frames are extracted from the image, and the features are input into a pre-trained environmental quality evaluation model to output a reliability score of visual positioning;
[0027] The IMU cumulative running time is recorded, the IMU working temperature is monitored, and the reliability score of the IMU is obtained by multiplying the temperature influence factor, the drift coefficient and the cumulative time, and then taking the negative value and calculating the natural exponential;
[0028] The difference between the odometry speed and the IMU integrated speed is compared to detect the slipping state, and when the speed difference exceeds the preset threshold, it is considered that slipping occurs and the reliability score of the odometry is reduced.
[0029] In a preferred embodiment, the deep learning model is used to predict the positioning quality trend, which comprises:
[0030] The pose uncertainty sequence, sensor reliability score sequence, observation data quality sequence, multi-source data consistency sequence, pose change acceleration sequence and environmental feature sequence are extracted, and the features are organized into a feature matrix in the time dimension;
[0031] The long short-term memory network LSTM is used as the backbone of the quality prediction model, the network structure includes an input layer, an LSTM layer, a fully connected layer and an output layer, and the model training adopts a supervised learning method;
[0032] When new data arrives, the feature matrix is updated, and the updated feature matrix is input into the trained LSTM model to perform forward propagation to obtain the quality prediction of the future time window;
[0033] The prediction uncertainty is estimated by Monte Carlo Dropout, and the Dropout layer activation is kept during prediction to perform multiple forward propagations to obtain multiple prediction results.
[0034] In a preferred embodiment, the time delay compensation and trajectory prediction by fusing the pose to build a motion model comprises:
[0035] The kinematic model of the differential drive AGV is established, the device state vector includes the x-direction position coordinate, the y-direction position coordinate, the rotation angle, the linear velocity and the angular velocity, the continuous time model is discretized and the process noise is added in the state transition equation to represent the uncertainty;
[0036] The extended Kalman filter (EKF) is used to fuse the motion model and sensor observations, the EKF includes a prediction step and an update step, and the optimal state estimation and covariance estimation are obtained by recursive calculation;
[0037] The multi-step prediction function of the EKF is used for time delay compensation, and multi-step prediction is performed continuously from the delay time, and the state transition is performed at each step using the control instruction at the corresponding time. After multi-step prediction, the state prediction value at the current time is obtained.
[0038] In a preferred embodiment, the digital twin model is driven according to the real-time synchronized pose estimation and the positioning quality change trend to realize high-fidelity synchronization and intelligent visualization, which comprises:
[0039] The optimized pose is converted from the local coordinate system of the work area to the world coordinate system of the digital twin scene, and the converted pose is decomposed into a position vector and a rotation quaternion and transmitted to the digital twin model;
[0040] The trajectory prediction is used to plan the model motion path in the rendering engine in advance, and when new observation data arrives, the predicted trajectory is compared with the actual pose. If the deviation is small, the transition is smooth, and if the deviation is large, the correction is fast.
[0041] The display effect of the digital twin model is dynamically adjusted according to the positioning quality prediction and the pose uncertainty, a multi-level quality indication scheme is designed, and an uncertainty ellipse is drawn around the model to visualize the uncertainty range of the pose estimation.
[0042] In a preferred embodiment, the real-time three-dimensional monitoring scene comprises:
[0043] According to the distance between the device and the camera, the device is divided into foreground layer, middle layer and far layer, different precision models are used for different layers, view frustum culling technology is used to render only the devices within the camera field of view, and occlusion culling technology is used to not render the devices completely occluded by other objects.
[0044] The motion trajectory is generated by connecting position points of the historical pose sequence, the quality heat map is drawn on the ground, the working area is divided into grids, the average positioning quality in each grid is counted, and the grid is colored according to the quality score to form the heat map.
[0045] Clicking any device model pops up a detailed information panel to display the real-time state of the device, provides a timeline control function to play back the historical running process, supports multiple device frame selection and batch operation.
[0046] In a preferred embodiment, a manifold optimization and two-dimensional code-based visual positioning system for performing the steps of the above-described manifold optimization and two-dimensional code-based visual positioning method comprises:
[0047] A data acquisition module is configured to acquire multi-source positioning data and perform spatio-temporal alignment to obtain a synchronized observation data sequence.
[0048] An anomaly detection module is configured to receive the observation data sequence, perform anomaly detection and quality evaluation on the observation data sequence, and obtain a filtered observation data set and a quality label.
[0049] A manifold optimization module is configured to obtain the filtered observation data set and the quality label, construct an incremental optimization framework of a manifold space, and obtain an optimized pose sequence.
[0050] An adaptive fusion module is configured to obtain the optimized pose sequence, perform adaptive weight fusion on the optimized pose sequence based on environment perception, and obtain a fused pose.
[0051] A quality prediction module is configured to obtain the fused pose, predict a positioning quality trend using a deep learning model, and obtain a positioning quality change trend.
[0052] A time delay compensation module is configured to obtain the fused pose, establish a motion model based on the fused pose to perform time delay compensation and trajectory prediction, and obtain real-time synchronized pose estimation.
[0053] A digital twin module is configured to drive a digital twin model based on the real-time synchronized pose estimation and the positioning quality change trend to achieve high-fidelity synchronization and intelligent visualization, and obtain a real-time three-dimensional monitoring scene.
[0054] The present application has the following advantages:
[0055] By constructing an incremental optimization framework of manifold space and combining with an adaptive weight fusion mechanism of environment perception, the fusion weight of each sensor can be dynamically adjusted according to the current environment state, reliable data sources can be automatically relied on in the harsh environment such as light change, two-dimensional code pollution and wheel skid, the robustness and precision of the positioning system are improved, meanwhile, the incremental optimization algorithm realizes efficient real-time solving by maintaining the QR decomposition of the information matrix, the calculation complexity is reduced, and the parallel optimization processing of large-scale equipment cluster is supported.
[0056] By establishing a kinematic model for multi-step prediction to realize time delay compensation, combining a deep learning quality prediction model to discover potential faults in advance, adopting a predictive synchronization strategy to drive a digital twin model, high-fidelity real-time synchronization of a virtual scene and physical equipment is realized, an intuitive and comprehensive three-dimensional visual monitoring interface and active early warning function are provided for operation and maintenance personnel, and the efficiency and decision quality of intelligent operation and maintenance supervision are improved. BRIEF DESCRIPTION OF DRAWINGS
[0057] Figure 1 is a flowchart of a visual positioning method based on manifold optimization and two-dimensional codes of the present application;
[0058] Figure 2 is a module diagram of a visual positioning system based on manifold optimization and two-dimensional codes of the present application;
[0059] Figure 3 is a flowchart of a visual positioning method based on manifold optimization and two-dimensional codes of the present application. DETAILED DESCRIPTION
[0060] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that the discussion of these implementations is merely meant to provide a better understanding of the subject matter described herein and can be changed in function and arrangement without departing from the scope of the present description. Various processes or components can be omitted, substituted, or added according to desired implementations. Additionally, features described in some examples can be combined in other examples.
[0061] At least one embodiment of the present application discloses a visual positioning method based on manifold optimization and two-dimensional codes, as shown in Figures 1 to 3 includes the following steps:
[0062] S1, collect multi-source positioning data and perform space-time alignment to obtain a synchronized observation data sequence;
[0063] Collect multi-source positioning data of mobile equipment through an industrial Internet of Things protocol, and perform timestamp calibration and space coordinate conversion to output a space-time aligned observation data sequence; specifically including the following steps:
[0064] S11, establish a real-time communication connection with the mobile device and collect raw sensor data;
[0065] The operation and maintenance supervision platform establishes a real-time communication connection with the vehicle-mounted controller of the mobile device through the MQTT protocol and subscribes to the sensor data topics published by each device. Each mobile device is equipped with a visual positioning module, an inertial navigation unit IMU, and a wheel odometer. The visual positioning module identifies the ground two-dimensional code identifier through the vehicle-mounted camera, outputs the two-dimensional code ID, image coordinates, recognition confidence, position coordinates, and rotation angle, and the sampling frequency is determined according to the camera frame rate, which is usually 10 Hz. The IMU outputs three-axis acceleration and three-axis angular velocity, and the sampling frequency is determined according to the sensor model, which is usually 100 Hz. The wheel odometer outputs the left and right wheel speeds and the cumulative mileage, and the sampling frequency is determined according to the encoder resolution, which is usually 50 Hz.
[0066] After the platform receives the data, it performs parsing and format conversion, unifies it into a standard data structure, and extracts the key parameters of each sensor for subsequent processing.
[0067] S12, network time synchronization and timestamp calibration;
[0068] Due to different clock sources of each sensor and network transmission delay, timestamp calibration is needed to ensure data time consistency. The platform uses the Network Time Protocol NTP to synchronize the clocks of the device end and the server end, calculates the network delay and clock bias by exchanging timestamp information. Local timestamps are added to each data packet at the data collection end, and server timestamps are recorded at the cloud end. According to the clock bias, all data timestamps are calibrated to the server time reference. To reduce the impact of network delay fluctuations, Kalman filtering is used to smooth the estimation of the clock bias.
[0069] S13, align data with different sampling frequencies to a unified time grid;
[0070] After timestamp calibration, different sensor data still has different sampling frequencies, which needs to be aligned to a unified time grid for fusion processing. Set the interval of the unified time grid, the interval value is determined according to the real-time requirements of the system and the computing resources, usually 10 milliseconds. For sensor data with a sampling frequency lower than the grid frequency, use interpolation method to generate intermediate time data; for sensor data with a sampling frequency higher than the grid frequency, use downsampling method to select the data point closest to the grid time.
[0071] Linear interpolation is used for position coordinates, and spherical linear interpolation is used for rotation angle considering the periodicity. When the angle crosses the positive and negative 180-degree boundary, normalization processing is performed. In the interpolation process, the data validity is considered, if there is no output data from the sensor in a certain period of time, the interpolated data at that time is marked as invalid and does not participate in subsequent fusion optimization.
[0072] S14, coordinate system conversion and spatial registration are performed;
[0073] Different sensor data are defined in different coordinate systems, and coordinate system conversion is needed to ensure spatial consistency. A conversion relationship between the global coordinate system and the body coordinate system is established. The global coordinate system takes the fixed point of the work area as the origin, and the body coordinate system takes the geometric center of the device as the origin. The conversion between the two coordinate systems is determined by the device pose, which includes rotation transformation and translation transformation.
[0074] For the acceleration and angular velocity in the body coordinate system measured by the IMU, rotation transformation is performed through the rotation matrix to obtain the corresponding values in the global coordinate system. Sensor installation position calibration is performed. The offset of each sensor relative to the geometric center of the device is obtained through offline calibration, and compensation is performed during data processing. The device center position is obtained by subtracting the offset vector after rotation from the sensor measurement position.
[0075] After coordinate system conversion and spatial registration, all sensor data are unified into the global coordinate system and correspond to the device geometric center position, providing a consistent spatial reference for subsequent fusion optimization.
[0076] S15, output the spatio-temporal aligned multi-source observation data sequence;
[0077] After the above processing, the spatio-temporal aligned multi-source observation data sequence is obtained. For each time grid point, a data structure containing timestamp, visual positioning data, IMU data, odometry data and control instructions is output. This data sequence serves as the input for the subsequent steps, supporting multi-source data fusion optimization and quality evaluation.
[0078] S2, receive the observation data sequence, perform anomaly detection and quality evaluation on the observation data sequence, and obtain the filtered observation data set and quality label;
[0079] To address the noise and outlier problems of visual positioning data in outdoor environments, a multi-level anomaly detection method is designed, and the quality of each observation data is evaluated, outputting a reliable observation data set and quality label after filtering; specific steps include:
[0080] S21, preliminary filtering based on recognition confidence;
[0081] The visual positioning module outputs a confidence score when recognizing the two-dimensional code, ranging from 0 to 1, indicating the reliability of the recognition result. The confidence score takes into account factors such as image quality, two-dimensional code edge definition, and decoding success rate. A confidence threshold is set, which is determined according to the recognition success rate statistics in the actual application scenario, usually taking 0.6 to 0.8. When the actual recognition confidence is lower than the threshold, it is considered that the frame positioning data is unreliable, and is marked as verification data. For verification data, it is not directly rejected, but its weight in subsequent fusion optimization is reduced, avoiding complete dependence on a single data source.
[0082] S22, jump detection based on motion continuity constraint;
[0083] The motion of the mobile device has continuity, and the position and velocity at adjacent time should not have sudden changes. For the visual positioning result at time t, the difference between it and the previous time pose is calculated. The position difference is obtained by calculating the Euclidean distance of the coordinate difference between the current time and the previous time. The angle difference is obtained by calculating the absolute value of the difference between the rotation angles of the current time and the previous time, considering the periodicity of the angle, when the angle difference is greater than π, the normalization processing is performed.
[0084] According to the maximum linear velocity and the maximum angular velocity of the device, reasonable position and angle change thresholds are calculated, and the maximum linear velocity and the maximum angular velocity are obtained from the technical parameters of the device. In a given time interval, the reasonable upper limit of the position change is the maximum linear velocity multiplied by the time interval, and the reasonable upper limit of the angle change is the maximum angular velocity multiplied by the time interval. If the position difference or the angle difference exceeds the reasonable upper limit, it is considered that the positioning data has a jump, and is marked as an abnormal value.
[0085] For the detected jump data, further analysis is performed on the reason. If multiple consecutive frames of data have jumps, it may be that the device has indeed moved quickly or collided, in which case the data is retained but the weight is reduced. If only individual frames have jumps and the surrounding frames are normal, it is considered to be a positioning error, and the interpolation result of the adjacent normal frames is used to replace it.
[0086] S23, rationality test based on map constraint;
[0087] A two-dimensional code map of the work area is constructed in advance, recording the ID and accurate position of each two-dimensional code, and the map data is stored in the database with the two-dimensional code ID as the index. When the visual positioning module recognizes a certain two-dimensional code, the standard global position of the two-dimensional code is queried from the map database according to the two-dimensional code ID, and at the same time the visual positioning module calculates the relative position of the two-dimensional code in the device coordinate system through image processing, and combines the current pose estimation to convert the relative position into a calculation position in the global coordinate system through coordinate transformation.
[0088] The map position and the calculated position difference are obtained by calculating the Euclidean distance of the coordinate difference of the two positions. If the position difference exceeds the preset threshold, it indicates that the positioning calculation is incorrect, and the data is marked as an abnormal value. The preset threshold is determined according to the accuracy of the two-dimensional code laying and the calibration accuracy of the visual positioning system, and is usually set to 0.2 to 0.5 meters. The map constraint test can effectively identify the positioning deviation caused by environmental interference.
[0089] S24, outlier elimination based on multi-frame consistency;
[0090] Single-frame anomaly detection may have false positives, and multi-frame data consistency analysis can improve detection reliability. A sliding window method is used to count the proportion of abnormal data within the time window, and the length of the time window is determined according to the device motion speed and the sampling frequency, usually 1 to 2 seconds. If the proportion of abnormal data in the window exceeds the threshold, which is usually 50% according to system fault tolerance requirements, the visual positioning system is considered to have failed as a whole during that period, and the visual positioning data in the entire window is marked as unusable, switching to the dead reckoning mode based on IMU and odometer. If only a small number of data in the window is abnormal, it is considered to be an occasional positioning error, and the weighted average or polynomial fitting of the normal data in the window is used for interpolation replacement. Multi-frame consistency test can distinguish between systematic failure and occasional error.
[0091] S25, calculating the quality weight of each observation data;
[0092] Based on the above abnormality detection results, the quality weight of each visual positioning observation data is calculated, which reflects the data reliability and is used for weighted processing in subsequent fusion optimization. The quality weight calculation considers factors such as recognition confidence, motion continuity, map consistency, etc. The total quality weight is obtained by multiplying the confidence-based weight, the motion continuity-based weight, and the map constraint-based weight. Figure 1 The confidence-based weight is obtained by dividing the actual recognition confidence by the confidence threshold. The motion continuity-based weight is set to 1 if no jump is detected, and if a jump is detected, the decay weight is calculated by taking the negative value of the position difference divided by the position change standard deviation and calculating the natural exponential. The position change standard deviation is obtained from historical data statistics. The map consistency-based weight is set to 1 if the map error is within the threshold, and if it exceeds the threshold, the decay weight is calculated by taking the negative value of the map position difference divided by the map error standard deviation and calculating the natural exponential. The map error standard deviation is determined according to the map calibration accuracy.
[0093] Figure 1
[0094] The final total quality weight is normalized to the range of 0 to 1, and the higher the weight, the more reliable the data. Data with a weight below the minimum threshold is directly eliminated and does not participate in subsequent processing, and the minimum threshold is determined according to the system's tolerance for low-quality data, usually 0.1.
[0095] S26, output reliable observation data set and quality label;
[0096] After anomaly detection and quality evaluation, the filtered visual positioning observation data set is output. For each time, the information containing visual positioning data, quality weight, anomaly flag and correction flag is output. The data set is used as the observation input of the subsequent manifold optimization, and the quality weight is used to adjust the weight coefficient in the optimization objective function, thereby improving the robustness of the fusion result.
[0097] In some embodiments, a machine learning-based anomaly detection method can be used instead of a rule-based method to improve the accuracy and generalization ability of anomaly detection. Specifically, historical operation data is collected, normal data and abnormal data are labeled, and a classification model such as a support vector machine or a random forest is trained. The input features of the model include multi-dimensional features such as recognition confidence, position change, angle change, map error, image quality indicators, and the output is an anomaly probability. Anomaly determination is performed by setting a probability threshold, and the probability threshold is determined according to the balance requirement of the false detection rate and the missed detection rate. The machine learning method can automatically learn the complex features of the anomaly pattern and adapt to different environmental conditions.
[0098] S3, obtain the filtered observation data set and quality label, construct an incremental optimization framework of the manifold space, and obtain the optimized pose sequence;
[0099] The pose of the mobile device is represented in the manifold space, and an incremental optimization framework based on the factor graph is constructed to realize efficient real-time pose solving, and the optimized pose sequence and covariance matrix are output; specifically including the following steps:
[0100] S31, establish the manifold representation of the pose on the special Euclidean group SE(2);
[0101] The pose of the mobile device on the plane is represented as an element of the special Euclidean group SE(2). SE(2) is composed of the plane rotation group SO(2) and the translation vector, and has group structure and differential manifold properties. The pose transformation matrix is a three-row three-column homogeneous transformation matrix, which is composed of a rotation matrix and a translation vector. The rotation matrix is a two-row two-column matrix, whose elements are composed of trigonometric values of the rotation angle, and the orthogonality constraint is satisfied to ensure that the rotation transformation does not change the length of the vector.
[0102] SE(2) constitutes a three-dimensional manifold, and its tangent space is Lie algebra se(2). The Lie algebra element is a three-dimensional vector, which contains the x-direction translation speed, the y-direction translation speed and the rotation angular velocity. The exponential mapping maps the Lie algebra element to the Lie group element, realizing the conversion from the velocity space to the pose space. The logarithmic mapping maps the Lie group element back to the Lie algebra element, realizing the inverse conversion from the pose space to the velocity space.
[0103] By manifold representation, the pose update operation is defined as Lie group multiplication. Pose update is achieved by converting pose increment via exponential mapping to Lie group form, and then performing Lie group multiplication with current pose to get updated pose. This updating method automatically maintains the rotation matrix orthogonality constraint, avoiding the parameterization singularity problem in Euclidean space.
[0104] S32, constructing a factor graph based optimization problem representation;
[0105] The factor graph structure is used to represent the pose optimization problem. The factor graph is a bipartite graph composed of variable nodes and factor nodes. The variable nodes represent the pose sequence to be optimized, and the factor nodes represent various observation constraints and motion constraints. The factor nodes include visual observation factors, IMU motion factors, odometry motion factors and prior factors. Each factor corresponds to an error function and a covariance matrix, and the covariance matrix is determined according to the sensor calibration parameters. The optimization goal is to minimize the weighted sum of squared errors of all factors.
[0106] The visual observation factor error is obtained by predicting the position of the two-dimensional code in the device coordinate system according to the current pose estimate and the known position of the two-dimensional code in the map, and then subtracting the predicted position from the actual observation value to obtain the observation error vector. The IMU motion factor error is obtained by calculating the relative transformation between the poses at adjacent two time points, calculating the expected rotation change according to the angular velocity measured by the IMU and the time interval, comparing the actual relative transformation with the expected rotation change and converting it to an error vector in the Lie algebra space through logarithmic mapping. The odometry motion factor error is obtained by calculating the linear and angular velocities of the device according to the left and right wheel speeds, the linear velocity being equal to the average of the left and right wheel speeds, and the angular velocity being equal to the difference between the left and right wheel speeds divided by the wheelbase, the wheelbase being obtained from the device mechanical parameters, and then calculating the relative pose change according to the linear and angular velocities and the time interval, comparing the actual relative pose transformation with the relative pose change calculated by the odometry and obtaining the error vector through logarithmic mapping.
[0107] S33, efficient solving is performed by using the incremental smoothing and mapping iSAM algorithm;
[0108] Traditional batch optimization methods need to solve the entire optimization problem at each time point, which has high computational complexity and is difficult to meet real-time requirements. The incremental smoothing and mapping iSAM algorithm reduces the computational complexity by maintaining the QR decomposition of the information matrix for incremental updating. The core idea of the iSAM algorithm is that when the factor graph structure changes, only the affected part is updated locally, rather than re-decomposing the entire matrix.
[0109] The optimization problem is represented as a least square problem, which is converted into a linear equation set by constructing an information matrix and normal equations. The information matrix is decomposed into the product of an orthogonal matrix and an upper triangular matrix through QR decomposition. The normal equations are simplified to an upper triangular equation set by using the characteristics of the orthogonal matrix, and the calculation complexity of solving the upper triangular equation set is much lower than directly solving the normal equations.
[0110] When new observation data arrives, new variable nodes and factor nodes are added in the factor graph, and QR decomposition is efficiently updated by numerical methods such as Givens rotation. When old data moves out of the sliding window, the corresponding variable node is marginalized by the Schur complement method, and the information of the marginalized variable is transmitted to the adjacent variables to maintain the consistency of the optimization problem. Through incremental updating and marginalization operations, the iSAM algorithm improves the calculation efficiency while maintaining the optimization accuracy, meeting the real-time requirements.
[0111] S34, parallel optimization of multiple devices is realized;
[0112] In the multi-device scenario, the pose optimization of each device is independent and can be processed in parallel. The platform uses multi-threading technology to allocate independent optimizer instances and calculation threads for each device. Each optimizer maintains its own factor graph structure and QR decomposition. When new data arrives for a device, the corresponding optimizer thread is awakened to perform incremental updating and solving operations, and the results are written to shared memory for the visualization module to read. To avoid thread competition and synchronization overhead, lock-free data structures and message queues are used for thread communication. Each device data stream is processed independently and there is no data dependency, enabling efficient parallel computing. In the case of sufficient hardware resources, parallel optimization can linearly improve system processing capacity and support larger-scale device clusters.
[0113] S35, output the optimized pose sequence and uncertainty estimate;
[0114] After incremental optimization and solving, the optimal pose estimation sequence at all times within the sliding window is obtained. Each pose is represented as an SE(2) group element, including the x component and y component of the position and the rotation angle. The covariance matrix of the pose is also extracted from the optimization process, which reflects the uncertainty of the pose estimate. The covariance matrix is obtained by inverting the information matrix, which is obtained by multiplying the transpose of the upper triangular matrix with the upper triangular matrix itself. The covariance matrix is a symmetric positive definite matrix with three rows and three columns. The diagonal elements represent the x-direction position variance, y-direction position variance, and rotation angle variance, respectively. The non-diagonal elements represent the correlation between variables. The covariance matrix is used for subsequent quality assessment and uncertainty visualization.
[0115] The output pose sequence and covariance matrix serve as inputs for subsequent steps, supporting positioning quality prediction, time delay compensation, and digital twin synchronization.
[0116] In some embodiments, an approximate method can be used to improve the calculation efficiency, reduce the calculation time while ensuring the accuracy. Specifically, when the angle change is small, the angle change threshold is usually 0.1 radian according to the accuracy requirement, and the first order approximation of Taylor expansion can be used. This approximate method avoids the calculation of trigonometric functions, and still uses accurate calculation to ensure accuracy when the angle change is large. By adaptively selecting the calculation method, the accuracy and efficiency are balanced in different scenarios.
[0117] S4, obtaining the optimized pose sequence, adaptively fusing the optimized pose sequence based on the environment perception to obtain a fused pose;
[0118] Taking the pose sequence and the covariance matrix output in step S3 as input, the reliability of each sensor is evaluated by perceiving the current environment state, the fusion weight is dynamically adjusted, the manifold optimization is re-executed, and the fused pose estimation result is output; Specifically, the following steps are included:
[0119] S41, extracting image features to evaluate the environmental quality of visual positioning;
[0120] The performance of the visual positioning system is greatly affected by factors such as environmental light, degree of barcode contamination, and motion blur. By analyzing the images captured by the camera, the degree of influence of the current environment on visual positioning is evaluated. Image brightness mean and variance, image sharpness, barcode region edge strength, similarity between consecutive frames, and other features are extracted from the image.
[0121] The image brightness mean and variance reflect the lighting conditions and are calculated by traversing all pixels of the image. The threshold values of the brightness mean that are too high or too low are determined according to the dynamic range of the camera, and the threshold value of the brightness mean that is too high is usually 200 and the threshold value of the brightness mean that is too low is usually 50, indicating that the lighting conditions are poor. The threshold value of the brightness variance that is too small is determined according to the scene contrast requirement, and is usually 100, indicating that the image contrast is low. The image sharpness is calculated by the Laplace operator, and the sharpness index is defined as the variance of the Laplace response. The sharpness threshold is determined according to the camera resolution and focusing accuracy. The barcode region edge strength is calculated by the Sobel operator, and the edge strength is defined as the average of the gradient amplitude. The similarity between consecutive frames is calculated by the normalized cross-correlation method, and the normalized cross-correlation value ranges from -1 to 1, and the closer the value is to 1, the more similar the two frames are.
[0122] The above features are input into a pre-trained environment quality evaluation model, which uses a support vector machine or a lightweight neural network and is trained offline through historical data. The model outputs a reliability score for visual positioning, ranging from 0 to 1, and the higher the score, the more suitable the current environment for visual positioning.
[0123] S42, evaluating the drift degree of the IMU based on the accumulated time and temperature;
[0124] The main sources of IMU measurement error are bias drift and random noise. Bias drift accumulates over time, causing the integrated pose estimate to gradually deviate from the true value. Temperature changes also affect IMU performance. Record the cumulative running time of the IMU, starting from the last successful visual positioning correction. IMU drift error is obtained by adding the initial error to the product of the drift coefficient and the cumulative time, and the drift coefficient is obtained by sensor calibration, reflecting the rate of IMU drift over time.
[0125] Monitor the working temperature of the IMU at the same time. When the temperature deviates significantly from the calibration temperature, the drift coefficient increases. The temperature influence factor is obtained by calculating the absolute value of the difference between the current temperature and the calibration temperature multiplied by the temperature sensitivity coefficient and then adding 1, and the temperature sensitivity coefficient is obtained by sensor temperature characteristic test. The IMU reliability score is obtained by multiplying the temperature influence factor, the drift coefficient and the cumulative time, taking the negative value and calculating the natural exponential, the score range is 0 to 1, the longer the cumulative time or the greater the temperature deviation, the lower the score.
[0126] When visual positioning is detected to be successful, reset the cumulative time to zero, indicating that the IMU drift is corrected by visual positioning. This periodic correction mechanism can effectively suppress long-term IMU drift and maintain the accuracy of fusion positioning.
[0127] S43, detect the slipping state of the odometer based on speed consistency;
[0128] The odometer calculates the device motion speed and displacement by measuring the wheel speed. When the wheels slip, the speed is inconsistent with the actual motion, causing the odometer measurement error to increase. The slipping state is detected by comparing the difference between the odometer speed and the IMU integrated speed. The odometer speed is obtained by adding the left wheel speed and the right wheel speed and dividing by 2 to get the average linear speed of the device. The IMU integrated speed is obtained by obtaining the speed estimate at the previous time, extracting the component of the IMU measured acceleration in the direction of device motion, and adding the acceleration component to the speed at the previous time to get the current time speed estimate.
[0129] The speed difference is obtained by subtracting the odometer speed from the IMU integrated speed and taking the absolute value. When the speed difference exceeds the preset threshold, it is considered that slipping occurs, and the slipping detection threshold is set according to the road surface state and wheel characteristics, usually 0.2 to 0.5 meters per second. Slip detection can also be combined with road surface state information, if the road surface types of different road segments are marked in the work area map, the road surface type can be queried according to the current position of the device to dynamically adjust the slipping detection threshold.
[0130] The odometer reliability score is obtained by dividing the speed difference by the speed difference standard deviation, taking the negative value and calculating the natural exponential, the speed difference standard deviation is obtained according to historical data statistics, the score range is 0 to 1, the larger the speed difference, the lower the score. When slipping is detected, the reliability score is reduced, and the odometer weight is reduced in fusion.
[0131] S44, dynamically adjusting the fusion weight in the optimization objective function;
[0132] Based on the reliability score of each sensor, the weight coefficient in the optimization objective function is dynamically adjusted. The optimization objective function contains multiple error terms, each corresponding to a weight coefficient. The visual observation factor weight coefficient is obtained by dividing the visual positioning reliability score by the square of the measurement standard deviation, which is determined according to the sensor calibration parameters. The IMU motion factor weight coefficient is obtained by dividing the IMU reliability score by the square of the measurement standard deviation. The odometer motion factor weight coefficient is obtained by dividing the odometer reliability score by the square of the measurement standard deviation.
[0133] Dynamic adjustment of the weight coefficient enables the fusion algorithm to automatically adapt to environmental changes and automatically rely on more reliable data sources when different sensor performance fluctuates, improving the robustness of the fusion result. To avoid sudden changes in weight causing pose jumps, exponential moving average is used to smooth the weight. Weight smoothing is achieved by multiplying the smoothing coefficient by the smoothed weight at the previous time, multiplying the result of 1 minus the smoothing coefficient by the original weight at the current time, and adding the two products to obtain the smoothed weight at the current time. The smoothing coefficient is determined according to the system dynamic response requirement and is usually taken as 0.8 to 0.9. The smoothing mechanism makes the weight change more gradual, avoiding discontinuity in the optimization result.
[0134] S45, re-executing incremental optimization to obtain adaptively fused pose;
[0135] The updated weight coefficient is applied to the optimization objective function, and incremental optimization is re-executed to solve it. Since an incremental framework is used, only the affected factor nodes need to be updated, and the computational load is small. The optimization objective function is obtained by adding the weighted error sum of squares of all visual observation factors, IMU motion factors, and odometer motion factors. By adjusting the weight coefficient, the contribution of different error terms to the total objective function changes, and the optimization result automatically shifts towards more reliable data sources. Solving the optimization problem obtains a new pose estimation sequence. Compared with the fusion result with fixed weights, the adaptively fused result is more stable when the environment changes and has higher positioning accuracy.
[0136] S46, outputting the robust pose estimation after adaptive fusion;
[0137] The pose sequence and covariance matrix sequence after adaptive weight fusion are output. At the same time, the reliability scores of each sensor are output, including the visual positioning reliability score, the IMU reliability score, and the odometer reliability score, which are used for subsequent quality assessment and fault diagnosis. The adaptive fusion mechanism improves the robustness of the positioning system in complex environments, enabling the fusion positioning to maintain stable performance under various environmental conditions.
[0138] In some embodiments, an online learning method can be employed to automatically adjust the parameters of the scoring model, improving the adaptive ability and generalization performance of the model. Specifically, sensor data and actual positioning errors during system operation are collected to construct online training samples. Recursive least squares or online gradient descent methods are used to update the parameters of the scoring model in real time. The online learning method enables the model to gradually adapt to the current application environment, without the need for manual recalibration, improving system usability and robustness.
[0139] S5, obtaining a fused pose, predicting a positioning quality trend using a deep learning model to obtain a positioning quality change trend;
[0140] The fused pose sequence, covariance matrix, and sensor reliability score output in step S4 are input to construct a positioning quality prediction model based on a long short-term memory network (LSTM), which analyzes the time sequence pattern of historical data to predict future positioning quality change trends, realizes active early warning, and outputs a quality prediction sequence and a warning signal; specifically including the following steps:
[0141] S51, constructing a multi-dimensional time sequence feature vector;
[0142] Positioning quality is affected by multiple factors, including pose uncertainty, sensor reliability, observation data quality, multi-source data consistency, etc. To accurately predict the quality change trend, a multi-dimensional time sequence feature containing these factors needs to be constructed. For data within a past time window, the time window length is determined according to the prediction length and system dynamic characteristics, usually 30 to 60 seconds.
[0143] The extracted features include a pose uncertainty sequence, a sensor reliability score sequence, an observation data quality sequence, a multi-source data consistency sequence, a pose change acceleration sequence, and an environmental feature sequence. The pose uncertainty sequence is obtained by adding the diagonal elements of the covariance matrix to obtain a pose uncertainty index. The observation data quality sequence includes the number of valid observations and the average confidence of visual positioning. The multi-source data consistency sequence is obtained by estimating the device pose based on visual positioning, IMU, and odometry single data sources respectively, and calculating the norm of the difference between visual pose and IMU pose and the difference between visual pose and odometry pose to obtain a consistency index. The pose change acceleration sequence is obtained by calculating the first-order difference of the pose change between adjacent time points to obtain the pose change speed, and then calculating the first-order difference of the pose change speed divided by the time interval to obtain the pose change acceleration. The environmental feature sequence includes image brightness mean, image variance, image sharpness, edge intensity, etc.
[0144] The above features are organized into a feature matrix in the time dimension, with the number of rows corresponding to the number of sampling points of the time window length, and the number of columns corresponding to the feature dimension.
[0145] S52, design LSTM network structure and offline training;
[0146] The long short-term memory network (LSTM) is used as the backbone of the quality prediction model. LSTM can effectively capture long-term dependencies in time series through a gating mechanism, making it suitable for handling dynamic change patterns of positioning quality. The network structure includes an input layer, an LSTM layer, a fully connected layer, and an output layer. The input layer receives a multi-dimensional time series feature matrix. The LSTM layer contains stacked LSTM units, with the number of layers and the hidden dimension of each layer determined by the model complexity and the amount of training data. Typically, there are 2 to 3 layers, with a hidden dimension of 64 to 128. The fully connected layer maps the hidden state at the last time step of LSTM to the prediction target space. A 2-layer fully connected network is used, with the first layer dimension determined by feature extraction requirements, usually 64, and the activation function ReLU. The second layer dimension is equal to the number of future prediction points. The output layer outputs the quality score prediction values for multiple future time points. The prediction length is determined by the early warning lead time requirement, usually 3 to 6 seconds.
[0147] The model training uses a supervised learning approach. Historical operation data is collected, including multi-dimensional feature sequences and corresponding quality score labels. Quality score labeling is based on post-analysis, considering factors such as positioning error, fault occurrence, and maintenance personnel evaluation. The loss function uses mean squared error. The Adam optimizer is used for training, with a learning rate of 0.001 determined by training convergence, a batch size of 32 determined by memory capacity, and a training round number of 50 to 100 rounds determined by validation set performance. The training data set is divided into training set, validation set and test set, with a ratio of 7:2:1. During training, the validation set loss is monitored, and when the validation loss no longer decreases, the training is stopped in advance to avoid overfitting.
[0148] S53, implement online sliding window prediction;
[0149] In practical applications, the model performs online prediction in a sliding window manner. When new data arrives, the feature matrix is updated, the oldest time data is removed from the window, and the latest time data is added, keeping the window length unchanged. The updated feature matrix is input into the trained LSTM model, and forward propagation is performed to obtain the quality prediction for the future time window. To improve prediction real-time performance, model quantization and pruning techniques are used to reduce model size. Quantization compresses model parameters from 32-bit floating-point numbers to 8-bit integers, and pruning removes connections with small weights.
[0150] S54, estimate prediction uncertainty using Monte Carlo Dropout;
[0151] The prediction of the deep learning model has uncertainty, especially in the scenario where the training data is insufficient. To quantify the prediction confidence, the Monte Carlo Dropout technique is used to estimate the prediction uncertainty. Monte Carlo Dropout keeps the Dropout layer active during prediction and performs multiple forward propagations, each time randomly dropping different neurons, to obtain multiple prediction results. For the same input, N forward propagations are performed, N is determined according to the computing resources and accuracy requirements, usually 10 to 50, to obtain N prediction sequences. The mean of these predictions is calculated as the final prediction value, and the variance is calculated as the prediction uncertainty. The prediction uncertainty reflects the confidence of the model for the current input. When the uncertainty is low, it means that the model has high confidence in the prediction result. When the uncertainty is high, it means that the current scenario may be beyond the range of the training data distribution, and the prediction result is less reliable.
[0152] S55, generating a hierarchical warning signal based on the prediction result;
[0153] The hierarchical warning signal is generated according to the quality score and uncertainty of the prediction. Multiple quality thresholds are set to divide the quality score into different levels. The thresholds are determined according to the system reliability requirements and historical fault statistics, and are usually divided into four levels: high quality, medium quality, low quality, and extremely low quality. Traverse the prediction sequence to find the time when the quality score first falls below the threshold. If the quality score of the prediction will be lower than the low quality threshold at a future time and the prediction uncertainty is low, indicating high prediction confidence, a warning signal is triggered.
[0154] The warning signal includes warning level, warning time, warning lead time, and warning cause. The warning level is determined according to the lowest quality score of the prediction, the warning time is the time when the quality score first falls below the threshold, the warning lead time is the time interval from the current time to the warning time, and the warning cause is determined by analyzing the main factors that cause the quality to decrease. The warning signal is sent to the operation and maintenance interface through the message queue to remind the operation and maintenance personnel. The operation and maintenance personnel can take preventive measures in advance to avoid failure according to the warning information.
[0155] S56, outputting the quality prediction sequence and the warning signal;
[0156] The positioning quality prediction sequence, prediction uncertainty sequence, and generated warning signal in the future time window are output. These information are important basis for operation and maintenance decision, supporting proactive preventive maintenance. Compared with traditional passive alarm method, the quality prediction based on deep learning can discover potential problems in advance, increasing the warning lead time.
[0157] In some embodiments, a transfer learning method can be employed to pre-train with other similar scenario data, accelerate model convergence and improve prediction performance in small sample scenarios. Specifically, AGV operation data from other ports or logistics parks is collected to train a general quality prediction model. When deployed in a new scenario, the pre-trained model is used as initialization, and only a small amount of local data is needed for fine-tuning to adapt to the characteristics of the new scenario. The LSTM layer parameters are frozen during fine-tuning, and only the fully connected layer parameters are trained to reduce the number of parameters to be learned. The transfer learning method can reduce the number of training samples required while maintaining similar prediction accuracy, reducing system deployment cost and time.
[0158] S6, obtain a fused pose, establish a motion model through the fused pose to perform time delay compensation and trajectory prediction, and obtain real-time synchronized pose estimation;
[0159] The fused pose sequence and covariance matrix output in step S4 are taken as inputs to establish a kinematic model of the mobile device, an extended Kalman filter is used for state estimation, time delay of data transmission and processing is compensated, current pose and future trajectory of the device are predicted, and real-time pose estimation after time delay compensation is output; specifically including the following steps:
[0160] S61, a kinematic model of a differential drive AGV is established;
[0161] The mobile device adopts a differential drive method to realize forward movement and turning by controlling the speed of left and right wheels. A kinematic model is established to describe the motion law of the device. The device state vector contains five components: x-direction position coordinate, y-direction position coordinate, rotation angle, linear velocity and angular velocity. The state transition equation describes the evolution of the state over time, the x-direction position change rate is equal to the linear velocity multiplied by the current angle cosine value, the y-direction position change rate is equal to the linear velocity multiplied by the current angle sine value, the angle change rate is equal to the angular velocity, the linear velocity change rate is equal to the linear acceleration, and the angular velocity change rate is equal to the angular acceleration.
[0162] The continuous time model is discretized, and the Euler method or Runge-Kutta method is used for numerical integration. The discretization method is determined according to the accuracy requirement and calculation efficiency. Considering that there are model errors and external disturbances in actual motion, process noise is added in the state transition equation to represent uncertainty. The process noise covariance matrix is determined according to the motion characteristics of the device and environmental conditions.
[0163] S62, an extended Kalman filter is designed for state estimation;
[0164] An extended Kalman filter (EKF) is used to fuse the motion model and sensor observations to achieve optimal state estimation. The EKF is an extension of the Kalman filter for nonlinear systems, which recursively estimates the state by linearizing the nonlinear function. The EKF includes a prediction step and an update step. In the prediction step, the state and covariance at the next time step are predicted based on the motion model. The state prediction is calculated by substituting the optimal state estimate at the current time step and the control input into the state transition function. The covariance prediction is calculated by computing the Jacobian matrix of the state transition function, multiplying the current covariance matrix by the Jacobian matrix on the left and the transpose of the Jacobian matrix on the right, and adding the process noise covariance matrix.
[0165] In the update step, the predicted state is corrected by fusing the new sensor observations. The observation equation maps the state vector to the observation space and adds the observation noise. The observation noise covariance matrix is determined based on the sensor calibration parameters. The Kalman gain is calculated by computing the Jacobian matrix of the observation function, computing the product of the predicted covariance matrix and the transpose of the observation Jacobian matrix, computing the triple product of the observation Jacobian matrix, the predicted covariance matrix, and the transpose of the observation Jacobian matrix, adding the observation noise covariance matrix to the sum matrix, and inverting the sum matrix. The state update is calculated by computing the expected observation value based on the predicted state, subtracting the expected observation value from the actual observation value to obtain the observation residual, multiplying the observation residual by the Kalman gain on the left, and adding the product to the predicted state to obtain the updated optimal state estimate. The covariance update is calculated by computing the product of the Kalman gain and the observation Jacobian matrix, subtracting the product from the identity matrix, and multiplying the resulting matrix by the predicted covariance matrix on the right.
[0166] By recursively computing the EKF, the motion model and multi-source observations are fused to obtain the optimal state estimate and covariance estimate.
[0167] S63, a multi-step prediction compensation system delay is realized;
[0168] Due to the time delay in data transmission and processing, the observation data received at the current time actually corresponds to the state of the device at a past time. The total delay is determined based on the network conditions and the computing load, and is usually 50 to 200 milliseconds. To realize real-time synchronization between the digital twin model and the physical device, the real pose of the device at the current time needs to be predicted. The multi-step prediction function of the EKF is used for time delay compensation. First, the observation data at the delayed time is used to update the state estimate to obtain the optimal state estimate and covariance matrix at the delayed time. Then, the number of steps to be predicted is calculated based on the total delay and the time step. From the delayed time, multi-step prediction is performed continuously, and the state transition is performed at each step using the control instruction at the corresponding time. After multi-step prediction, the state prediction value at the current time is obtained.
[0169] The covariance also needs to be propagated, by calculating the cumulative Jacobian matrix of multi-step transition, left multiplying the covariance matrix at the delay time by the cumulative Jacobian matrix, right multiplying the transpose of the cumulative Jacobian matrix, and adding the cumulative process noise covariance matrix to obtain the predicted covariance matrix at the current time. The historical observation data is extrapolated to the current time through multi-step prediction, the system time delay is compensated, and the pose estimation is kept synchronized with the physical device.
[0170] S64, adaptive process noise is designed to improve prediction accuracy;
[0171] The process noise covariance matrix reflects the uncertainty of the motion model. The model error is different in different motion states, and the process noise needs to be adaptively adjusted to improve the prediction accuracy. When moving at a constant speed in a straight line, the motion model prediction accuracy is higher and the process noise is smaller. When accelerating, decelerating or turning, the motion state changes dramatically and the model error increases, so the process noise should be increased accordingly. The process noise matrix is dynamically adjusted according to the acceleration and angular acceleration of the device.
[0172] Linear acceleration is obtained by subtracting the linear velocity at the previous time from the linear velocity at the current time and then dividing by the time interval. Angular acceleration is obtained by subtracting the angular velocity at the previous time from the angular velocity at the current time and then dividing by the time interval. The motion intensity index is obtained by adding the absolute value of the linear acceleration and the absolute value of the angular acceleration, which reflects the degree of change in the motion state. The adaptive process noise is obtained by multiplying the motion intensity index by the adjustment coefficient, then adding 1 to the product, and finally multiplying the obtained value by the reference process noise covariance matrix. The adjustment coefficient is determined according to the dynamic response characteristics of the system, usually taking 1 to 5.
[0173] When the motion intensity is large, the process noise increases, the EKF increases the trust degree of the observation data and reduces the trust degree of the model prediction, thereby reducing the influence of the model error. When the motion intensity is small, the process noise is close to the reference value, maintaining normal filtering characteristics. The adaptive process noise mechanism enables the EKF to automatically adjust the filtering parameters according to the motion state, and maintains high estimation accuracy in various motion modes.
[0174] S65, generate a short-time trajectory prediction in the future;
[0175] In addition to compensating for the current time delay, the trajectory in the future short time needs to be predicted for smooth rendering and operation decision support of the digital twin model. The future trajectory is generated by performing multi-step prediction based on the current state estimation and future control instruction sequence. The prediction length is usually 1 to 3 seconds, determined according to the rendering smoothness requirement and the decision time window. Each prediction point is calculated recursively by the state transition function, and the covariance of each prediction point is calculated to reflect the prediction uncertainty, which gradually increases as the prediction length increases. The future control instruction can be obtained from the task planning module of the vehicle-mounted controller, and if the future control instruction cannot be obtained, a constant speed model is used for prediction. The trajectory prediction result is used for predictive synchronization of the digital twin model, and the model motion path is planned in advance in the rendering engine to achieve smooth animation effect.
[0176] S66, outputting the real-time pose after time delay compensation and the predicted trajectory;
[0177] The current time pose estimation after time delay compensation is output, including position, angle, speed and angular velocity. At the same time, the covariance matrix of the pose is output to reflect the estimation uncertainty. The future short-time trajectory prediction sequence and the corresponding covariance sequence are output, which provides forward-looking information for smooth synchronization and operation decision of the digital twin. The time delay compensation mechanism improves the synchronization accuracy and real-time performance.
[0178] In some embodiments, unscented Kalman filter (UKF) can be used instead of EKF to improve the estimation accuracy of nonlinear systems. UKF uses unscented transformation to select a set of deterministic sampling points called Sigma points, which are propagated through nonlinear functions and then the mean and covariance are calculated, avoiding the calculation of Jacobian matrix and linearization error. Specifically, Sigma points are generated based on the current state estimation and covariance, and the number of Sigma points is 2n+1, where n is the state dimension. These Sigma points are propagated through the state transition function and the observation function to obtain the predicted Sigma points, and the weighted mean and covariance of the predicted Sigma points are calculated to obtain the state prediction and covariance prediction. The calculation complexity of UKF is slightly higher than that of EKF, but the estimation accuracy is higher, especially in high nonlinear scenarios such as rapid turning of equipment.
[0179] S7, driving the digital twin model with the real-time pose and quality prediction sequence based on time delay compensation to achieve high-fidelity synchronization and intelligent visualization, and obtaining a real-time three-dimensional monitoring scene;
[0180] Taking the real-time pose and predicted trajectory output in step S6 and the quality prediction and warning signal output in step S5 as input, the pose data is mapped to the digital twin model, and the predictive synchronization strategy and intelligent rendering technology are used to realize real-time update and interactive visualization of the virtual scene, output an intuitive three-dimensional monitoring interface, and support operation decision; specifically including the following steps:
[0181] S71, coordinate system conversion and pose mapping are performed;
[0182] The optimized pose is defined in the local coordinate system of the work area, and needs to be converted to the world coordinate system of the digital twin scene. The conversion relationship between the two coordinate systems is calibrated in advance, and the conversion matrix is obtained through offline calibration. The calibration method is to select several control points in the work area and measure their coordinates in the two coordinate systems, and then solve the conversion matrix by least squares method. For the pose of the device in the local coordinate system, the pose in the scene coordinate system is calculated through the conversion matrix. The converted pose is decomposed into a position vector and a rotation quaternion, and the three-dimensional rendering engine usually uses quaternions to represent rotation to avoid the gimbal lock problem of Euler angles. The position vector and the rotation quaternion are passed to the digital twin model to update the transformation matrix of the model and drive the position and attitude of the model in the scene to be updated synchronously.
[0183] S72, a predictive synchronization strategy is adopted to realize smooth transition;
[0184] The trajectory prediction output by step S6 is used to plan the model motion path in advance in the rendering engine. When new observation data arrives, the predicted trajectory is compared with the actual pose. If the deviation is small, smooth transition is adopted, and if the deviation is large, rapid correction is adopted. The pose deviation is obtained by obtaining the predicted pose and the actual observed pose, calculating the product of the inverse transformation of the actual observed pose and the predicted pose to obtain the relative transformation between the two, and then converting the relative transformation into a deviation vector in the Lie algebra space through logarithmic mapping. The deviation is obtained by calculating the Euclidean norm of the deviation vector.
[0185] If the deviation norm is less than the preset threshold, it is considered that the prediction is accurate, and interpolation correction is adopted. The preset threshold is determined according to the rendering smoothness requirement and is usually 0.1 to 0.2 meters. The interpolation correction is achieved by gradually increasing the interpolation parameter from 0 to 1. For each interpolation parameter value, the deviation vector is multiplied by the parameter, and the scaled deviation vector is converted into Lie group form through exponential mapping. The interpolated and corrected pose is obtained by right multiplying the Lie group element with the predicted pose. The interpolation process is completed within several rendering frames, so that the model motion is continuous and smooth, avoiding sudden changes and jitter. If the deviation norm is greater than or equal to the threshold, it is considered that the prediction deviation is large, and the model pose is directly jumped to the actual observed pose, and an abnormal event is triggered for analysis by the operation and maintenance personnel. The predictive synchronization strategy combines the foresight of prediction and the accuracy of observation, and realizes smooth visual effect while ensuring synchronization accuracy.
[0186] S73, dynamically adjusting the model display effect according to the positioning quality;
[0187] According to the positioning quality prediction output in step S5 and the pose uncertainty output in step S6, the display effect of the digital twin model is dynamically adjusted to intuitively show the equipment health state. A multi-level quality indication scheme is designed, and the quality score threshold is determined according to the system reliability requirement. When the quality score is high, the model is displayed in normal color, indicating that the positioning system is working normally; when the quality score is medium, the model color changes to warning color and a semi-transparent effect is added to prompt the operation and maintenance personnel to pay attention to the equipment; when the quality score is low, the model color changes to alarm color and the transparency is reduced, and a flashing animation effect is added, and a warning icon is displayed above the model to strongly prompt that there is an abnormality.
[0188] The uncertainty ellipse is drawn around the model to visualize the uncertainty range of the pose estimation, and the shape and size of the ellipse are determined by the covariance matrix. The uncertainty ellipse drawing is performed by extracting the position part of the covariance matrix to perform eigenvalue decomposition to obtain eigenvalues and eigenvectors, the long and short axis directions of the ellipse are determined by the eigenvectors, and the ellipse axis length is obtained by taking the square root of the eigenvalue and multiplying it by a confidence coefficient. The confidence coefficient is determined according to the visualization confidence requirement, and is usually taken as 3 corresponding to a confidence of 99.7%. The ellipse is drawn in a semi-transparent manner with a color consistent with the model color to intuitively show the uncertainty area of the pose. The quality indication and uncertainty visualization enable the operation and maintenance personnel to quickly identify the problem equipment, evaluate the reliability of the positioning result, and make accurate judgments and decisions.
[0189] S74, hierarchical rendering optimization of large-scale equipment cluster is implemented;
[0190] In a multi-device scenario, rendering dozens of high-precision three-dimensional models simultaneously requires high computing resources. Hierarchical rendering strategy and performance optimization techniques are used to ensure rendering frame rate and visual effect. According to the distance between the equipment and the camera, the equipment is divided into foreground layer, middle layer and far layer, different layers use different precision models, and the model precision and the number of patches are determined according to the rendering performance requirement. According to the camera view angle and the equipment position, the rendering level of each equipment is dynamically adjusted, and the level switching uses gradual transition to avoid visual jump caused by sudden model switching.
[0191] The view frustum culling technique is used to render only the equipment within the camera field of view, and the spatial index structure of the scene is constructed to quickly query the equipment within the view frustum. The occlusion culling technique is used to not render the equipment that is completely occluded by other objects, and the visibility of the equipment is determined through depth buffer and occlusion query. Through hierarchical rendering and culling optimization, the rendering frame rate is improved under the premise of ensuring visual quality to support smooth real-time monitoring and interactive operation.
[0192] S75, drawing equipment motion trajectory and quality heat map;
[0193] To show the historical running conditions and positioning quality distribution of the devices, motion trajectories and quality heat maps are drawn in the three-dimensional scene. The motion trajectories are generated by connecting the position points of the historical pose sequence, and the trajectory data is stored in the GPU buffer by using the GPU-accelerated particle system or line rendering technology. The trajectory line is generated by using a vertex shader and a geometry shader. The color of the trajectory line is encoded according to the time or the quality score, and the width of the trajectory line can be adjusted according to the speed or the uncertainty.
[0194] The quality heat map is drawn on the ground, the working area is divided into grids, the average positioning quality in each grid is counted, and the grid size is determined according to the size of the working area and the visualization accuracy requirement, and is usually 1 meter by 1 meter. The grid is colored according to the quality score to form a heat map, and the heat map can intuitively show the performance distribution of the positioning system in the working area and identify the difficult positioning area to provide a basis for environmental optimization. The drawing of the trajectory and the heat map adopts an asynchronous rendering technology to generate graphic data in a background thread to avoid blocking the main rendering process and ensure the smoothness of the interface.
[0195] S76, providing interactive analysis and decision support functions;
[0196] To support the in-depth analysis and decision-making of the operation and maintenance personnel, an interactive function is provided. Clicking any device model pops up a detailed information panel to display the real-time state of the device. The information panel is displayed in the form of a floating window above the scene, and supports dragging and zooming. A time axis control function is provided, and the operation and maintenance personnel can drag the time axis slider to play back the historical running process. The time axis is displayed at the bottom of the interface, and key time points and events are marked. During the playback process, the device model is updated according to the historical trajectory motion quality indication and the trajectory color, which helps the operation and maintenance personnel analyze the spatio-temporal background and reasons of the fault occurrence.
[0197] Multi-device framing and batch operations are supported. The operation and maintenance personnel can frame multiple devices in the scene, batch query the state, compare and analyze, and export data. The batch operation results are displayed in the form of a table or a chart, and sorting, filtering and exporting are supported. A multi-perspective switching function is provided, including a top view, a side view and a follow-up perspective. The perspective switching adopts a smooth animation transition to improve the user experience. The interactive function makes the operation and maintenance platform not only a passive monitoring tool, but also an active analysis and decision support system, which improves the work efficiency and decision-making quality of the operation and maintenance personnel.
[0198] S77, outputting a real-time three-dimensional digital twin monitoring scene;
[0199] The output complete digital twin three-dimensional monitoring scene displays the accurate position motion state, positioning quality and predicted trajectory of all devices in real time. The scene is constructed based on the WebGL technology, supports browser-side operation and does not need to install a special client, thereby reducing the deployment cost. The scene includes interactive interface elements such as a three-dimensional terrain model of a work area, three-dimensional digital models of all mobile devices, motion trajectory lines and quality heat maps of the devices, uncertainty ellipses and quality indication icons, an information panel, a time axis control button and the like. The scene is rendered in real time in response to the interactive operation of the user, and the rendering frame rate is determined according to the hardware performance, usually 60 FPS, to ensure a smooth visual experience.
[0200] The digital twin monitoring scene provides an intuitive, comprehensive and real-time device monitoring view for the operation and maintenance personnel, supports rapid discovery of abnormalities, analysis of problems, decision making, and improves the efficiency and quality of operation and maintenance supervision.
[0201] In some embodiments, cloud rendering technology can be used to transfer the rendering task to a cloud server, supporting a wider range of client devices. Specifically, a high-performance GPU and a rendering engine are deployed on the cloud server to perform rendering calculation of the three-dimensional scene to generate a rendering image, the rendering image is transmitted to the client browser through video stream encoding and compression, and the client only needs to decode the video stream and display the image without performing complex three-dimensional rendering calculation. The user's interactive operation is transmitted to the cloud in real time, and the cloud updates the scene and re-renders according to the interaction. The cloud rendering technology transfers the computing burden to the cloud, and the client only needs a low hardware configuration to obtain a high-quality three-dimensional visualization experience. The cloud rendering scheme is particularly suitable for mobile devices and low-end PC access scenarios, expanding the application scope of the system.
[0202] The visual positioning method and system based on manifold optimization and two-dimensional code proposed in the application realize high-precision real-time positioning and intelligent operation and maintenance supervision of a mobile device cluster through multi-source data fusion, incremental optimization, environment adaptation, deep learning prediction, time delay compensation and digital twin synchronization, and have achieved certain application effects.
[0203] Compared with traditional single visual positioning methods, the multi-source fusion positioning accuracy of the application is improved, and the position error and angle error are reduced. In a harsh environment such as strong light pollution, the success rate of adaptive weight fusion positioning is improved, and the robustness is enhanced.
[0204] In terms of computing efficiency, incremental manifold optimization reduces the pose solving time of a single device, and the computing efficiency is greatly improved. In the scene of parallel optimization of multiple devices, the system can maintain a high real-time update frequency to meet the real-time requirements of digital twins.
[0205] In terms of fault early warning, the quality prediction model based on deep learning can predict the performance decline of the positioning system in advance, with high early warning accuracy and low false positive rate. Compared with the traditional passive warning method, the early warning time is increased, providing sufficient time for operation and maintenance intervention to effectively avoid job interruption caused by equipment failure.
[0206] In terms of synchronization accuracy, the time delay compensation and predictive synchronization strategy based on the motion model reduce the pose deviation and time delay between the digital twin model and the physical device, achieving high-fidelity real-time synchronization. Operation and maintenance personnel can accurately grasp the real-time state of the device through the digital twin scene, improving the accuracy and reliability of monitoring.
[0207] In terms of operation efficiency, the intuitive three-dimensional visualization interface and intelligent early warning function shorten the fault positioning time, greatly improving the operation response efficiency. The historical data playback and statistical analysis function provides data support for optimizing device scheduling strategy and improving operation process, improving the overall operation efficiency.
[0208] The present application provides a systematic technical solution for intelligent operation and maintenance supervision of mobile device clusters, with important practical value and promotion prospects.
[0209] As shown in Table 1, an example of multi-source positioning data collected by an AGV in the first time period is shown:
[0210] Table 1, multi-source positioning data
[0211]
[0212] As shown in Table 2, the pose estimation and quality evaluation results of the AGV after optimization and fusion in the second time period are shown:
[0213] Table 2, pose estimation and quality evaluation results
[0214]
[0215] By comparing the data of the two time periods, it can be seen that the multi-source fusion positioning method of the present application can effectively handle the noise and outliers of visual positioning data, outputting smoother and more accurate pose estimation. The quality evaluation model can monitor the performance change of the positioning system in real time and issue an early warning when the quality decreases, providing a reliable decision basis for operation and maintenance personnel.
[0216] The application focuses on the application field of AGV cluster intelligent operation and maintenance supervision in large port container transfer area. The application effect of the application is demonstrated by taking the actual deployment of a large port as an example. The container transfer area of the port is deployed with multiple AGVs to perform unmanned transfer tasks, each AGV is equipped with a visual positioning module, an IMU, a wheel odometer and other sensors, and a visual positioning system based on a two-dimensional code is used for navigation. The operation area is an outdoor open space, and the ground is paved with two-dimensional code markers.
[0217] The operation and maintenance supervision platform is deployed in the data center of the port and establishes real-time communication with the AGV through a wireless network. The platform adopts a distributed micro-service architecture, and the data acquisition module, optimization calculation module and visual rendering module are respectively deployed on different servers to support horizontal expansion.
[0218] By comparing the data before and after optimization, it can be seen that the multi-source fusion positioning method of the application can effectively process the noise and abnormal values of the visual positioning data, and output more smooth and accurate pose estimation. The quality evaluation model can monitor the performance change of the positioning system in real time, and timely issue a warning when the quality decreases, providing reliable decision basis for operation and maintenance personnel.
[0219] In actual operation, the platform successfully accesses the real-time data of multiple AGVs, realizes real-time synchronization and three-dimensional visualization of key information such as device pose, speed, battery capacity, task state, etc. The operation and maintenance personnel can intuitively master the real-time state and historical trajectory of the whole field device through the monitoring large screen, and quickly identify abnormal devices and problem areas.
[0220] During the trial operation, the health evaluation model of the platform has a high prediction accuracy for common faults such as battery capacity anomaly, positioning accuracy decline and communication delay, and the fault warning time is advanced, effectively avoiding multiple work interruption events caused by device failure. The fault positioning time is shortened, the operation and maintenance response efficiency is greatly improved, the overall operation efficiency is improved, and the economic benefit is improved.
[0221] The above describes the embodiments of the application, but the embodiments are not limited to the specific implementation described above, which is only illustrative and not limiting, and those skilled in the art can make more forms of equivalent embodiments under the inspiration of the embodiments, which are all within the protection scope of the embodiments.
Claims
1. A manifold optimization and QR code based visual positioning method, characterized in that, The method comprises the following steps: S1, collecting multi-source positioning data and performing space-time alignment to obtain a synchronized observation data sequence; S2, receiving the observation data sequence, performing abnormality detection and quality evaluation on the observation data sequence to obtain a screened observation data set and a quality label; S3, obtaining the screened observation data set and the quality label, constructing an incremental optimization framework of a manifold space, and obtaining an optimized pose sequence; S4, obtaining the optimized pose sequence, performing adaptive weight fusion on the optimized pose sequence based on environment perception to obtain a fused pose; S5, obtaining the fused pose, predicting a positioning quality trend by using a deep learning model to obtain a positioning quality change trend; S6, obtaining the fused pose, establishing a motion model through the fused pose to perform time delay compensation and trajectory prediction to obtain real-time synchronized pose estimation; S7, driving a digital twin model based on the real-time synchronized pose estimation and the positioning quality change trend to realize high-fidelity synchronization and intelligent visualization, and obtaining a real-time three-dimensional monitoring scene. 2.The method of claim 1, wherein, The collecting multi-source positioning data and performing space-time alignment comprises: collecting original sensor data of a visual positioning module, an inertial navigation unit (IMU) and a wheel odometer; implementing clock synchronization between a device end and a server end by using a network time protocol (NTP), calculating network delay and clock deviation by exchanging timestamp information, and calibrating all data timestamps to a server time reference according to the clock deviation; setting a uniform time grid interval, generating intermediate time data by using an interpolation method for sensor data with a sampling frequency lower than a grid frequency, and selecting data points closest to grid time by using a downsampling method for sensor data with a sampling frequency higher than the grid frequency. 3.The method of claim 1, wherein, The abnormality detection and quality evaluation on the observation data sequence comprises: setting a confidence threshold, marking data to be verified and reducing the weight of the data in subsequent fusion optimization when the actual recognition confidence is lower than the threshold; calculating reasonable position and angle change thresholds according to the maximum linear velocity and the maximum angular velocity of the device, marking an abnormal value when the position difference or the angle difference exceeds the reasonable upper limit, and pre-constructing a two-dimensional code map of a work area, comparing the map position with the calculated position difference when the visual positioning module recognizes the two-dimensional code, and marking the observation data corresponding to the two-dimensional code as an abnormal value when the position difference exceeds a preset threshold. 4.The method of claim 1, wherein, The incremental optimization framework of the manifold space comprises: representing the pose of the mobile device on the plane as an element of a special Euclidean group SE(2), mapping the Lie algebra element to the Lie group element through an exponential mapping, and mapping the Lie group element back to the Lie algebra element through a logarithmic mapping; representing the pose optimization problem by using a factor graph structure, the factor nodes including a visual observation factor, an IMU motion factor, a wheel odometer motion factor and a prior factor, and the optimization target being to minimize the weighted error sum of squares of all factors; solving the pose optimization problem by using an incremental smoothing and mapping (iSAM) algorithm, updating the QR decomposition of the information matrix to realize incremental updating when new observation data arrives, and updating the QR decomposition by using a Givens rotation numerical method.
5. The manifold optimization and QR code based visual positioning method according to claim 1, wherein, The adaptive weight fusion on the optimized pose sequence based on environment perception comprises: The image brightness mean and variance, image definition, two-dimensional code region edge strength, and similarity between consecutive frames are extracted from the image, and the features are input into a pre-trained environment quality evaluation model to output a reliability score of visual positioning; An IMU cumulative running time is recorded, an IMU working temperature is monitored, and an IMU reliability score is calculated by multiplying a temperature influence factor, a drift coefficient and the cumulative time, taking a negative value and taking a natural exponential; A skid state is detected by comparing the difference between the odometer speed and the IMU integral speed, and when the speed difference exceeds a preset threshold, it is considered that the skid occurs and the odometer reliability score is reduced. 6.The method of claim 1, wherein, The prediction of the positioning quality trend by using the deep learning model comprises: A sequence of pose uncertainty, a sequence of sensor reliability scores, a sequence of observation data quality, a sequence of multi-source data consistency, a sequence of pose change acceleration and a sequence of environmental features are extracted, and the features are organized into a feature matrix in the time dimension; A long short-term memory network (LSTM) is used as the backbone of the quality prediction model, and the network structure comprises an input layer, an LSTM layer, a fully connected layer and an output layer, and the model training adopts a supervised learning method; When new data arrives, the feature matrix is updated, and the updated feature matrix is input into the trained LSTM model to perform forward propagation to obtain a quality prediction of a future time window; A Monte Carlo Dropout is used to estimate the prediction uncertainty, and the Dropout layer activation is maintained during prediction to perform multiple forward propagations to obtain multiple prediction results.
7. The manifold optimization and QR code based visual positioning method according to claim 1, wherein, The time delay compensation and trajectory prediction by fusing the pose to establish a motion model comprise: A kinematic model of a differential drive AGV is established, a device state vector comprises an x-direction position coordinate, a y-direction position coordinate, a rotation angle, a linear speed and an angular speed, a continuous time model is discretized, and process noise is added to the state transition equation to represent uncertainty; An extended Kalman filter (EKF) is used to fuse the motion model and sensor observations, the EKF comprises a prediction step and an update step, and the optimal state estimation and covariance estimation are obtained by recursive calculation; The multi-step prediction function of the EKF is used for time delay compensation, and multi-step prediction is continuously performed from the delay time, each step uses the control command at the corresponding time to perform state transition, and the state prediction value at the current time is obtained after multi-step prediction. 8.The method of claim 1, wherein, The driving of the digital twin model according to the real-time synchronized pose estimation and the positioning quality change trend to realize high-fidelity synchronization and intelligent visualization comprises: The optimized pose is converted from the local coordinate system of the work area to the world coordinate system of the digital twin scene, and the converted pose is decomposed into a position vector and a rotation quaternion and transmitted to the digital twin model; The trajectory prediction is used to plan the model motion path in the rendering engine in advance, and when new observation data arrives, the predicted trajectory is compared with the actual pose, if the deviation is small, the transition is smooth, and if the deviation is large, the deviation is quickly corrected; The display effect of the digital twin model is dynamically adjusted according to the positioning quality prediction and the pose uncertainty, a multi-level quality indication scheme is designed, and an uncertainty ellipse is drawn around the model to visualize the uncertainty range of the pose estimation. 9.The method of claim 1, wherein, The real-time three-dimensional monitoring scene comprises: Devices are divided into foreground, mid-ground and background layers according to the distance between the device and the camera, different layers use different precision models, view frustum culling technology is used to render only the devices within the camera's field of view, and occlusion culling technology is used to not render the devices that are completely occluded by other objects; The motion trajectory is generated by connecting the position points of the historical pose sequence, and the quality heat map is drawn on the ground to divide the work area into grids to count the average positioning quality in each grid, and then color the grid according to the quality score to form a heat map; Clicking on any device model pops up a detailed information panel showing the real-time status of the device, provides time axis control function for historical running process playback, supports multi-device frame selection and batch operation.
10. A manifold optimization and QR code based visual positioning system for performing the steps of a manifold optimization and QR code based visual positioning method according to any one of claims 1 to 9, characterized by Comprise: a data acquisition module for acquiring multi-source positioning data and performing spatio-temporal alignment to obtain a sequence of synchronized observation data; an anomaly detection module for receiving the sequence of observation data, performing anomaly detection and quality evaluation on the sequence of observation data, and obtaining a filtered observation data set and a quality label; a manifold optimization module for obtaining the filtered observation data set and the quality label, constructing an incremental optimization framework of a manifold space, and obtaining an optimized pose sequence; an adaptive fusion module for obtaining the optimized pose sequence, adaptively fusing the optimized pose sequence based on environmental perception, and obtaining a fused pose; a quality prediction module for obtaining the fused pose, predicting a positioning quality trend using a deep learning model, and obtaining a positioning quality change trend; a time delay compensation module for obtaining the fused pose, establishing a motion model through the fused pose to perform time delay compensation and trajectory prediction, and obtaining real-time synchronized pose estimation; a digital twin module for driving a digital twin model based on the real-time synchronized pose estimation and the positioning quality change trend to achieve high-fidelity synchronization and intelligent visualization, and obtaining a real-time three-dimensional monitoring scene.