Augmented reality method and system for inspection of sewage treatment plant

By combining marker recognition and feature comparison with visual-inertial joint optimization, the problem of low efficiency of traditional inspections in sewage treatment plants has been solved, and efficient and accurate equipment identification and fault location in complex environments have been achieved, improving inspection efficiency and accuracy.

CN120656096APending Publication Date: 2025-09-16QINGDAO UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510793314.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Traditional manual inspection methods are inefficient in sewage treatment plants and make it difficult to quickly identify and locate equipment failures. Traditional visual recognition technology is susceptible to interference and failure in environments with dim light, water mist obstruction, and complex equipment layouts, and is unable to adapt to on-site differences.

Method used

A dual strategy of landmark recognition and feature comparison is adopted, with the help of landmarks to quickly and accurately associate device types and information. The scene-based pose estimation mechanism is used to ensure the accurate superposition of digital twin models and real devices. Combined with visual-inertial joint optimization, full-scene device recognition is achieved, and device features are extracted and compared with the database to cover scenes where landmarks are missing.

Benefits of technology

It achieves efficient and accurate equipment identification and fault location in complex environments, improves inspection efficiency and accuracy, provides real-time early warning functions, and reduces the impact of faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656096A_ABST
    Figure CN120656096A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital twinning, and provides an augmented reality method and system for inspection of a sewage treatment plant, and the method comprises the steps: carrying out the recognition of a marker for each frame in video data, and determining the type of equipment based on the marker and calling the information of the equipment if the marker is successfully recognized; if the marker is not identified, extracting the equipment characteristics of each frame and comparing the equipment characteristics with the equipment characteristic information in the database to determine the equipment type and call the equipment information; retrieving a digital twin model in a database based on the device type; according to a scene-based pose estimation mechanism, a digital twinborn model is superposed on a reality device. It is guaranteed that the digital twinborn model is continuously and accurately superposed to real equipment, virtual-real fusion visualization is achieved, and operation and maintenance personnel are assisted in visually mastering the state of the equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of digital twin technology, and in particular relates to an augmented reality method and system for sewage treatment plant inspection. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] As cities continue to expand, the demand for urban sewage treatment continues to grow, and accordingly, the number and scale of sewage treatment plants are also increasing. Faced with the ever-increasing scale of sewage treatment plants, relying solely on manual inspections is no longer sufficient to meet production needs. The sewage treatment process is complex, involving numerous steps and equipment. Traditional manual inspections require inspectors to invest significant time in reading, recording, and comparing equipment data. Failure of a device in the process often results in anomalies in multiple related data points, making it difficult for inspectors to accurately and promptly locate the problematic device, thus impacting treatment efficiency and production safety.

[0004] Augmented reality and digital twin technologies are effective solutions to these problems and represent a key trend in the development of intelligent industry. Augmented reality uses cameras, sensors, and other devices to overlay virtual information onto real-world scenes, providing an immersive experience. Since its introduction in the 1960s, the technology has been widely used in healthcare, education, manufacturing, and other fields, and has been deeply integrated with technologies such as 5G and AI. Digital twin technology replicates physical entities through digital models, enabling real-time simulation, optimization, and monitoring, improving efficiency and reducing costs. It is currently being applied in urban management, energy, transportation, and healthcare. Compared to traditional manual inspections, inspection methods based on augmented reality and digital twin technology are not restricted by physical entities and can quickly identify and locate anomalies in sewage treatment plant operations, significantly improving efficiency and accuracy.

[0005] However, in the process of matching the digital twin model with the actual equipment, traditional visual recognition is easily interfered and fails due to problems such as dim lighting (such as underground pump rooms), water mist obstruction (around the treatment unit), and complex equipment layout (intertwined pipes) in sewage treatment plants. Moreover, simply relying on model import cannot adapt to on-site differences. Summary of the Invention

[0006] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides an augmented reality method and system for sewage treatment plant inspection, which adopts a dual strategy of marker recognition + feature comparison. When there are markers, the model is used to quickly and accurately associate equipment types and information, which is efficient and accurate; when there are no markers, the equipment features are extracted and compared with the database, covering the scenes where markers are missing, ensuring that all scenes of equipment are identified without omission, and adapting to the inspection needs of various equipment in sewage treatment plants. Through the scene-based pose estimation mechanism, the digital twin model is ensured to be continuously and accurately superimposed on the real equipment, realizing the visualization of virtual and real fusion, and assisting operation and maintenance personnel to intuitively grasp the equipment status.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions: A first aspect of the present invention provides an augmented reality method for sewage treatment plant inspection, comprising: Obtain video data and IMU inertial data during sewage treatment plant inspections; For each frame of the video data, the landmark recognition model is used to identify the landmark. If the landmark is successfully identified, the device type is determined based on the landmark and the device information is retrieved. If the landmark is not identified, the device features of each frame are extracted and compared with the device feature information in the database to determine the device type and retrieve the device information. Retrieve digital twin models from the database based on device type; When there are identical markers in two adjacent frames, the feature points of the markers are extracted to estimate the device pose. Combined with the three-dimensional feature points of the markers in the digital twin model, the camera coordinate system and the world coordinate system are converted by solving the camera pose matrix to achieve alignment between the digital twin model and the real device. When the markers in two adjacent frames are different or the markers are not recognized, the initial pose of the device is estimated by extracting the feature points of each frame, and the pose increment is estimated based on the IMU inertial data. The pose increment is accumulated and fused with the initial pose of the device to obtain the absolute pose. After that, a globally consistent pose sequence is obtained through visual-inertial joint optimization to achieve alignment between the digital twin model and the real device. The digital twin model is superimposed on the real device.

[0008] Furthermore, it also includes: obtaining the operating data of real equipment, detecting abnormal values ​​of the operating data model, and when abnormal values ​​appear, marking the corresponding digital twin model as abnormal.

[0009] Furthermore, the marker recognition model includes a backbone network; After the backbone network processes the image frame through a convolution layer, it extracts multi-scale features through alternating processing of several convolution and feature fusion units, and compresses and fuses the extracted multi-scale features through spatial pyramid pooling and attention mechanism.

[0010] Furthermore, the landmark recognition model includes a neck network and a detection head; The multi-scale features output by the backbone network are obtained, and the multi-scale features include first-scale features, second-scale features, third-scale features, and fourth-scale features with gradually decreasing sizes; the neck network upsamples the fourth-scale features, fuses them with the third-scale features through a bidirectional feature pyramid, and processes them through C2f to obtain a first shallow feature; the first shallow feature is upsampled, fused with the second-scale features through a bidirectional feature pyramid, and processed through C2f to obtain a second shallow feature; the second shallow feature is upsampled, fused with the first-scale features through a bidirectional feature pyramid, and processed through C2f to obtain a second shallow feature; The scale features are fused and processed by C2f to obtain the third shallow feature; the third shallow feature is convolved, fused with the second scale feature and the second shallow feature through the bidirectional feature pyramid, and processed by C2f to obtain the first deep feature; the first deep feature is convolved, fused with the third scale feature and the first shallow feature through the bidirectional feature pyramid, and processed by C2f to obtain the second deep feature; the second deep feature is convolved, fused with the fourth scale feature through the bidirectional feature pyramid, and processed by C2f to obtain the third deep feature; The detection head detects the third shallow feature, the first deep feature, the second deep feature and the third deep feature respectively.

[0011] A second aspect of the present invention provides an augmented reality system for sewage treatment plant inspection, comprising: A data acquisition module configured to: acquire video data and IMU inertial data during a sewage treatment plant inspection; A device identification module is configured to: perform landmark recognition on each frame of the video data using a landmark recognition model; if the landmark is successfully recognized, determine the device type based on the landmark and retrieve device information; if no landmark is recognized, extract the device features of each frame and compare them with the device feature information in the database to determine the device type and retrieve device information; A digital twin module is configured to: retrieve a digital twin model from a database based on a device type; The tracking and registration module is configured as follows: when the same marker exists in two adjacent frames, the feature points of the marker are extracted to estimate the device pose, and the three-dimensional feature points of the marker in the digital twin model are combined to solve the camera pose matrix to convert the camera coordinate system and the world coordinate system to achieve alignment between the digital twin model and the real device; when the markers in two adjacent frames are different or the markers are not recognized, the initial pose of the device is estimated by extracting the feature points of each frame, and the pose increment is estimated based on the IMU inertial data, and the pose increment is accumulated and fused with the initial pose of the device to obtain the absolute pose, and then a globally consistent pose sequence is obtained through visual-inertial joint optimization to achieve alignment between the digital twin model and the real device; the digital twin model is superimposed on the real device.

[0012] Furthermore, the digital twin module is also configured to: obtain the operating data of the real device, detect abnormal values ​​of the operating data model, and when abnormal values ​​appear, mark the corresponding digital twin model as abnormal.

[0013] Furthermore, the marker recognition model includes a backbone network; After the backbone network processes the image frame through a convolution layer, it extracts multi-scale features through alternating processing of several convolution and feature fusion units, and compresses and fuses the extracted multi-scale features through spatial pyramid pooling and attention mechanism.

[0014] Furthermore, the landmark recognition model includes a neck network and a detection head; The multi-scale features output by the backbone network are obtained, and the multi-scale features include first-scale features, second-scale features, third-scale features, and fourth-scale features with gradually decreasing sizes; the neck network upsamples the fourth-scale features, fuses them with the third-scale features through a bidirectional feature pyramid, and processes them through C2f to obtain a first shallow feature; the first shallow feature is upsampled, fused with the second-scale features through a bidirectional feature pyramid, and processed through C2f to obtain a second shallow feature; the second shallow feature is upsampled, fused with the first-scale features through a bidirectional feature pyramid, and processed through C2f to obtain a second shallow feature; The scale features are fused and processed by C2f to obtain the third shallow feature; the third shallow feature is convolved, fused with the second scale feature and the second shallow feature through the bidirectional feature pyramid, and processed by C2f to obtain the first deep feature; the first deep feature is convolved, fused with the third scale feature and the first shallow feature through the bidirectional feature pyramid, and processed by C2f to obtain the second deep feature; the second deep feature is convolved, fused with the fourth scale feature through the bidirectional feature pyramid, and processed by C2f to obtain the third deep feature; The detection head detects the third shallow feature, the first deep feature, the second deep feature and the third deep feature respectively.

[0015] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the augmented reality method for sewage treatment plant inspection as described above.

[0016] The fourth aspect of the present invention provides a computer device comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, wherein when the processor executes the program, the steps of the augmented reality method for sewage treatment plant inspection as described above are implemented.

[0017] Compared with the prior art, the present invention has the following beneficial effects: This invention adopts the dual strategy of marker recognition + feature comparison. When there are markers, the model quickly and accurately associates the device type and information, which is efficient and accurate. When there are no markers, the device features are extracted and compared with the database, covering the scenarios where markers are missing (such as damaged markers and undeployed areas), ensuring that all scenarios are recognized without omissions, and adapting to the diverse equipment inspection needs of sewage treatment plants. The present invention designs a scene-based pose estimation mechanism. When adjacent frames have the same markers, the marker feature points are used to solve the camera pose matrix, and the three-dimensional information of the markers in the digital twin is used for fast and accurate alignment. When the markers are different or there are no markers, the initial pose is estimated first, and then the equipment feature matching is optimized to cope with complex situations such as marker changes and occlusions during inspections, ensuring that the digital twin model is continuously and accurately superimposed on the real equipment, realizing virtual-reality fusion visualization, and assisting operation and maintenance personnel to intuitively grasp the equipment status.

[0018] This method collects real-world equipment operating data and detects outliers, then associates them with a digital twin model to annotate these anomalies. This transforms physical equipment anomalies into a virtual model, visualizing them for rapid location and identification of faulty equipment. Compared to traditional manual inspections, which rely on post-facto discovery, this method provides real-time early warning, speeding up the response to equipment failures in sewage treatment plants and minimizing their impact.

[0019] The backbone network of this invention extracts image features at multiple scales through a combination of convolutional layers, alternating convolution and feature fusion units, spatial pyramid pooling, and an attention mechanism. The convolution and feature fusion units enhance the feature hierarchy, spatial pyramid pooling expands the receptive field, and the attention mechanism focuses on key features. This allows landmark recognition to accurately capture landmark details even in complex sewage plant environments (such as pipes and sewage interference), improving recognition accuracy.

[0020] The neck network of this invention generates multi-scale deep and shallow features through multiple rounds of upsampling, bidirectional feature pyramid fusion, and C2f processing. The detection head detects features at different scales separately. This adapts to markers of varying sizes and distances in sewage treatment plants (such as large equipment logos and small instrument labels), effectively identifying both close-up shots and distant small objects, ensuring full coverage of equipment recognition scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0022] Figure 1 This is a flow chart of an augmented reality method for sewage treatment plant inspection according to the first embodiment of the present invention; Figure 2 This is a flowchart of the tracking and registration operation of the first embodiment of the present invention; Figure 3 This is a diagram of the improved model structure based on YOLOv8s according to the first embodiment of the present invention; Figure 4 It is a structural diagram of a computer device according to the fourth embodiment of the present invention. DETAILED DESCRIPTION

[0023] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0024] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0025] Example 1 This embodiment provides an augmented reality method for sewage treatment plant inspection.

[0026] This embodiment provides an augmented reality method for sewage treatment plant inspections, which uses a visual-inertial SLAM algorithm to achieve virtual tracking, completes three-dimensional registration through a hybrid method of markerless and marker coexistence, uses a deep learning-based computer vision algorithm for target recognition, and projects the digital twin model onto the corresponding real-world device through augmented reality technology, thereby helping inspectors to quickly complete their inspection work.

[0027] This embodiment provides an augmented reality method for sewage treatment plant inspection, such as Figure 1 As shown, the following steps are included: S1. Real-time data collection.

[0028] In this embodiment, real-time data acquisition utilizes visual cameras and IMUs. Visual cameras include, but are not limited to, monocular cameras, binocular cameras, and RGB-D cameras; IMUs include, but are not limited to, accelerometers and gyroscopes. These sensors enable real-time acquisition of video data and IMU inertial data.

[0029] In this embodiment, a visual camera captures real-time video data within the user's field of view, which serves as the basic input for aligning the digital twin model with the real environment. Simultaneously, an IMU collects real-time acceleration and angular velocity data during the user's movement, which serves as input for tracking and registration. S2. Device identification.

[0030] Equipment identification combines computer vision and deep learning technologies to perform real-time detection and identification of sewage treatment equipment in collected video data.

[0031] like Figure 2 As shown, the steps of device identification include: first, feature extraction and feature matching are performed on the video data to detect whether the marker exists; if the marker is successfully identified, the device information and the marker boundary box in the two image frames are directly passed to step S3; when the marker is not detected, the device type is determined by extracting the overall features of the device and comparing them with the device feature information in the database, and then the device information is passed to step S3.

[0032] Among them, the markers are pre-pasted on the sewage treatment equipment.

[0033] Landmark recognition uses the Path Aggregation Network (PANet), and the PANet structure is optimized based on YOLOv8s: a bidirectional weighted feature pyramid network (BiFPN) structure is adopted; at the same time, an additional 160×160 detection head is added to the detection layer and deeply integrated with the BiFPN structure; the feature splicing method in the neck network is improved, using the BiFPN splicing mechanism, and an upsampling operation is added to the seventeenth layer to obtain shallower feature information, thereby better integrating the feature expression of small targets; after feature splicing and C2f module processing, the feature information is input into the newly added P2 detection layer (Detect); the P2 detection layer can process higher-resolution feature maps, further improving the detection accuracy of small targets and effectively reducing the occurrence of missed and false detections.

[0034] like Figure 3As shown in the figure, the landmark recognition model consists of the Backbone network, the Neck network, and the Head. A 640×640×3 RGB image is input; the Backbone first extracts basic features, the Neck then fuses multi-scale features, and the Head finally outputs detection results at different scales.

[0035] (1) Backbone (backbone feature extraction, refining features layer by layer): It consists of Conv (convolution) + C2f (feature fusion unit) + SPPF (spatial pyramid pooling) + EMA, and progresses from "downsampling → feature extraction → compression fusion": Initial convolution: 2 layers of Conv first reduce the dimensionality of the input image (compress the spatial size and increase the number of channels) and initially extract basic features such as edges and textures; C2f alternating convolution: Multiple C2f modules (similar to CSPNet variants, with the core being feature splitting and fusion) are used to enhance feature extraction capabilities while controlling computational complexity. SPPF Compression: Use SPPF for spatial pyramid pooling, using multi-scale MaxPool2d + Concat to quickly compress feature map size and expand the receptive field (allowing the model to "see" a wider range of context). EMA Conclusion: For Backbone output features, the EMA attention mechanism significantly enhances the ability to extract key features by dynamically adjusting feature channel weights. It also effectively suppresses the interference of background noise and improves the robustness of the model in complex scenarios.

[0036] The C2f microstructure includes: input → Conv (convolutional dimensionality reduction) → Split (feature diversion, split into two paths); one path: directly transmit basic features; the other path: through n Bottleneck layers (bottleneck layers, first compress the channel → then expand it, reducing the amount of computation while retaining key features) → Concat (splicing the two features) → final Conv (convolutional integration) output.

[0037] Function: Through "diversion + bottleneck layer + splicing", it efficiently integrates features at different levels, preserving details (shallow diversion) and enhancing semantics (bottleneck layer extraction).

[0038] (2) Neck (feature fusion, BiFPN-dominated multi-scale connection).

[0039] Core logic: BiFPN (Bidirectional Feature Pyramid).

[0040] Goal: To fuse the multi-scale features output by Backbone (feature maps of different layers have different sizes and semantics, such as shallow layers with more fine-grained details and deep layers with stronger semantic information), so that the model can simultaneously use "details + semantics" to detect objects of different sizes.

[0041] Top-down: Starting from the deep large receptive field features (small-size feature maps), they are unsampled (upsampled and restored to size), concat-fused with the features of the previous layer (e.g., BiFPN Concat module), and passed to the shallower layers.

[0042] Bottom-up: The fused shallow features are then processed by C2f + Conv and fused with the deep features for a second time to strengthen cross-level information interaction.

[0043] Loop structure: Repeat "upsampling fusion → downsampling enhancement" multiple times (the three groups of BiFPN processes in the figure), gradually refine the multi-scale features, and output four groups of feature maps of different scales (corresponding to the four detection branches of the Head).

[0044] (3) Head (detection head, output target prediction).

[0045] Corresponding to the 4 sets of feature maps output by Neck, each branch performs Detect independently to adapt to targets of different scales: The top layer Detect: processes the largest feature map (with the deepest semantic information) and is suitable for detecting small targets at long distances (such as the detection map at the bottom right, where the vehicle is smaller and farther away). The bottom layer Detect: processes the smallest feature map (with the richest details) and is suitable for detecting large targets at close range (such as the detection map at the top right, where the vehicle is larger and closer).

[0046] The Detect microstructure (classification + regression) consists of the following steps: input features → two-way Conv (convolutional branch): the Bbox Loss branch predicts the target bounding box coordinates (x, y, w, h), using regression loss to optimize positioning accuracy; the CLs.Loss branch predicts the target category probability (such as vehicle, pedestrian, etc.), using classification loss to optimize category judgment.

[0047] Function: Through the "parallel convolution branch", the position and category of the target are output simultaneously to complete the detection task.

[0048] SPPF (Spatial Pyramid Pooling, efficient compression): Input features → Convolution (dimensionality reduction) → 3 layers of MaxPool2d (maximum pooling, using 5×5, 9×9, and 13×13 pooling kernels, respectively, to achieve multi-scale pooling through padding + stride) → Concat (concatenating multi-scale pooling results) → Final Convolution (convergence) output. This method minimizes computational effort, rapidly expanding the receptive field (allowing the model to "see" a wider range of scenes) while preserving multi-scale features and adapting to objects of varying sizes.

[0049] The standard Conv pipeline (basic feature transformation) follows: Conv2d (convolutional layer, extracting spatial features) → BatchNorm2d (batch normalization, stabilizing training and reducing gradient oscillation) → SiLU (activation function, a Sigmoid variant, introducing nonlinearity and providing smoother results than ReLU). Function: Each convolution layer undergoes the "convolution → normalization → activation" cycle to ensure the stability of feature transformation and the ability to express nonlinearities.

[0050] S3. Digital twin.

[0051] By accurately measuring the operating equipment of the sewage treatment plant, a high-precision independent BIM model is built for each device, and then a digital twin model of the entire plant equipment is built based on these BIM models.

[0052] A 3D digital twin model was constructed using Unity3D software, using high-precision BIM models of sewage treatment plant equipment. All 3D digital twin models are stored in a database. The database uses an efficient index structure, supporting fast retrieval and dynamic updates, enabling rapid and accurate retrieval of the digital twin model corresponding to the identified equipment.

[0053] The digital twin model uses C# scripts to achieve bidirectional communication with the data synchronization module. In actual use, the data synchronization module collects device operating data from sensors, transmits this data via the PLC, and saves it to the database. The digital twin model uses C# scripts to periodically obtain real-time data from the database and synchronize the data with the digital twin model. The augmented reality display interface intuitively presents real-time operating data in a table format in the upper left corner of the user interface.

[0054] When an outlier appears in the collected data, the device's digital twin model is marked as abnormal. When the flow rate of a device exceeds the normal range, the color of the corresponding device's digital twin model will be marked red.

[0055] After receiving the device type identified in step S2, the database is queried and the retrieved digital twin model and the bounding box data passed in step S2 are provided to step S4 for use.

[0056] S4. Track registration.

[0057] The real-world coordinate system is constructed using the visual-inertial joint initialization algorithm through real-time collected video data and IMU inertial data.

[0058] The IMU part uses the pre-integration method and combines the acceleration and angular velocity The data is used to estimate the pose increment, and the pre-integration calculation formula is as follows: ; ; ; in, 、 、 Represents the rotation increment, speed increment and position increment respectively, is the starting time of the pre-integration window, is the end time of the pre-integration window, The duration between two IMU sampling intervals. is an exponential map, For the The deviation of the angular velocity at the moment, For the The measurement noise of the angular velocity at each moment is For the The measured value of the angular velocity at a moment, For the The measured value of IMU acceleration at each sampling moment, is the deviation of acceleration at the kth moment, For the The measurement noise of acceleration at each moment, Indicates from time At the time The rotation increment, Indicates from time At the time speed increment.

[0059] The visual component selects the corresponding tracking and registration method based on the natural feature information of the landmark or device identified in step S2. When the same landmark is identified in two adjacent image frames, the ORB feature points of the landmark within the bounding box of the two adjacent image frames, transmitted in step S3, are used in conjunction with the PnP algorithm to perform pose estimation. Combined with the three-dimensional feature points of the landmark in the digital twin model, the camera pose matrix is ​​solved to achieve initial alignment between the camera coordinate system and the world coordinate system. When the landmarks in the two adjacent image frames are different or no landmark is identified, the initial pose estimation is performed by extracting the ORB feature points of the device in the image and using the PnP algorithm. The camera pose matrix is ​​solved to achieve initial alignment between the camera coordinate system and the world coordinate system. After obtaining the initial pose, the pose increments obtained by IMU pre-integration between the two frames are cumulatively fused with the PnP estimation results to update the absolute pose of the current frame. This fused pose is then used as input for the subsequent visual-inertial joint optimization.

[0060] After integrating visual initialization with inertial data, the system first performs joint visual-inertial optimization within a local sliding window. Leveraging visual observations between adjacent frames (such as feature matching and reprojection errors) and IMU pre-integration constraints, the system adjusts the rotation matrix and displacement vector of each keyframe within the window. This reduces the visual residual error and the IMU residual error simultaneously, eliminating short-term drift and ensuring local pose consistency. After completing this window optimization, the pose of the keyframe in the current window is output to subsequent modules.

[0061] After the sliding window optimization, the system will perform loop detection on the output frame: the current key frame and the historical key frame library will be quickly searched to screen out candidate frames with high similarity, and then the matching pairs between the candidate frames and the current frame will be verified for geometric consistency, and the loop constraints that are established will be added to the global pose graph; if no loop is detected, the local pose output by the sliding window will be directly retained and the next round of window sliding will continue.

[0062] Once a loop is established, the system constructs a global pose graph consisting of all keyframe nodes, visual-inertial constraints for adjacent frames, and the newly added loop constraints. Through graph optimization, the global poses of all keyframes are jointly adjusted to correct the accumulated drift of the entire trajectory, thereby outputting a globally consistent pose sequence with minimal loop closure error. If no valid constraints are found during the loop closure detection phase, global optimization is skipped and the local results from the sliding window phase are used.

[0063] After the global map is optimized and the final globally consistent pose sequence is output, the system will align these poses with the existing point cloud map. The ICP (Iterative Closest Point) algorithm is used to achieve high-precision point cloud registration. The formula is: ; in, and are the corresponding points in the source point cloud and the target point cloud, Represents the number of target point clouds and original point clouds, and points Since visual and inertial sensors may experience cumulative errors or positioning drift in dynamic and complex environments, ICP further optimizes the accuracy of 3D registration by fine-tuning the positional relationship between the source and target point clouds, ensuring precise alignment of the digital twin model with the real-world device position.

[0064] S5. Data synchronization.

[0065] Data synchronization involves four steps: data acquisition, transmission, processing, and distribution. Various sensors installed on sewage treatment equipment (such as temperature and pressure sensors) are connected to the PLC via a 485 bus. The PLC collects sensor data sequentially through a polling mechanism, performs preliminary analysis and packaging, and generates data frames that conform to a custom protocol.

[0066] The PLC sends packaged data frames to the DTU (Data Transmission Unit) via the 485 interface. After receiving the data frames, the DTU establishes a connection with the data server using the TCP / IP protocol and sends the data to the data server. The DTU is equipped with a heartbeat detection mechanism to monitor communication status in real time. If a communication anomaly or data packet loss is detected, the DTU triggers a retransmission mechanism, requesting the PLC to resend the unacknowledged data frames to ensure data integrity.

[0067] The data server receives operational data from the DTU via the OPC / UA protocol. Upon receiving the data, the server first formats the raw data, converting it into a standardized JSON format for subsequent processing and use. The server then annotates any abnormal data that falls outside the normal range to ensure data accuracy.

[0068] Finally, the processed data is stored in a local MySQL database, and the index table is updated in real time. The augmented reality display module sends data requests to the server through a RESTful interface. The server retrieves the corresponding data and returns it to the display, realizing dynamic visualization of the digital twin model.

[0069] S5. Augmented reality display module.

[0070] The augmented reality display module includes a high-resolution display device. High-resolution display devices include but are not limited to tablets, smartphones, AR glasses, or head-mounted displays. These devices enable precise alignment of the digital twin model with the real environment and dynamic visualization.

[0071] The real-time pose data of the digital twin model calculated in step S4 can continuously update the pose of the device, ensure that the pose of the digital twin model is always accurate, and achieve accurate registration and virtual tracking of the digital twin model.

[0072] Supports user interaction. When a user makes a gesture to "zoom in" or "rotate," the system recognizes the user's gesture and generates corresponding operation instructions, thereby triggering the corresponding operation of the digital twin model, such as rotation, scaling, or switching displays, to achieve natural and efficient human-computer interaction in the scene.

[0073] With the above functions, it can provide real-time equipment status and information display functions, provide users with a comfortable and fast operating experience, and effectively improve the efficiency of inspections.

[0074] This embodiment realizes virtual tracking and three-dimensional registration based on marker assistance by identifying markers pre-pasted on the sewage treatment equipment. When the same marker is detected in two adjacent image frames, the digital twin model will be virtually tracked and three-dimensionally registered based on the features of the markers identified in the two image frames; when the markers in the two adjacent image frames are different or no markers are identified, virtual tracking and three-dimensional registration based on natural features will be performed. The registered digital twin model is projected into the real scene through augmented reality display to achieve high-precision registration and tracking. Compared with extracting natural features, artificial markers have richer and clearer features, so feature matching based on markers can significantly improve the speed and accuracy of the system.

[0075] This example uses precise measurements of operating sewage treatment plant equipment to construct high-precision, independent BIM models. Based on these models, a digital twin model of the entire plant is generated. The digital twin model is embedded in a C# script, enabling two-way data communication with the server, synchronizing real-time data and presenting this data in a tabular form on the user interface. Detected abnormal data points are automatically identified and the corresponding BIM model of the equipment or pipeline is marked in red, significantly improving the efficiency of sewage treatment plant inspections.

[0076] Example 2 This embodiment provides an augmented reality system for sewage treatment plant inspection, which specifically includes: A data acquisition module configured to: acquire video data and IMU inertial data during a sewage treatment plant inspection; A device identification module is configured to: perform landmark recognition on each frame of the video data using a landmark recognition model; if the landmark is successfully recognized, determine the device type based on the landmark and retrieve device information; if no landmark is recognized, extract the device features of each frame and compare them with the device feature information in the database to determine the device type and retrieve device information; A digital twin module is configured to: retrieve a digital twin model from a database based on a device type; The tracking and registration module is configured as follows: when the same marker exists in two adjacent frames, the feature points of the marker are extracted to estimate the device pose, and the three-dimensional feature points of the marker in the digital twin model are combined to solve the camera pose matrix to convert the camera coordinate system and the world coordinate system to achieve alignment between the digital twin model and the real device; when the markers in two adjacent frames are different or the markers are not recognized, the initial pose of the device is estimated by extracting the feature points of each frame, and the pose increment is estimated based on the IMU inertial data, and the pose increment is accumulated and fused with the initial pose of the device to obtain the absolute pose, and then a globally consistent pose sequence is obtained through visual-inertial joint optimization to achieve alignment between the digital twin model and the real device; the digital twin model is superimposed on the real device.

[0077] Furthermore, the digital twin module is also configured to: obtain the operating data of the real device, detect abnormal values ​​of the operating data model, and when abnormal values ​​appear, mark the corresponding digital twin model as abnormal.

[0078] Furthermore, the marker recognition model includes a backbone network; After the backbone network processes the image frame through a convolution layer, it extracts multi-scale features through alternating processing of several convolution and feature fusion units, and compresses and fuses the extracted multi-scale features through spatial pyramid pooling and attention mechanism.

[0079] Furthermore, the landmark recognition model includes a neck network and a detection head; The multi-scale features output by the backbone network are obtained, and the multi-scale features include first-scale features, second-scale features, third-scale features, and fourth-scale features with gradually decreasing sizes; the neck network upsamples the fourth-scale features, fuses them with the third-scale features through a bidirectional feature pyramid, and processes them through C2f to obtain a first shallow feature; the first shallow feature is upsampled, fused with the second-scale features through a bidirectional feature pyramid, and processed through C2f to obtain a second shallow feature; the second shallow feature is upsampled, fused with the first-scale features through a bidirectional feature pyramid, and processed through C2f to obtain a second shallow feature; The scale features are fused and processed by C2f to obtain the third shallow feature; the third shallow feature is convolved, fused with the second scale feature and the second shallow feature through the bidirectional feature pyramid, and processed by C2f to obtain the first deep feature; the first deep feature is convolved, fused with the third scale feature and the first shallow feature through the bidirectional feature pyramid, and processed by C2f to obtain the second deep feature; the second deep feature is convolved, fused with the fourth scale feature through the bidirectional feature pyramid, and processed by C2f to obtain the third deep feature; The detection head detects the third shallow feature, the first deep feature, the second deep feature and the third deep feature respectively.

[0080] It should be noted here that the various modules in this embodiment correspond one-to-one to the various steps in Example 1, and the specific implementation processes are the same, which will not be repeated here.

[0081] Example 3 This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the augmented reality method for sewage treatment plant inspection as described in the first embodiment above are implemented.

[0082] Example 4 This embodiment provides a computer device, such as Figure 4 As shown, the system includes a computer-readable storage medium 1003, a processor 1001, a communication interface 1002, and a computer program stored on the computer-readable storage medium 1003 and executable on the processor 1001. The processor 1001, the communication interface 1002, and the computer-readable storage medium 1003 may be connected via a bus or other means. The communication interface 1002 is configured to receive and transmit data, and when the processor 1001 executes the program, the steps of the augmented reality method for sewage treatment plant inspection described in the first embodiment are implemented.

[0083] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. An augmented reality method for sewage treatment plant inspection, characterized in that: include: Obtain video data and IMU inertial data during sewage treatment plant inspections; For each frame of the video data, the landmark recognition model is used to identify the landmark. If the landmark is successfully identified, the device type is determined based on the landmark and the device information is retrieved. If the landmark is not identified, the device features of each frame are extracted and compared with the device feature information in the database to determine the device type and retrieve the device information. Retrieve digital twin models from the database based on device type; When there are identical markers in two adjacent frames, the feature points of the markers are extracted to estimate the device pose. Combined with the three-dimensional feature points of the markers in the digital twin model, the camera coordinate system and the world coordinate system are converted by solving the camera pose matrix to achieve alignment between the digital twin model and the real device. When the markers in two adjacent frames are different or the markers are not recognized, the initial pose of the device is estimated by extracting the feature points of each frame, and the pose increment is estimated based on the IMU inertial data. The pose increment is accumulated and fused with the initial pose of the device to obtain the absolute pose. After that, a globally consistent pose sequence is obtained through visual-inertial joint optimization to achieve alignment between the digital twin model and the real device. The digital twin model is superimposed on the real device.

2. The augmented reality method for sewage treatment plant inspection according to claim 1, characterized in that: Also includes: Obtain the operating data of real equipment, detect abnormal values ​​of the operating data model, and when abnormal values ​​appear, mark the corresponding digital twin model as abnormal.

3. The augmented reality method for sewage treatment plant inspection according to claim 1, characterized in that: The marker recognition model includes a backbone network; After the backbone network processes the image frame through a convolution layer, it extracts multi-scale features through alternating processing of several convolution and feature fusion units, and compresses and fuses the extracted multi-scale features through spatial pyramid pooling and attention mechanism.

4. The augmented reality method for sewage treatment plant inspection according to claim 1, characterized in that: The landmark recognition model includes a neck network and a detection head; The multi-scale features output by the backbone network are obtained, and the multi-scale features include first-scale features, second-scale features, third-scale features, and fourth-scale features with gradually decreasing sizes; the neck network upsamples the fourth-scale features, fuses them with the third-scale features through a bidirectional feature pyramid, and processes them through C2f to obtain a first shallow feature; the first shallow feature is upsampled, fused with the second-scale features through a bidirectional feature pyramid, and processed through C2f to obtain a second shallow feature; the second shallow feature is upsampled, fused with the first-scale features through a bidirectional feature pyramid, and processed through C2f to obtain a second shallow feature; The scale features are fused and processed by C2f to obtain the third shallow feature; the third shallow feature is convolved, fused with the second scale feature and the second shallow feature through the bidirectional feature pyramid, and processed by C2f to obtain the first deep feature; the first deep feature is convolved, fused with the third scale feature and the first shallow feature through the bidirectional feature pyramid, and processed by C2f to obtain the second deep feature; the second deep feature is convolved, fused with the fourth scale feature through the bidirectional feature pyramid, and processed by C2f to obtain the third deep feature; The detection head detects the third shallow feature, the first deep feature, the second deep feature and the third deep feature respectively.

5. An augmented reality system for sewage treatment plant inspection, characterized in that: include: A data acquisition module configured to: acquire video data and IMU inertial data during a sewage treatment plant inspection; A device identification module is configured to: perform landmark recognition on each frame of the video data using a landmark recognition model; if the landmark is successfully recognized, determine the device type based on the landmark and retrieve device information; if no landmark is recognized, extract the device features of each frame and compare them with the device feature information in the database to determine the device type and retrieve device information; A digital twin module is configured to: retrieve a digital twin model from a database based on a device type; The tracking and registration module is configured as follows: when the same marker exists in two adjacent frames, the feature points of the marker are extracted to estimate the device pose, and the three-dimensional feature points of the marker in the digital twin model are combined to solve the camera pose matrix to convert the camera coordinate system and the world coordinate system to achieve alignment between the digital twin model and the real device; when the markers in two adjacent frames are different or the markers are not recognized, the initial pose of the device is estimated by extracting the feature points of each frame, and the pose increment is estimated based on the IMU inertial data, and the pose increment is accumulated and fused with the initial pose of the device to obtain the absolute pose, and then a globally consistent pose sequence is obtained through visual-inertial joint optimization to achieve alignment between the digital twin model and the real device; the digital twin model is superimposed on the real device.

6. The augmented reality system for sewage treatment plant inspection according to claim 5, characterized in that: The digital twin module is also configured to: obtain the operating data of the real device, detect abnormal values ​​of the operating data model, and when abnormal values ​​appear, mark the corresponding digital twin model as abnormal.

7. The augmented reality system for sewage treatment plant inspection according to claim 5, characterized in that: The marker recognition model includes a backbone network; After the backbone network processes the image frame through a convolution layer, it extracts multi-scale features through alternating processing of several convolution and feature fusion units, and compresses and fuses the extracted multi-scale features through spatial pyramid pooling and attention mechanism.

8. The augmented reality system for sewage treatment plant inspection according to claim 5, characterized in that: The landmark recognition model includes a neck network and a detection head; The multi-scale features output by the backbone network are obtained, and the multi-scale features include first-scale features, second-scale features, third-scale features, and fourth-scale features with gradually decreasing sizes; the neck network upsamples the fourth-scale features, fuses them with the third-scale features through a bidirectional feature pyramid, and processes them through C2f to obtain a first shallow feature; the first shallow feature is upsampled, fused with the second-scale features through a bidirectional feature pyramid, and processed through C2f to obtain a second shallow feature; the second shallow feature is upsampled, fused with the first-scale features through a bidirectional feature pyramid, and processed through C2f to obtain a second shallow feature; The scale features are fused and processed by C2f to obtain the third shallow feature; the third shallow feature is convolved, fused with the second scale feature and the second shallow feature through the bidirectional feature pyramid, and processed by C2f to obtain the first deep feature; the first deep feature is convolved, fused with the third scale feature and the first shallow feature through the bidirectional feature pyramid, and processed by C2f to obtain the second deep feature; the second deep feature is convolved, fused with the fourth scale feature through the bidirectional feature pyramid, and processed by C2f to obtain the third deep feature; The detection head detects the third shallow feature, the first deep feature, the second deep feature and the third deep feature respectively.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the augmented reality method for sewage treatment plant inspection according to any one of claims 1 to 4 are implemented.

10. A computer device comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, wherein: When the processor executes the program, the steps of the augmented reality method for sewage treatment plant inspection according to any one of claims 1 to 4 are implemented.