Video real scene map management and dispatching system and method based on AR technology

By using a software-based AR technology system, combined with visual SLAM and IMU data, dynamic label following and multi-source data fusion were achieved, solving the problems of high-cost hardware dependence and poor interactivity, and improving management efficiency and continuity.

CN122334709APending Publication Date: 2026-07-03CHINA TOWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA TOWER CO LTD
Filing Date
2026-05-26
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing AR camera systems are expensive, have strong hardware dependencies, are incompatible with existing cameras, have static tag management leading to information loss, and have poor interactivity due to the separation of video and map.

Method used

The video real-scene map management and scheduling system based on AR technology implements AR functions through software, combines visual SLAM technology, IMU data and visual information, dynamically follows tags, and supports multi-source data fusion and real-time interaction.

Benefits of technology

It reduces deployment costs, ensures that tags automatically follow camera movement, enhances management continuity and interactivity, and improves management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122334709A_ABST
    Figure CN122334709A_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of AR technology, and particularly relates to a video real scene map management and scheduling system and method based on AR technology. The present disclosure improves the adaptability of the video label, ensures that the label automatically follows the camera movement, enhances the management continuity and reliability, solves the problem of passive watching and lack of interaction of the video monitoring system, realizes dynamic business information labeling and real-time interaction of the video picture, and improves the management efficiency. The present disclosure is decoupled from the hardware, uses the existing stock cameras, avoids the purchase of professional AR equipment, reduces the hardware cost, and the software deployment shortens the project cycle. The label following technology is based on a space perception algorithm, ensures the information continuity, reduces the human intervention, improves the system reliability, and is particularly suitable for high-speed change scenes such as emergency command. Through the AR label dynamic management and interactive query, the video is upgraded from passive monitoring to active management tool, the management efficiency is improved by more than 50%, and the problem of poor interaction of the traditional system is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of AR technology, and in particular to a video real-scene map management and scheduling system and method based on AR technology. Background Technology

[0002] Existing technologies, such as video annotation systems based on professional AR cameras, achieve AR tag overlay through high-end hardware. However, this requires the purchase of dedicated equipment, resulting in high costs and incompatibility with existing cameras. The drawbacks are strong hardware dependence, high costs, and inability to utilize existing video equipment.

[0003] A static video tag management system fixes tags to specific positions on the video screen. When the camera pans or zooms, the tags cannot follow the target, resulting in information loss. The drawback is that the static tag management makes it susceptible to information misalignment due to camera movement.

[0004] Traditional video surveillance platforms only support basic preview and playback, lacking map fusion and interactive query functions, requiring users to switch between multiple systems. A drawback is the separation of video and map, resulting in poor interactivity.

[0005] In summary, this disclosure provides a hardware-decoupled solution that implements AR functionality through software, reducing deployment costs; enables dynamic tag following to ensure continuous visibility of business information; and integrates video real-scene maps with AR tags to support one-stop command and dispatch. Summary of the Invention

[0006] To address the aforementioned issues, this disclosure provides a video real-scene map management and scheduling system and method based on AR technology.

[0007] Firstly, a video real-scene map management and scheduling system based on AR technology includes: a terminal presentation unit, a business capability unit, a technical support unit, and an infrastructure unit; The terminal presentation unit is used to display the global video real-view map and interactive operations; The business capability unit includes a video intelligent access module, an AR tag management engine, a real-scene map fusion engine, a command and dispatch engine, and a data processing module. The video intelligent access module acquires video streams, the AR tag management engine overlays AR tags with spatial location information onto the video screen, and the real-scene map fusion engine fuses and displays multi-source heterogeneous data to achieve command and dispatch based on real-scene maps. The technical support unit is used to provide video streaming media processing, spatial location calculation, and real-time data communication capabilities; Infrastructure units are used to provide computing, networking, and storage resources.

[0008] Furthermore, the AR tag management engine employs visual SLAM technology to achieve dynamic tag following, specifically including: A hybrid visual odometry method combining feature point method and direct method is used for pose estimation, and a dynamic feature point filtering mechanism is introduced to eliminate interference from dynamic objects. A visual-inertial fusion framework is adopted to combine IMU data with visual information for localization; A loop closure detection mechanism combining semantic information and geometric verification is adopted to detect loop closures by visually recognizing the distribution of semantic objects and their spatial relationships in the scene.

[0009] Furthermore, a hybrid visual odometry method combining feature point method and direct method is used for pose estimation, and a dynamic feature point filtering mechanism is introduced to eliminate interference from dynamic objects, including: For pose estimation, preliminary pose estimation is performed through ORB feature extraction and optical flow tracking, and then the estimation results are optimized using a direct method. For processing dynamic objects, the motion consistency of dynamic feature points is analyzed, and feature points with motion consistency are identified through a deep learning network; dynamic feature points identified as dynamic objects are removed, and only static objects are used for localization.

[0010] Furthermore, localization is achieved by combining IMU data with visual information, including: When the image is blurry or obstructed, the position is predicted using an inertial measurement unit (IMU). When the image is clear and the inertial measurement unit (IMU) shows drift, visual position correction is used.

[0011] Furthermore, the AR tag management engine also includes: A multi-class label recognition is achieved using an improved YOLOv5 architecture. An adaptive spatial feature fusion mechanism is introduced into the feature pyramid network, and feature maps of different scales are fused through weight learning. A labeled multi-Bernoulli filter (LMB) is used as the multi-sensor data fusion framework, and distributed optimization fusion is performed in combination with the generalized covariance intersection fusion rule (GCI). Pose graph optimization is adopted, which models the relationship between label position, camera pose and map points as a graph structure, and achieves global consistency optimization by minimizing the error function.

[0012] Furthermore, the video intelligent access module supports protocol adaptive technology, automatically identifying device fingerprint characteristics and selecting the corresponding protocol for access; the video stream processing adopts adaptive bitrate technology, dynamically adjusting the video bitrate according to network bandwidth.

[0013] Furthermore, the real-scene map fusion engine supports the seamless fusion of 2D electronic maps, 3D oblique photogrammetry models, point cloud data, and panoramic images; Two-dimensional electronic maps are based on a tile pyramid structure; The 3D oblique photogrammetry model uses Level of Detail (LOD) technology to dynamically adjust the model's accuracy based on the viewpoint distance. VR panoramic image generation is based on the SIFT feature matching algorithm. It generates panoramic images through image preprocessing, feature matching, image registration, and fusion optimization steps.

[0014] Furthermore, VR panoramic image generation is based on the SIFT feature matching algorithm. Through image preprocessing, feature matching, image registration, and fusion optimization steps, panoramic images are generated, including: Image preprocessing includes distortion correction and color equalization, and geometric and color normalization of the original image. Feature matching is performed using the SIFT algorithm to extract feature points. Image registration is performed based on feature points, and the RANSAC algorithm is used to eliminate mismatches. The registered images are fused and optimized, and the seams are eliminated through multi-band fusion.

[0015] Furthermore, the command and dispatch engine adopts a rule-based workflow engine, supporting visual process orchestration; The system monitors for abnormal events on the video real-view map and automatically activates the corresponding emergency response plan when a preset abnormal event is detected.

[0016] Furthermore, a data processing module is used for video data stream processing and spatial data management; The video data stream processing adopts a pipelined architecture, sequentially performing data acquisition, decoding and transcoding, image enhancement, target detection, and data encapsulation. Spatial data management employs PostGIS extended storage for geospatial data and uses the R-Tree algorithm for spatial indexing. The data processing module uses WebRTC technology to achieve real-time audio and video communication, and the signaling service is based on the WebSocket protocol.

[0017] Secondly, a video real-scene map management and scheduling method based on AR technology is provided, which manages and commands the video real-scene map based on AR technology as described above.

[0018] This disclosure includes at least the following beneficial effects: This disclosure reduces the cost of AR technology applications by reusing existing cameras through software solutions, avoiding the need for specialized hardware procurement and reducing overall investment. It improves the adaptability of video tags, ensuring they automatically follow camera movement, enhancing management continuity and reliability. It addresses the issues of passive viewing and lack of interactivity in video surveillance systems by enabling dynamic annotation of business information and real-time interaction on video footage, thereby improving management efficiency.

[0019] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objects and other advantages of this disclosure may be realized and obtained by means of the structures pointed out in the description and the accompanying drawings. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the system architecture of an embodiment of this disclosure. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0023] like Figure 1 As shown, a video real-scene map management and scheduling system based on AR technology includes: a terminal presentation unit 101, a business capability unit 102, a technical support unit 103, and an infrastructure unit 104. The terminal presentation unit 101 is used to display a global video real-view map and interactive operations; Business capability unit 102 includes a video intelligent access module, an AR tag management engine, a real-scene map fusion engine, a command and dispatch engine, and a data processing module. The video intelligent access module acquires video streams, the AR tag management engine overlays AR tags with spatial location information onto the video screen, and the real-scene map fusion engine fuses and displays multi-source heterogeneous data to achieve command and dispatch based on real-scene maps. Technical support unit 103 is used to provide video streaming media processing, spatial location calculation and real-time data communication capabilities; Infrastructure unit 104 is used to provide computing, networking and storage resources.

[0024] The specific implementation details are as follows: The architecture adopts a microservices design and is divided into four layers: terminal presentation layer, business capability layer, technical support layer, and infrastructure layer.

[0025] The terminal presentation layer supports various terminal devices, including command center large screens, PC workstations, and mobile terminals. The large screen uses ultra-high-definition resolution (3840×2160 and above) to display a global video real-view map, supporting multi-touch and gesture operation; the PC provides a complete configuration management interface and implements 3D rendering based on WebGL technology; the mobile terminal supports iOS and Android systems and adapts to different screen sizes through responsive design.

[0026] The business capability layer comprises five core service modules: intelligent video access module, AR tag management engine, real-scene map fusion engine, command and dispatch engine, and data processing module. Each service communicates using a RESTful API interface, supporting distributed deployment and horizontal scaling.

[0027] The technical support layer provides fundamental technical capabilities, including a video streaming media processing framework, a spatial location computing engine, real-time data communication middleware, and security management components. Among these, the video streaming media processing framework supports tens of thousands of concurrent video streams with latency controlled within 200ms.

[0028] The infrastructure layer includes cloud computing resources, network equipment, and storage systems. The system supports a hybrid cloud deployment model, where critical business data can be deployed on a private cloud, while non-core business data can be deployed on a public cloud. The storage system adopts a tiered storage architecture, using SSDs for hot data and distributed object storage for cold data.

[0029] Video intelligent access module: This module employs protocol adaptive technology, supporting standard protocols such as ONVIF Profile S / T, GB / T 28181-2016, and RTSP 1.0, while also being compatible with proprietary protocols such as Hikvision iVMS-8700, Dahua DHSP, and Huawei eSight. The module has a built-in automatic protocol identification engine that can automatically select the optimal access method based on device fingerprint characteristics (such as device model, manufacturer information, and port characteristics).

[0030] The video stream processing employs adaptive bitrate technology, dynamically adjusting the video bitrate based on network bandwidth (adjustable from 256Kbps to 8Mbps). It supports multiple encoding formats such as H.265 / HEVC, H.264 / AVC, and MPEG-4, and improves processing efficiency through hardware acceleration technologies (such as GPU decoding). The module also features an intelligent reconnection mechanism, automatically restoring the connection in case of network anomalies to ensure video continuity.

[0031] AR tag management engine: The tag management system adopts a layered architecture, including a data layer, a logic layer, and a presentation layer. The data layer uses a time-series database to store the spatiotemporal information of tags, supporting millisecond-level timestamp accuracy; the logic layer implements tag lifecycle management, including full-process control of creation, update, and destruction; the presentation layer uses WebGL technology to render tags, supporting mixed display of 2D / 3D tags.

[0032] The label-following algorithm employs an improved visual SLAM technique: This algorithm belongs to the field of computer vision and intelligent information processing technology. Traditional label following algorithms suffer from problems such as positioning drift, tracking loss, and high computational complexity in dynamic environments and large-scale scenes. This algorithm achieves accurate and stable label recognition and tracking functions by improving visual SLAM technology.

[0033] The core innovation of this algorithm lies in combining an improved visual SLAM front-end, deep learning methods, and multi-sensor fusion technology to construct a complete label recognition and tracking system. The system achieves precise localization and environmental mapping through visual SLAM technology, while utilizing deep learning models for multi-category label recognition. Finally, it achieves stable label tracking and following through multimodal data fusion and optimization algorithms.

[0034] This algorithm incorporates a multi-Bernoulli filter to handle multi-target tracking scenarios and significantly reduces computational complexity through a distributed optimization fusion strategy. Compared to traditional methods, this algorithm offers significant improvements in tracking accuracy, real-time performance, and environmental adaptability, making it suitable for various applications such as robot vision navigation, intelligent video surveillance, and human-computer interaction.

[0035] Hybrid visual odometry based on feature points and direct methods: This algorithm employs a hybrid visual odometry design combining the feature point method and the direct method. This maintains the stability of the feature point method in motion estimation while leveraging the effectiveness of the direct method in weakly textured regions. Specifically, the system achieves preliminary pose estimation through ORB feature extraction and optical flow tracking, and then optimizes the estimation results using the direct method, significantly improving robustness in environments with uniform texture or varying illumination.

[0036] To address the critical issue of dynamic object interference in traditional SLAM systems, this algorithm introduces a dynamic feature point filtering mechanism. By analyzing the motion consistency of feature points and combining prior information provided by the semantic segmentation network, it effectively identifies and excludes feature points on dynamic objects, thereby reducing pose estimation errors and improving system stability.

[0037] Visual-inertial fusion SLAM architecture: To address the issue of tracking loss during rapid motion in pure vision-based SLAM, this algorithm employs a vision-inertial fusion framework, tightly coupling IMU (Inertial Measurement Unit) data with visual information. The IMU provides high-frequency motion prediction, assisting the vision system in maintaining tracking continuity even under blurred or transient occlusion conditions. Through sliding window optimization techniques, the system achieves accurate localization over long periods and over large areas with limited computational resources.

[0038] In the loop closure detection module, this algorithm not only employs the traditional bag-of-words model but also introduces a loop closure detection mechanism that combines semantic information with geometric verification. By analyzing the distribution and spatial relationships of semantic objects in the scene, the accuracy of loop closure detection in complex environments is significantly improved, effectively reducing the cumulative error of the SLAM system.

[0039] Tag recognition and tracking module design: This deep learning-based multi-class label recognition algorithm employs an improved YOLOv5 architecture to achieve real-time multi-class label recognition. Addressing the issue of traditional label recognition models' poor performance in detecting small targets, this algorithm introduces an adaptive spatial feature fusion mechanism into the feature pyramid network. By fusing feature maps of different scales through weight learning, it significantly improves the recognition accuracy of small-sized labels.

[0040] To improve the generalization ability of label recognition, this algorithm employs advanced techniques such as multi-scale training, mosaic data augmentation, and adversarial training during the training phase. Furthermore, for specific application scenarios, the system supports online learning, enabling it to dynamically adjust model parameters based on newly acquired sample data and continuously optimize recognition performance.

[0041] Label tracking and adaptive matching algorithm: In terms of label tracking, this algorithm designs a label tracker based on improved visual SLAM, treating labels as special map points in the SLAM system for tracking. By combining motion model prediction with appearance feature matching, stable label tracking is achieved even under brief occlusion or rapid movement.

[0042] To address the identity switching problem in multi-label scenarios, this algorithm introduces an adaptive matching mechanism. It comprehensively considers the visual features, motion trajectory, and spatial location information of the labels, and determines the optimal match through a weighted scoring function. Furthermore, the algorithm includes a dedicated trajectory management module to handle label creation, disappearance, and recurrence, ensuring the continuity of the tracking process. The label feature matching weight coefficients are shown in Table 1.

[0043] Table 1

[0044] Data fusion based on multi-Bernoulli filters: This algorithm employs a Labeled Multi-Bernoulli Filter (LMB) as the core framework for multi-sensor data fusion. This filter describes the target state using random finite set theory, effectively handling uncertainties and missed detections in sensor data. Compared to traditional filters, the LMB filter not only estimates the number and state of targets but also maintains the trajectory label for each target, making it highly suitable for multi-label tracking scenarios.

[0045] In a distributed sensor network environment, this algorithm introduces a Generalized Covariance Intersection (GCI) fusion rule to prevent the "double counting" problem during the fusion process. By recording fusion mapping information using group variables and sensor unique identifiers (IDs), historical information is utilized in subsequent fusion processes to reduce computational complexity from O(N^3) to O(N), significantly improving the system's real-time performance.

[0046] Global optimization based on pose graph: This algorithm employs a pose graph optimization method in the backend optimization, modeling the relationship between label positions, camera poses, and map points as a graph structure, and achieving global consistency optimization by minimizing the error function. For optimization problems in large-scale environments, this algorithm introduces a keyframe mechanism and sliding window optimization to control computational complexity while maintaining accuracy.

[0047] In particular, this algorithm proposes a label-aware optimization strategy that incorporates the semantic information of the labels as constraints into the optimization function. For example, for labels such as "No Entry," the system considers their semantic constraints during the optimization process, making the final trajectory planning more reasonable and safer. This semantic SLAM approach greatly improves the system's intelligence level in complex scenarios.

[0048] Reality map fusion engine: The map engine employs multi-source data fusion technology, supporting seamless integration of 2D electronic maps, 3D oblique photogrammetry models, point cloud data, and panoramic images. The 2D map is based on a tile pyramid structure and supports multi-level zoom; the 3D model uses LOD (Level of Detail) technology to dynamically adjust model accuracy based on viewpoint distance.

[0049] VR panorama generation is based on an improved SIFT feature matching algorithm, achieved through the following steps: ① Image preprocessing: distortion correction and color equalization; ② Feature matching: extracting feature points using the SIFT algorithm; ③ Image registration: eliminating mismatches using the RANSAC algorithm; ④ Fusion optimization: multi-band fusion to eliminate seams. The final generated panoramic image can reach a resolution of up to 8K, supporting 360° horizontal and 180° vertical panoramic browsing.

[0050] Command and dispatch engine: The scheduling engine employs a rule-based workflow engine, supporting visual process orchestration. The contingency plan management module provides a graphical configuration interface, allowing administrators to define response procedures via drag-and-drop. The video linkage module supports intelligent rule triggering, automatically activating contingency plans for abnormal events such as area intrusion and crowd gatherings.

[0051] The system supports multiple linkage modes: ① Sequential linkage: Executes multiple actions in chronological order; ② Conditional linkage: Triggers corresponding operations based on environmental parameters; ③ Manual linkage: Operators actively trigger the linkage process. Linkage delay is controlled within 1 second, meeting real-time command requirements.

[0052] Detailed explanation of the data processing module: Video data stream processing employs a pipelined architecture, including the following processing stages: ① Data Acquisition: Obtain the raw video stream through the SDK or protocol interface; ② Decoding and transcoding: Hardware decoding is performed using the FFmpeg framework, and the data is uniformly converted to YUV420 format; ③ Image enhancement: Adaptive histogram equalization is used to improve image quality; ④ Target detection: Real-time target detection is achieved based on the YOLOv5 algorithm; ⑤ Data encapsulation: Pack metadata and video stream together for transmission.

[0053] Spatial data management utilizes a PostGIS extension for the spatial database, supporting efficient storage and retrieval of geospatial data. The spatial index employs the R-Tree algorithm, enabling millisecond-level spatial queries. The coordinate system adopts the WGS84 coordinate system, supporting data exchange with other geographic information systems.

[0054] The system employs WebRTC technology for real-time audio and video communication, supporting both point-to-point and multi-party conferencing. Signaling services are based on the WebSocket protocol, ensuring low latency and high reliability. The data channel supports SRTP encrypted transmission, ensuring communication security.

[0055] System working mechanism introduction: AR tag dynamic following mechanism: Tag following is based on visual inertial odometry (VIO) technology, fusing visual information and IMU data to achieve accurate tracking. Specific implementation includes: ① Initialization phase: Establishing the world coordinate system and feature point map; ② Tracking stage: Feature points are tracked using the KLT optical flow method; ③ Optimization phase: Use graph optimization algorithms to optimize camera trajectory; ④ Relocation mechanism: When tracking fails, fast relocation is achieved through the DBoW2 algorithm.

[0056] Multi-map fusion mechanism: Map fusion employs an adaptive weighting algorithm, dynamically calculating fusion weights based on map accuracy, timeliness, and resolution. The fusion process includes: ① Coordinate unification: Convert different map data to a unified coordinate system; ② Data registration: Precise alignment is achieved through control point matching; ③ Conflict resolution: Data conflicts are handled using a priority strategy; ④ Real-time updates: Establish an incremental update mechanism to ensure data timeliness.

[0057] To enable those skilled in the art to better understand this disclosure, it is described below in conjunction with application scenarios: This algorithm can be widely applied in fields such as intelligent robot navigation, augmented reality, intelligent video surveillance, and human-computer interaction. In robot navigation, the algorithm enables robots to recognize and follow specific tags (such as people and equipment), making it suitable for service robots in environments such as hospitals and warehouses.

[0058] In the field of intelligent video surveillance, this algorithm can continuously track and analyze the behavior of multiple targets, and achieve abnormal behavior detection and early warning by combining scene semantic information. Experimental results show that in a 5-target tracking scenario, the OSPA error of this algorithm is reduced by 30% compared with traditional methods, and it can still maintain stable tracking performance in a 10-target scenario.

[0059] Extensive experimental verification demonstrates that this algorithm exhibits significant advantages in tracking accuracy, real-time performance, and robustness. Particularly in scenarios where the number of targets changes dynamically, the improved label-based multi-Bernoulli distributed optimization fusion method employed in this algorithm accurately estimates changes in the number of targets, avoiding false tracking and missed tracking. Algorithm performance comparison data is shown in Table 2.

[0060] Table 2

[0061] In summary, this algorithm, by improving visual SLAM technology and combining deep learning with multi-sensor fusion, achieves efficient and accurate label recognition and tracking, providing reliable technical support for the application of intelligent systems in complex environments. Typical application scenarios are implemented as follows: In smart park applications, the system implements the following key features: ①Personnel trajectory tracking: Achieve target tracking across cameras using ReID technology; ② Facility status monitoring: AR tags display equipment operating parameters in real time; ③ Intelligent inspection: Abnormal behavior recognition based on deep learning; ④ Emergency Command: Multiple contingency plans are executed in parallel, and resources are intelligently scheduled.

[0062] Highway management, addressing the specific needs of highway scenarios, the system implements: ① Event detection: Automatic accident detection based on spatiotemporal context analysis; ② Traffic parameter extraction: Real-time statistics of parameters such as traffic flow and average speed; ③ Route planning: Intelligent route planning based on real-time traffic conditions; ④ Information dissemination: Disseminate early warning information through the VMS system.

[0063] Performance metrics and reliability assurance: Under typical configuration (server: Intel Xeon Silver 4210 processor, 64GB memory; network: Gigabit Ethernet), the system achieves the following performance metrics: Video input capability: Supports simultaneous input of 1000 channels of 1080p video; Processing latency: End-to-end latency less than 500ms; Tag tracking accuracy: pixel-level error less than 3 pixels; System availability: 99.99% high availability design.

[0064] Reliability assurance measures include: ① Redundancy design: Key components adopt a primary-backup redundant architecture; ② Load balancing: Request distribution is achieved through Nginx; ③ Fault self-healing: Automatic fault recovery based on Kubernetes; ④ Data backup: Combining real-time incremental backup with regular full backup.

[0065] Through the detailed technical solution described above, those skilled in the art can fully understand the system's design concept, implementation methods, and key technical indicators. This system, through its innovative technical architecture and algorithm design, achieves an intelligent upgrade of the video surveillance system, demonstrating significant technological advancement and practicality.

[0066] This disclosure decouples hardware from existing cameras, avoids the purchase of specialized AR equipment, reduces hardware costs by 30%-50%, and shortens project cycles by 40% through software deployment. The tag-following technology, based on spatial perception algorithms, ensures information continuity, reduces human intervention, and improves system reliability, making it particularly suitable for rapidly changing scenarios such as emergency command. Through dynamic management and interactive querying of AR tags, video monitoring is upgraded from passive surveillance to an active management tool, improving management efficiency by over 50% and solving the problem of poor interactivity in traditional systems.

[0067] This disclosure presents a dynamic tracking method for AR tags based on video spatial perception. It uses computer vision algorithms to calculate camera motion parameters in real time, enabling tags to automatically follow the movement of the calibrated target and ensuring persistent visibility of information.

[0068] This disclosure adopts a multi-protocol video access and AR tag integration system, supporting standard protocols such as ONVIF and GB / T 28181 as well as proprietary protocols, realizing the reuse of existing equipment and the decoupling of AR function software and hardware.

[0069] This disclosure uses a video real-scene map fusion engine, combining 2D maps, AR real-scene and VR panoramas to provide multi-view interactive maps, supporting on-demand tag uploading and switching between high and low point perspectives.

[0070] This disclosure adopts video linkage and contingency plan management methods to realize picture-in-picture patrol, access to surrounding videos, and automatic triggering of linkage based on contingency plan rules, thereby improving command efficiency.

[0071] This disclosure employs VR panoramic map tagging technology, uses the krpano engine to stitch panoramic images, and supports tag search, positioning, and querying, replacing the need for high-point video.

[0072] This disclosure adopts a mobile AR tag interaction and information delivery mechanism, which allows users to directly operate video tags via mobile devices to achieve task allocation and real-time feedback, and supports offline operation.

[0073] This disclosure adopts a method of automatic tag generation and business data association, which automatically creates tags based on business rules and associates them with rich media information (such as text and video), reducing manual configuration.

[0074] This disclosure adopts a low-cost AR video processing architecture, which replaces professional AR hardware with software modules, reducing the total system cost and supporting rapid iteration.

[0075] This disclosure employs multi-terminal consistent display technology to ensure synchronized display of AR tags on large screens, PCs, and mobile devices, and supports multi-user shared sessions.

[0076] This disclosure employs a low-latency overlay method of real-time video stream and AR tags to optimize encoding and processing, reduce tag rendering latency, and improve user experience and response speed.

[0077] Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure.

Claims

1. An AR technology-based video real-scene map management and dispatching system, characterized in that, include: Terminal presentation unit, business capability unit, technical support unit, and infrastructure unit; The terminal presentation unit is used to display the global video real-view map and interactive operations; The business capability units include a video intelligent access module, an AR tag management engine, a real-scene map fusion engine, a command and dispatch engine, and a data processing module; The video stream is acquired through the video intelligent access module, AR tags with spatial location information are superimposed on the video screen using the AR tag management engine, and multi-source heterogeneous data is fused and displayed through the real-scene map fusion engine to realize command and dispatch based on the real-scene map. The technical support unit is used to provide video streaming media processing, spatial location calculation, and real-time data communication capabilities; Infrastructure units are used to provide computing, networking, and storage resources.

2. The video real-scene map management and scheduling system based on AR technology according to claim 1, characterized in that, The AR tag management engine uses visual SLAM technology to achieve dynamic tag following, specifically including: A hybrid visual odometry method combining feature point method and direct method is used for pose estimation, and a dynamic feature point filtering mechanism is introduced to eliminate interference from dynamic objects. A visual-inertial fusion framework is adopted to combine IMU data with visual information for localization; A loop closure detection mechanism combining semantic information and geometric verification is adopted to detect loop closures by visually recognizing the distribution of semantic objects and their spatial relationships in the scene.

3. The video real-scene map management and scheduling system based on AR technology according to claim 2, characterized in that, A hybrid visual odometry method combining feature point method and direct method is used for pose estimation, and a dynamic feature point filtering mechanism is introduced to eliminate interference from dynamic objects, including: For pose estimation, preliminary pose estimation is performed through ORB feature extraction and optical flow tracking, and then the estimation results are optimized using a direct method. For processing dynamic objects, the motion consistency of dynamic feature points is analyzed, and feature points with motion consistency are identified through a deep learning network; dynamic feature points identified as dynamic objects are removed, and only static objects are used for localization.

4. The video real-scene map management and scheduling system based on AR technology according to claim 2, characterized in that, Localization involves combining IMU data with visual information, including: When the image is blurry or obstructed, the position is predicted using an inertial measurement unit (IMU). When the image is clear and the inertial measurement unit (IMU) is drifting, visual position correction is used.

5. A video real-scene map management and scheduling system based on AR technology according to claim 2, characterized in that, The AR tag management engine also includes: A multi-class label recognition is achieved using an improved YOLOv5 architecture. An adaptive spatial feature fusion mechanism is introduced into the feature pyramid network, and feature maps of different scales are fused through weight learning. A labeled multi-Bernoulli filter (LMB) is used as the multi-sensor data fusion framework, and distributed optimization fusion is performed in combination with the generalized covariance intersection fusion rule (GCI). Pose graph optimization is adopted, which models the relationship between label position, camera pose and map points as a graph structure, and achieves global consistency optimization by minimizing the error function.

6. The video real-scene map management and scheduling system based on AR technology according to claim 1, characterized in that, The intelligent video access module supports protocol adaptive technology, automatically identifying device fingerprint characteristics and selecting the corresponding protocol for access; the video stream processing adopts adaptive bitrate technology, dynamically adjusting the video bitrate according to network bandwidth.

7. A video real-scene map management and scheduling system based on AR technology according to claim 1, characterized in that, The real-scene map fusion engine supports the seamless fusion of 2D electronic maps, 3D oblique photogrammetry models, point cloud data, and panoramic images; Two-dimensional electronic maps are based on a tile pyramid structure; The 3D oblique photogrammetry model uses Level of Detail (LOD) technology to dynamically adjust the model's accuracy based on the viewpoint distance. VR panoramic image generation is based on the SIFT feature matching algorithm. It generates panoramic images through image preprocessing, feature matching, image registration, and fusion optimization steps.

8. A video real-scene map management and scheduling system based on AR technology according to claim 7, characterized in that, VR panoramic image generation is based on the SIFT feature matching algorithm. Through image preprocessing, feature matching, image registration, and fusion optimization steps, it generates panoramic images for browsing, including: Image preprocessing includes distortion correction and color equalization, and geometric and color normalization of the original image. Feature matching is performed using the SIFT algorithm to extract feature points. Image registration is performed based on feature points, and the RANSAC algorithm is used to eliminate mismatches. The registered images are fused and optimized, and the seams are eliminated through multi-band fusion.

9. A video real-scene map management and scheduling system based on AR technology according to claim 1, characterized in that, The command and dispatch engine adopts a rule-based workflow engine and supports visual process orchestration. The system monitors for abnormal events on the video real-view map and automatically activates the corresponding emergency response plan when a preset abnormal event is detected.

10. A video real-scene map management and scheduling system based on AR technology according to claim 1, characterized in that, The data processing module is used for video data stream processing and spatial data management; The video data stream processing adopts a pipelined architecture, sequentially performing data acquisition, decoding and transcoding, image enhancement, target detection, and data encapsulation. Spatial data management employs PostGIS extended storage for geospatial data and uses the R-Tree algorithm for spatial indexing. The data processing module uses WebRTC technology to achieve real-time audio and video communication, and the signaling service is based on the WebSocket protocol.

11. A video real-scene map management and scheduling method based on AR technology, characterized in that, The video real-scene map management and scheduling system based on AR technology, as described in any one of claims 1-10, is used for management and command scheduling.