Multi-view real-time rendering and digital-view fusion method based on ultra-high-definition panoramic AR (Augmented Reality) video
Through 360-degree panoramic acquisition, multi-video stitching and fusion, and virtual data layer overlay using AR technology, the problems of insufficient real-time, accuracy, and flexibility in ultra-high-definition panoramic video generation have been solved, efficient multi-perspective collaboration and interactive management have been achieved, and user experience and data accuracy have been improved.
Patent Information
- Application Number
- CN202510935194.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-17
AI Technical Summary
The existing ultra-high-definition panoramic video generation process has problems with real-time performance, accuracy, and flexibility, especially in multi-perspective rendering, lighting changes, and poor environmental adaptability, resulting in insufficient user interactivity and inability to achieve immersion.
Ultra-high-definition panoramic videos are generated using 360-degree panoramic acquisition, multi-video stitching and fusion, feature point detection and matching, and video compression encoding. AR technology is used to overlay virtual data layers in a virtual environment. Combined with multi-dimensional data virtual-reality fusion technology, a deep integration of reality and virtuality is achieved. The spatial segmentation cone clipping algorithm is used to eliminate invalid rendering, and thread separation and asynchronous mechanisms are used to optimize the rendering process.
It achieves efficient multi-perspective seamless collaboration and interactive management, improves real-time performance and precision, ensures data accuracy and interactive fluency, provides rich visual experience and interactive management functions, and is suitable for smart cities, industrial simulation, virtual tourism and other fields.
Smart Images

Figure CN120812232A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of AR video rendering and digital-visual fusion, and particularly relates to a multi-view real-time rendering and digital-visual fusion method based on super-high-definition panoramic AR video. BACKGROUND
[0002] At present, most safety management systems are based on GIS (geographic information system) maps, and the command and dispatch functions are realized by superimposing relevant elements on the maps, such as the public security video monitoring command and management system developed for the needs of social public safety management. However, the command and dispatch mode based on the plane GIS map still has some deficiencies in use and experience, such as the traditional GIS map command mode based on the BS / CS (browser / client server) architecture, which has two resource display modes: one is the resource tree mode, which classifies resources by placing different resources under different organizations through multi-level organization; the other resource display mode is to add resources directly to the map, and find the corresponding geographical resources through the geographical information of the map; the effects brought by these two resource display modes are not very good when the resource distribution is dense. Therefore, it is particularly important to use digital-visual fusion technology to clearly display the safety management elements contained in the area through real-time monitoring pictures based on real-time video monitoring of key areas, so that the presentation of resources is more intuitive.
[0003] In addition, GIS cannot achieve active perception, but can only rely on other systems to discover and then present through GIS. This passive perception method may delay the command opportunity when encountering an emergency. Therefore, it is an important research direction to improve safety perception by realizing real-time monitoring of key areas, realizing resource joint linkage, and being able to discover and view resources in time, combining the data associated with resources with business. Secondly, the traditional GIS map command mode also has the problem of weak overall control of the region, and multiple pictures are relatively independent and incoherent.
[0004] The current industrial safety production related system only displays the site information through a two-dimensional plane map, does not combine with algorithm alarm results, and only informs the alarm result related information in the form of text, pictures, etc., so that the relevant operation and maintenance personnel cannot quickly perceive and locate the place where the alarm occurs, so as to be unable to respond efficiently. Therefore, there is an urgent need for an AR (augmented reality) rendering technology that combines video and system map to improve dynamic perception by combining real-time video with algorithm alarm results, but there are problems such as insufficient real-time performance, precision and flexibility in the existing super-high-definition panoramic video generation process.
[0005] (1) Insufficient real-time performance.
[0006] Processing latency: The generation of ultra-high-definition panoramic videos requires processing a large amount of data, especially at high frame rates (such as 60fps or higher), which puts a high demand on computing resources. Real-time rendering involves the synthesis and rendering of multiple perspectives, and any computational latency can cause the final picture to lag, affecting user experience; for example, in AR applications, users' actions need immediate feedback, but latency can cause the picture to not match the user's actual position.
[0007] Multi-perspective rendering: Multi-perspective rendering techniques require simultaneous processing of data from multiple cameras, which increases the complexity of real-time processing. Although some optimization algorithms have been proposed, it is still difficult to achieve ideal real-time performance in extreme cases such as rapid motion or complex scenes.
[0008] (2) Lack of precision.
[0009] Parallax and lighting issues: During multi-perspective video synthesis, parallax between different perspectives can cause unnatural effects in the synthesized picture, and changes in lighting conditions can also affect the final visual effect. Current algorithms still have deficiencies in handling lighting changes and parallax, resulting in a decrease in the realism and precision of the synthesized video.
[0010] Image quality loss: During video compression and transmission, especially in low-bandwidth environments, image quality may be lost, causing picture blurring or compression artifacts, which is particularly detrimental to ultra-high-definition videos that rely on the clarity and precision of details.
[0011] Poor environmental adaptability: In different environmental conditions (such as changes in light, weather effects, etc.), the generation accuracy of ultra-high-definition panoramic videos may be significantly reduced; for example, changes in strong light or shadows during video acquisition can cause inconsistencies in the image quality captured by the sensor, affecting the final synthesis effect.
[0012] (3) Lack of flexibility.
[0013] Poor scene adaptability: Existing panoramic video generation techniques often rely on pre-set environmental conditions and configurations, lacking the ability to adapt to dynamic environments; for example, unexpected situations in mobile shooting (such as crowds, smoke, etc.) can cause the system to fail to adapt in real time, affecting the stability and continuity of the picture.
[0014] Lack of interactivity: In augmented reality applications, users expect to interact with the content, but existing technologies have limited capabilities in dynamic content generation and real-time feedback, resulting in insufficient flexibility for users when interacting and failing to achieve the expected immersion. SUMMARY
[0015] In view of the above, the application provides a multi-view real-time rendering and digital-video fusion method based on super-high-definition panoramic AR video, which adopts an AR rendering technology of fusing video and a system map to improve dynamic perception by combining real-time video and algorithm alarm results, and solves the problems of real-time performance, precision and flexibility in the existing super-high-definition panoramic video generation process.
[0016] A multi-view real-time rendering and digital-video fusion method based on super-high-definition panoramic AR video comprises the following steps: (1) generating super-high-definition panoramic video by 360-degree panoramic acquisition, multi-video splicing and fusion, feature point detection and matching, and video compression and encoding; (2) creating a virtual environment by using an AR content generation technology, virtually digitizing a subject object by AR technology to enable flexible operation and in-depth analysis of the subject object in the virtual environment, then layering virtual data on the super-high-definition panoramic video and enhancing the fusion of video and application by multi-dimensional data virtual-real fusion technology to realize the deep combination of reality and virtuality; (3) realizing seamless collaboration and interactive management from a high point and a large scene to local details and from outdoor panorama to indoor view by multi-view collaborative technology; (4) for the rendering process of a large-scale panoramic virtual scene, removing objects that do not need to be rendered by a view frustum clipping algorithm based on spatial segmentation to reduce invalid rendering and realize lightweight loading.
[0017] Further, the specific implementation of step (1) is as follows: 1.1 according to the scene size, complexity and resolution requirement, realizing 360-degree panoramic coverage by a multi-camera combination array to collect video data in real time; during the collection process, the positions and angles of the cameras need to be accurately configured to ensure that the view coverage ranges of the cameras seamlessly connect, and the data collection of the cameras must be synchronized to avoid time deviation in the subsequent splicing process; 1.2 pre-processing the video image by grayscale and filtering to reduce noise caused by different light intensities and parameters between different cameras, then analyzing the background and structure of the video image based on a deep learning algorithm to select key frames with less dynamic objects and clear geometric structures from the video sequence to provide clearer scene data for subsequent modeling and splicing; 1.3 Build 3D field expression system, according to the camera parameter calculation image pixel corresponding 3D scene coordinates, so as to register the video image to 3D scene, and then through SURF(Speed Up Robust Feature, accelerated robust feature) and FLANN(Fast Library for Approximate Nearest Neighbors, fast approximate nearest neighbor library) algorithm to detect and match the feature points of the image, based on the matched feature points, the video images collected by multiple cameras are seamlessly spliced in the same plane through the weighted average image fusion algorithm; 1.4 The video encoding standard H.265 or HEVC(high efficiency video coding) is used to compress the spliced super high definition panoramic video, reduce the file size, and improve the transmission efficiency; according to the scene change, the encoding parameters are dynamically adjusted, and under the premise of ensuring the video quality, the efficient compression of the video is realized.
[0018] Further, the specific implementation of step (2) is as follows: 2.1 A series of AR technologies including the use of Fabric technology framework(based on HTML5 Canvas open source graphics operation framework) to build visual drawing board, integration of MP4Box(multimedia packaging tool) video player plug-in, dependence on ZLMediaKit(open source streaming media live server) for video streaming, use of three.js(Javascript library for creating and rendering 3D graphics in browser) three-dimensional graphics engine, and application of spherical coordinate space conversion algorithm are used to virtually digitize the subject object, so that it can be flexibly operated and deeply analyzed in the virtual environment, and then the generated virtual data is layered on the super high definition panoramic video to realize the effect of virtual and real fusion; 2.2 Build a virtual information layer, draw virtual data and apply point, line and face marking to the collected multimedia materials including image, text, video and audio, and use Fabric technology framework for visual presentation; then, the real-time video stream from the camera array is synchronized and fused with these virtual data to construct an AR scene; the interactive actions of the subject object in the initially constructed AR scene are captured, which will trigger scene update, and then based on the updated data, the conversion of coordinate system is performed, from global coordinate system to local coordinate system, and then back to the updated spatial coordinates, to ensure that the AR scene can respond and update in real time; 2.3 Deep development of label plug-in using ARRM (Automatic Resource and Label Management) component to realize creation, storage, display, data integration and webpage display of labels, so that the subject object can flexibly design label content and customize the appearance of the label including shape, color, size and spatial position adjustment according to actual scene and business needs, further promoting the deep integration and linkage of video and various applications; 2.4 Real-time analysis of video content, precise marking of abnormal events, behaviors and objects using AR marking technology, precise binding and dynamic following effect of labels and objects, and then optimizing the classification, arrangement and management process of video content through systematic label and metadata management mechanism; 2.5 Generate a fusion image to realize the deep combination of reality and virtuality by calculating and integrating geographic location information.
[0019] Further, the step 2.2 calculates the global spherical coordinates according to the field of view of the camera ball machine, the inclination of the holder and the horizontal rotation angle during the construction of the virtual information layer, and then maps the coordinates on the screen to the panoramic video space coordinates to realize the precise positioning of the virtual elements and preliminarily form the AR scene. In addition, by integrating the geographic location information of the camera group, the enhanced AR scene is seamlessly integrated into the panoramic map to generate a fusion image, realizing the deep integration of video and virtual content. At the same time, the orientation and visual range of the camera group are accurately calculated using the latitude and longitude and the north point information of the map, and the corresponding virtual data is intelligently matched based on the visual range, and these matched virtual data are fed back to the AR scene for display, thereby providing the subject object with a more rich and interactive virtual and real combination experience.
[0020] Further, the specific implementation of step 2.5 is as follows: First, the holder pitch angle and field of view angle of the current camera ball machine need to be obtained, and then the loop stage is entered. In the preliminary AR scene, the scene change data triggered by the subject object interaction is collected, and then the coordinate system conversion operation is performed using these data to establish a localized coordinate reference system. Next, the local coordinate system is seamlessly connected to the global coordinate system framework to calculate the updated spatial position coordinates. According to these latest coordinate data, the AR scene is drawn and updated in real time to ensure that the instant feedback of the subject object interaction and the dynamic change of the scene are consistent. Finally, in the display stage, the geographic coordinate information of the camera group is collected and integrated, and the AR scene is seamlessly integrated into the system map using this information to finally generate a fusion image containing both real elements and virtual elements.
[0021] Further, the specific implementation of step (3) is as follows: 3.1 Create interactive tags that can be updated and reflected in real-time in high point videos by tag linkage and interactive presentation technology. These tags need to be dynamically bound with low point video sources, face recognition systems, and vehicle camera data to achieve unified management and display of multi-dimensional data. 3.2 To achieve picture-in-picture linkage of high and low level videos, use timestamp alignment and resampling technology based on buffer to achieve synchronization, superposition and rendering of high and low point videos. Through CDN (Content Delivery Network) and edge computing, reduce the delay of data transmission, and use adaptive streaming technology to dynamically adjust the playback speed of video stream to match the data update frequency according to network conditions. 3.3 Use 3D zoom and tag synchronization to achieve high-high linkage (such as outdoor high point one-key switching to indoor high point), high-low linkage (such as tag point warning linkage high-altitude ball machine focusing), and panoramic detail linkage. Use three.js three-dimensional graphics engine to ensure smooth rendering of images. Through high and low perspective switching technology and preloading target video, maintain the coherence of the video when switching perspectives, minimize the delay when switching, and use the easing algorithm to avoid visual discomfort caused by delay or inconsistency during the switching process. 3.4 Build a collaborative management system of three-dimensional large scenes and local details to achieve full cooperation between high and low perspectives and between internal and external scenes. Specifically, based on GIS map information, track and synchronize the location information of mobile video equipment in real time, use GPS positioning technology to ensure the information authenticity and dynamic of panoramic map, update the location of equipment in real time, support the display of real-time data, and realize seamless cooperation of multiple perspectives, providing users with rich visual experience and interactive management functions.
[0022] Further, the implementation process of step 3.1 needs to rely on precise timestamp and GIS spatial position calibration technology to ensure the synchronization and consistency of different data sources. Specifically, use timestamp synchronization, i.e. all data sources must have a unified time reference, use NTP (Network Time Protocol) to ensure the time consistency between different data sources, and each frame of video and each piece of data is attached with a timestamp for accurate matching. Also use GIS spatial position calibration method, use GIS technology for position calibration, combine the position of video source with the coordinates in geographic information system, and use Kalman filtering algorithm to correct the position data in real time to ensure the accurate correspondence between video stream and geographic information.
[0023] Further, the specific implementation mode of the step (4) is as follows: firstly, the whole scene is divided into several spatial partitions by using a spatial partitioning technique, and then objects in each partition are processed; by calculating the intersection relationship between the bounding box of each object and the view cone, objects completely located outside the view cone are quickly removed to reduce unnecessary calculation burden; then, the objects to be rendered are converted to the view cone coordinate system for collision test, the bounding box of the object in the world coordinate system is compared with the view cone, if all the vertices are found to be located outside the view cone, the object is completely clipped; if there is an intersection, the projection points of the object vertices on the near clipping plane and the far clipping plane are compared to determine whether depth clipping is needed; after clipping, the rendering result is converted back to the world coordinate system for subsequent rendering and processing; the view cone clipping algorithm introduces the design concept of thread separation, and combines the asynchronous mechanism for event loop processing, thereby avoiding the problem of main process blocking, improving the rendering frame rate, and ensuring the smooth experience of users in the interaction process.
[0024] A computer device comprises a memory and a processor, the memory has a computer program stored therein, and the processor is configured to execute the computer program to implement the multi-view real-time rendering and data-visual fusion method based on the ultra-high-definition panoramic AR video.
[0025] A computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the multi-view real-time rendering and data-visual fusion method based on the ultra-high-definition panoramic AR video.
[0026] Based on the above technical solutions, the present application has the following beneficial technical effects: 1. Comprehensive coordination to achieve efficient control of the scene: the present application constructs a three-dimensional scene and local detail coordination management system, relies on GIS map information and GPS positioning technology, realizes comprehensive coordination of high and low perspectives, internal and external scenes, can track the position of the mobile video device in real time, dynamically update the panoramic map, and ensure seamless switching of multiple perspectives. Users can quickly locate and view key scene information in a rich visual experience through powerful interactive management functions, and can efficiently complete both macro scene control and micro detail observation.
[0027] 2. Two-dimensional calibration to ensure accurate and reliable data: the present application uses precise timestamp and GIS spatial position calibration technology to unify the time reference of different data sources with NTP, and adds timestamps to videos and data to achieve accurate matching in the time dimension; at the same time, GIS technology is used in combination with Kalman filtering algorithm to calibrate and correct the video source position in real time, ensuring accurate correspondence in the geographical space dimension. The two-dimensional calibration mechanism effectively avoids the confusion of data due to time and space deviation, and provides a solid data foundation for subsequent analysis and decision-making based on real-time data.
[0028] 3. Intelligent Rendering Optimization for Improved Interaction Fluency: This invention uses spatial segmentation and a frustum clipping algorithm to rapidly remove unnecessary objects outside the frustum, significantly reducing computational complexity. It also employs thread separation combined with an asynchronous rendering mechanism to avoid main process blockage and significantly improve rendering frame rates. During ultra-high-definition panoramic AR video interactions, users experience smooth, lag-free visuals. For example, when navigating virtual scenes or viewing complex models, the screen responds quickly, ensuring a seamless and natural interactive experience.
[0029] 4. Deep integration of software and hardware promotes widespread application of the technology: The computer device and computer-readable storage medium designed in conjunction with this invention effectively implement multi-perspective real-time rendering and digital-visual fusion methods. The computer device's processor executes programs stored in memory, ensuring stable operation of the technical solution; the computer-readable storage medium ensures reliable storage and flexible access to programs. This deep integration of software and hardware ensures excellent compatibility and scalability of the invention, making it widely applicable in fields such as smart cities, industrial simulation, and virtual tourism. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a technical architecture diagram of multi-perspective real-time rendering and digital-visual fusion based on ultra-high-definition panoramic AR video in the present invention. DETAILED DESCRIPTION
[0031] In order to describe the present invention more specifically, the technical solution of the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0032] like Figure 1 As shown in FIG, the present invention is based on a multi-view real-time rendering and digital-visual fusion method of ultra-high-definition panoramic AR video, comprising the following steps: Step 1: Efficient generation process of ultra-high-definition panoramic videos.
[0033] This study aims to solve the problems of insufficient real-time performance, accuracy and flexibility in the existing ultra-high-definition panoramic video generation process, and to design and implement an innovative method for rapid generation of ultra-high-definition panoramic videos by splicing and fusion of multiple videos.
[0034] 1.1 Efficient Video Capture: Based on the scene's size, complexity, and resolution, a multi-camera array achieves 360-degree panoramic coverage and captures real-time video data. The acquisition process requires precise configuration of camera positions and angles to ensure seamless coverage. Furthermore, camera data acquisition must be synchronized to avoid time lag during the subsequent stitching process.
[0035] 1.2 Panoramic video stitching and rendering: An innovative panoramic video stitching architecture is designed, which can efficiently integrate multiple video sources, generate comprehensive images with high spatiotemporal consistency, and provide seamless interaction experience with 3D scenes. First, the video images are preprocessed by methods such as grayscale processing and filtering to reduce noise such as different lighting levels, different parameters, etc. between different devices; then, based on deep learning algorithms, the video images are analyzed for background and structure, and key frames with less dynamic objects and clear geometric structures are selected from the video sequence to provide clearer scene data for subsequent modeling and stitching. Next, a 3D scene expression system is designed, based on camera parameters and other information, the 3D world coordinates corresponding to the image pixels are calculated, and the image is registered to the 3D world for subsequent feature point matching; feature point detection and matching are performed on the image through SURF, FLANN, and deep learning algorithms; finally, based on the matched feature points, multiple video images are seamlessly stitched on the same plane through a weighted average image fusion algorithm.
[0036] 1.3 Compression and encoding optimization: The synthesized ultra-high-definition panoramic video is optimized through video compression and encoding technology to improve transmission and real-time rendering efficiency. The invention uses efficient video encoding standards such as H.265 / HEVC to compress the synthesized ultra-high-definition panoramic video, reducing file size and improving transmission efficiency; dynamically adjusts encoding parameters according to scene changes to achieve more efficient compression while ensuring video quality. Use of caching technology to improve video playback smoothness and reduce stuttering caused by network delay; use of GPU for video decoding and rendering to improve real-time rendering efficiency and ensure smooth user experience.
[0037] Step 2: AR content generation and multi-dimensional data virtual-real fusion.
[0038] 2.1 Technical support: The core of AR real scene rendering and multi-dimensional data fusion is built on a variety of technologies and methods, including but not limited to using the Fabric framework to build a visual palette, integrating mp4box video playback plug-in solutions, relying on zlmediakit for video streaming, using three.js three-dimensional graphics engine, and using complex spherical coordinate space conversion algorithms (including gimbal control, field of view angle, and vector calculation), with this comprehensive AR technology, the characteristics of the subject object can be accurately digitized and flexibly operated and analyzed in a virtual environment, and by cleverly layering these virtual data on top of the real scene video, the effect of virtual-real fusion is achieved.
[0039] 2.2 AR real-time scene synchronization and display based on virtual data rendering: First, a virtual information layer is constructed by rendering virtual data and applying point, line, and surface marking methods to collected multimedia materials such as images, text, videos, and audio, and visualizing the results using Fabric technology. Then, real-time video streams from the camera array are synchronized and fused with the virtual data to create an AR environment. During this process, the field of view of the spherical camera (ball machine) in the camera, the tilt and horizontal rotation angle of the gimbal, and the precise global spherical coordinate calculation are used to map the coordinates on the screen to the panoramic video space coordinates, enabling accurate positioning of virtual elements and the initial formation of an AR scene. Subsequently, the system captures the subject's interactive actions in the initially constructed AR environment, which can trigger scene updates. Based on these updated data, the coordinate system is converted from the global coordinate system to the local coordinate system and then back to the updated spatial coordinates, ensuring that the virtual scene can respond and update in real time. In addition, the geographical location information of the camera group is integrated to seamlessly fuse the enhanced virtual scene onto the panoramic map, generating a fused image and achieving deep integration of video and virtual content. To further enhance the subject's experience, the system uses the latitude and longitude and north point information of the map to accurately calculate the orientation and visual range of the camera group. Based on this visual range, the system intelligently matches the corresponding virtual data and immediately feeds back the matched data to the AR scene for display, providing the subject with a more rich and interactive virtual and real-world experience.
[0040] 2.3 Label management: To meet the diverse needs of the subject, the present application allows for individual customization of the appearance of the label, including shape, color, size, and spatial position (distance). In addition, it also provides a free adjustment option for the size of the label icon, and the customization of label content is flexible and diverse, divided into two modes: basic customization and advanced customization. In the basic customization mode, the subject can easily configure the label through a friendly Web interface, covering various elements such as text, pictures, hyperlinks, and videos. In the advanced customization mode, the powerful functions of the ARRM component are utilized to support the subject in developing label plugins in depth, achieving comprehensive functions such as label creation, storage, display, data integration, and web display. The subject can flexibly design the label content according to the actual scene and business needs, and closely associate it with relevant information. By operating these labels, the subject can efficiently display the required information, further promoting the deep integration and linkage of video and various applications, and achieving the maximization of information utilization and the improvement of comprehensive benefits.
[0041] 2.4 Coordinate calculation: Traditional video technology is limited to simple superimposition or information labeling of page content, and cannot flexibly respond to specific labeling needs of PTZ cameras. In contrast, AR tagging technology, with its deep integration with digital twin scenarios, achieves precise binding and dynamic following effects of labels and objects. This innovation has great potential in the field of video surveillance and security. The technology can analyze video content in real time, accurately mark abnormal events, behaviors, and objects, thereby giving the monitoring system the ability of intelligent identification and early warning. Through systematic label and metadata management mechanisms, not only the classification, organization, and management processes of video content are optimized, but also the efficiency is improved, and revolutionary changes are brought to the management and maintenance of large-scale video databases. In addition, in-depth video tagging and metadata analysis can mine valuable information and deep insights from videos, providing strong support for decision-making.
[0042] 2.5 In order to realize the acquisition of spatial coordinates, first of all, the current PTZ camera's pan tilt angle and the current PTZ camera's field of view angle need to be obtained. The current PTZ camera's pan tilt angle is (radian): ptz The current PTZ camera's field of view angle is
[0043] (field of view radian):
[0044] where the local position Local _ pos is:
[0045] The calculation process of position rotation is as follows: z-axis rotation: local_pos.applyAxisAngle(new THREE.Vector3(0, 0, 1), current_ptz_r.z) y-axis rotation: local_pos.applyAxisAngle(new THREE.Vector3(0, 0, 1), current_ptz_r.y) where the local coordinate is:
[0046] For the plane coordinate:
[0047] where: height is the video height, widthFor video width, Pixel coord For plane coordinates.
[0048] Then enters the loop phase, in the preliminary AR environment, by collecting the subject interaction triggered scene change data, and then using these data to perform coordinate system conversion operation, to establish a localized coordinate reference system. Next, this local coordinate system is seamlessly connected to the global coordinate system framework, so as to calculate the updated spatial position coordinates. Finally, according to these latest scene update data, real-time rendering and updating in the virtual environment, to ensure the immediate feedback of the subject interaction and the dynamic change of the scene consistent. Finally, in the display stage, collect and integrate the precise geographic coordinate information of the camera group, and then use these information to seamlessly integrate the carefully constructed virtual scene into the system map. This process realizes the deep integration of video content and virtual scene, and finally generates a fusion image that contains both real elements and virtual elements.
[0049] Step 3: Multi-perspective coordination and interaction management.
[0050] Multi-perspective coordination technology aims to achieve seamless coordination and interaction management from high point large scene to local details, from outdoor panorama to indoor view through highly integrated technology system. This technology system integrates multiple video sources and data streams to ensure that users can achieve real-time and dynamic information management and display in various environments.
[0051] 3.1 First of all, the creation and application of interactive tags, through the tag linkage and interactive presentation technology, create interactive tags in high point video that can update and reflect various data in real time. The characteristics of interactive tags include: ① Dynamic binding: the tag is not just static text, but can be dynamically bound to multiple data sources (such as low point video source, face recognition system, vehicle card data, etc.); through the use of API and data stream interface, the tag can real-time access and update information.
[0052] ② Data update mechanism: the system will periodically pull data from various data sources, and through real-time communication protocols such as WebSocket, update the data on the tag in real time; each tag can display information related to it, such as face recognition results, device running information, etc.
[0053] Ensuring data synchronization and consistency, realizing unified management and display of multi-dimensional data relies on precise timestamp and GIS spatial position calibration technology. The invention adopts timestamp synchronization, all data sources must have a unified time reference, uses high-precision time synchronization protocol (such as NTP) to ensure the time consistency between different data sources; each frame of video and each piece of data is attached with timestamp for accurate matching. The invention also adopts GIS spatial position calibration method, uses GIS technology for position calibration, combines the position of video source with the coordinates in geographic information system, and corrects the position data in real time through algorithm (such as Kalman filter), to ensure the accurate correspondence between video stream and geographic information.
[0054] 3.2 Video synchronization algorithm and network delay optimization, in order to realize the picture-in-picture linkage of high and low level videos, develop video synchronization algorithm and network delay optimization technology, the invention develops efficient synchronization algorithm, including timestamp alignment and resampling technology based on buffer, to ensure that high and low videos can be synchronized, superimposed and rendered. The algorithm should consider network delay, can automatically adjust the video playback speed to match the data update frequency, reduce the delay of data transmission by using CDN and edge computing. At the same time, use adaptive streaming technology to dynamically adjust the quality of video stream according to network conditions, to ensure smooth viewing experience.
[0055] 3.3 Panoramic detail linkage and perspective switching optimization level, in order to realize panoramic detail linkage and effective switching of high and low perspectives, the system will use 3D zoom and label synchronization, as well as high-high linkage and high-low linkage. When viewing high point video, users can choose to zoom in on a specific area in 3D, while synchronously displaying relevant label information of the area. Use three.js and other three-dimensional graphics engines to ensure smooth rendering of images. When switching, users can switch from outdoor high point video to indoor high point video with one key. The system will preload target video to minimize delay during switching, and support label point pre-warning linkage, use high-altitude ball machine to focus on specific position, and realize efficient alarm.
[0056] Research on switching technology of high and low perspectives to ensure perspective optimization, maintain the coherence of video when switching perspectives, use Easing algorithm to avoid visual discomfort caused by delay or inconsistency during switching process.
[0057] 3.4. Construct a collaborative management system for building a large scene and local details, to ensure all-round coordination between high and low perspectives, and between internal and external scenes. Specific measures include: real-time tracking and synchronization of data based on GIS map information to track and synchronize the location information of mobile video equipment in real time; through GPS and other positioning technologies, ensure the information authenticity and dynamic of panoramic map, real-time update the location of equipment, support real-time data display. Ultimately, through the integration of the above technologies, seamless coordination of multiple perspectives is achieved, providing users with richer visual experience and interactive management functions; users can freely switch perspectives according to their needs to obtain the best viewing experience, while dynamic labels and multi-dimensional data display improve information availability and intuitiveness.
[0058] Step 4: Lightweight loading and updating of large-scale scenes.
[0059] To solve the performance problems such as frame freezing that occur during panoramic model rendering, the present application proposes a space partition-based view frustum clipping algorithm. This algorithm effectively eliminates objects that do not need to be rendered by determining whether the object is inside the view frustum, significantly improving rendering efficiency. Specifically: first, use space partitioning technology to divide the entire scene into several spatial partitions, and then process the objects within each partition; next, by calculating the intersection relationship between the bounding box of each object and the view frustum, quickly eliminate those objects that are completely outside the view frustum to reduce unnecessary computational burden.
[0060] However, when dealing with large panoramic scenes, the large amount of data may cause objects to be cut and split, so further optimization of the video loading and updating process is needed. In the specific implementation of the view frustum clipping, the present application can effectively clip the polygons by accurately detecting and calculating their depths, thereby reducing invalid rendering and further improving performance. Specifically: First, convert the objects to be rendered to the view frustum coordinate system for subsequent processing; then, perform collision testing by comparing the bounding box of the rendering model in the world coordinate system with the view frustum. If all vertices are found to be outside the view frustum, the object is completely clipped; if there is an intersection, perform depth clipping. Specifically, by comparing the projection points of the object's vertices on the near and far clipping planes, determine whether depth clipping is needed.
[0061] Finally, after clipping, convert the rendering result back to the world coordinate system for subsequent rendering and processing. To further optimize computational overhead, the algorithm introduces a thread separation design concept combined with an asynchronous mechanism for event loop processing, effectively avoiding the blocking problem of the main process and significantly improving the rendering frame rate, ensuring a smooth experience for users during interaction.
[0062] The above description of the embodiments is for the purpose of enabling one of ordinary skill in the art to make and use the application and is not intended to limit the application as construed in the broadest scope possible. Inasmuch as modifications to the above described embodiments can readily be made by persons of ordinary skill in the art, it is intended that the application not be limited to the embodiments described above but should be construed in the broadest scope possible.
Claims
1. A multi-view real-time rendering and digital-visual fusion method based on ultra-high-definition panoramic AR video, characterized in that: The steps include: (1) Generate ultra-high-definition panoramic videos through 360-degree panoramic acquisition, multi-video splicing and fusion, feature point detection and matching, and video compression encoding; (2) Using AR content generation technology to create a virtual environment, using AR technology to virtually digitize the main objects so that they can be flexibly operated and deeply analyzed in the virtual environment, and then overlaying the virtual data on the ultra-high-definition panoramic video and enhancing the integration of video and application through multi-dimensional data virtual-real fusion technology; (3) Seamless collaboration and interactive management from high-angle scenes to local details, from outdoor panoramas to indoor perspectives, through multi-perspective collaborative technology; (4) For the rendering process of large-scale panoramic virtual scenes, objects that do not need to be rendered are eliminated through a cone clipping algorithm based on space segmentation.
2. The multi-perspective real-time rendering and digital-visual fusion method based on ultra-high-definition panoramic AR video according to claim 1 is characterized by: The specific implementation of step (1) is as follows: 1.1 Based on the scene size, complexity, and resolution requirements, a multi-camera array is used to achieve 360-degree panoramic coverage and collect video data in real time. During the collection process, the camera positions and angles must be precisely configured to ensure seamless coverage of each camera's viewing angle and synchronized data collection. 1.2 Preprocess the video images through grayscale conversion and filtering to reduce noise caused by varying lighting levels and parameters between different cameras. Then, use a deep learning algorithm to analyze the background and structure of the video images, selecting key frames from the video sequence with fewer dynamic objects and clear geometric structures. 1.3 Build a 3D scene representation system. Calculate the 3D scene coordinates corresponding to image pixels based on camera parameters, thereby registering video images into the 3D scene. Then, use the SURF and FLANN algorithms to detect and match feature points in the image. Based on the matched feature points, use a weighted average image fusion algorithm to seamlessly stitch video images captured by multiple cameras onto the same plane. 1.4 Use the video coding standard H.265 or HEVC to compress the stitched ultra-high-definition panoramic video to reduce file size and improve transmission efficiency; dynamically adjust the coding parameters according to scene changes.
3. The multi-perspective real-time rendering and digital-visual fusion method based on ultra-high-definition panoramic AR video according to claim 1 is characterized by: The specific implementation of step (2) is as follows: 2.1 By using a series of AR technologies, including the Fabric technology framework to build a visual sketchpad, integrating the MP4Box video playback plug-in, relying on ZLMediaKit for video streaming, adopting the three.js 3D graphics engine, and applying the spherical coordinate space conversion algorithm, the main object is virtually digitized, allowing it to be flexibly operated and deeply analyzed in the virtual environment. The generated virtual data is then layered on top of the ultra-high-definition panoramic video; 2.2 Constructing a virtual information layer: By mapping virtual data and applying point, line, and surface annotation to collected multimedia materials, including images, text, video, and audio, the Fabric technology framework is used for visualization. The real-time video stream from the camera array is then synchronously integrated with this virtual data to construct an AR scene. Capturing the subject's interactive actions in the preliminarily constructed AR scene. These actions will trigger scene updates, and then perform coordinate system transformations based on the updated data, transitioning from the global coordinate system to the local coordinate system and then converting back to the updated spatial coordinates. 2.3 Use ARRM components to conduct in-depth development of label plug-ins to achieve label creation, storage, display, data integration and web page display, and personalize the appearance of labels, including adjustment of shape, color, size and spatial position; 2.4 Real-time analysis of video content, using AR tagging technology to accurately mark abnormal events, behaviors, and objects, achieving precise binding of tags and objects and dynamic tracking effects. Furthermore, through a systematic tag and metadata management mechanism, the classification, organization, and management processes of video content are optimized; 2.5 By calculating and integrating geographic location information, a fused image is generated to achieve a deep integration of reality and virtuality.
4. The multi-perspective real-time rendering and digital-visual fusion method based on ultra-high-definition panoramic AR video according to claim 3 is characterized by: In the process of constructing the virtual information layer, step 2.2 calculates the global spherical coordinates based on the field of view of the dome camera, the inclination of the pan / tilt, and the horizontal rotation angle, and then maps the coordinates on the screen to the panoramic video space coordinates to initially form an AR scene. In addition, by integrating the geographic location information of the camera group, the enhanced AR scene is seamlessly integrated into the panoramic map to generate a fused image. At the same time, the latitude and longitude and north point information of the map are used to accurately calculate the orientation and visual range of the camera group, and the corresponding virtual data is intelligently matched based on the visual range, and the matched virtual data is immediately fed back to the AR scene for display.
5. The multi-view real-time rendering and digital-visual fusion method based on ultra-high-definition panoramic AR video according to claim 3 is characterized by: The specific implementation of step 2.5 is as follows: first, the pan / tilt angle and field of view angle of the current camera dome camera need to be obtained, and then a loop phase is entered. In the preliminary AR scene, scene change data triggered by the subject-object interaction is collected, and then the coordinate system conversion operation is performed using this data to establish a localized coordinate reference system. Next, this local coordinate system is seamlessly connected to the global coordinate system framework to calculate the updated spatial position coordinates. Based on these latest coordinate data, the AR scene is drawn and updated in real time to ensure that the immediate feedback of the subject object interaction is consistent with the dynamic changes of the scene; finally, in the display stage, the geographic coordinate information of the camera group is collected and integrated, and this information is used to seamlessly integrate the AR scene into the system map, ultimately generating a fused image that contains both real and virtual elements.
6. The multi-view real-time rendering and digital-visual fusion method based on ultra-high-definition panoramic AR video according to claim 1 is characterized by: The specific implementation of step (3) is as follows: 3.1 Use tag linkage and interactive presentation technology to create interactive tags that can update in real time and reflect various data in high-point videos. These tags need to be dynamically bound to low-point video sources, facial recognition systems, and vehicle checkpoint data; 3.2 To achieve picture-in-picture linkage between high- and low-level videos, buffer-based timestamp alignment and resampling technology are used to synchronize, overlay, and render high- and low-level videos. Through CDN and edge computing, adaptive streaming technology is used to dynamically adjust the playback speed of video streams based on network conditions to match the data update frequency. 3.3 3D zoom and label synchronization are used to achieve high-high linkage, high-low linkage, and panoramic detail linkage. The three.js 3D graphics engine ensures smooth image rendering. Through high-low perspective switching technology and preloading target videos, the video continuity of the subject is maintained when switching perspectives, ensuring that the delay during switching is minimized. At the same time, an easing algorithm is used to avoid visual discomfort caused by delays or inconsistencies during the switching process. 3.4 Establish a three-dimensional collaborative management system for large-scale scenes and local details to achieve all-round collaboration between high and low perspectives and between internal and external scenes. Specifically: Based on GIS map information, the location information of mobile video devices is tracked and synchronized in real time. GPS positioning technology is used to ensure the authenticity and dynamic nature of panoramic map information, and the location of devices is updated in real time to support the display of real-time data.
7. The multi-perspective real-time rendering and digital-visual fusion method based on ultra-high-definition panoramic AR video according to claim 6, characterized in that: The implementation process of step 3.1 needs to rely on accurate timestamp and GIS spatial position calibration technology to ensure the synchronization and consistency of different data sources. Specifically: timestamp synchronization is adopted, that is, all data sources must have a unified time base, and NTP is used to ensure the time consistency between different data sources. Each frame of video and each piece of data is accompanied by a timestamp for accurate matching; at the same time, GIS spatial position calibration is also adopted, and GIS technology is used for position calibration. The position of the video source is combined with the coordinates in the geographic information system, and the position data is corrected in real time through the Kalman filter algorithm.
8. The multi-perspective real-time rendering and digital-visual fusion method based on ultra-high-definition panoramic AR video according to claim 1 is characterized by: The specific implementation of step (4) is as follows: first, the entire scene is divided into several spatial partitions using space segmentation technology, and then the objects in each partition are processed; by calculating the intersection relationship between the bounding box of each object and the view cone, those objects that are completely outside the view cone are quickly eliminated; then the object to be rendered is converted to the view cone coordinate system for collision testing, and the bounding box of the object in the world coordinate system is compared with the view cone. If it is found that all vertices are outside the view cone, the object is completely clipped; If there is an intersection, the projection points of the object's vertices on the near clipping plane and the far clipping plane are compared to determine whether depth clipping is required. After clipping is completed, the rendering result is converted back to the world coordinate system for subsequent rendering and processing; the frustum clipping algorithm introduces the design concept of thread separation and combines the asynchronous mechanism for event loop processing.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: The processor is used to execute the computer program to implement the multi-perspective real-time rendering and digital-visual fusion method based on ultra-high-definition panoramic AR video as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, it implements the multi-perspective real-time rendering and digital-visual fusion method based on ultra-high-definition panoramic AR video as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Full space-time three-dimensional visualization method
CN103795976A
Dynamic video space-time virtual-real fusion method and system based on three-dimensional geographic information
CN109068103A
AR+3DGIS based hybrid reality three-dimensional dynamic space-time visual system and method
CN109561295A
Space-time position intelligent analysis method and system based on three-dimensional geographic information
CN111429584A
Full space-time video enhancement management and control system based on 3D live-action model
CN112367507A
Cited By
Video rendering method for virtual reality equipment and electronic equipment
CN121940589A