Panoramic monitoring camera video processing method and device based on cloud storage
By encoding, compressing, calibrating and multi-layer cloud storage management of panoramic surveillance camera video data, massive video data storage and processing problems are solved, efficient video processing and behavior prediction are achieved, and the performance and user experience of the surveillance system are improved.
Patent Information
- Application Number
- CN202510513387.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The massive high-definition video data generated by panoramic surveillance cameras are difficult to effectively store, process and analyze, resulting in limited storage capacity, lack of computing resources, and problems such as splicing gaps and inconsistent brightness of the monitoring screen, affecting the monitoring effect and user experience.
The video processing method of panoramic surveillance camera based on cloud storage is adopted, and the original video data is encoded and compressed and layered, and the initial data set is obtained, and the calibration frame parameter optimization calculation is performed to obtain the camera calibration parameters. Data is allocated to a three-layer cloud storage structure of the hot data layer, the temperature data layer and the cold data layer, and multi-band fusion and triangulation calculation are performed to generate a three-dimensional scene digital model, and behavior prediction analysis and design of adaptive data preloading strategies.
The differentiated modeling and behavior prediction of different regions by panoramic monitoring is realized, providing a smooth user experience, reducing video access latency and operation complexity, and significantly improving storage efficiency and system response speed.
Smart Images

Figure CN120047543A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and in particular, to a method and apparatus for video processing of a panoramic surveillance camera based on cloud storage. Background Art
[0002] The huge amount of high-definition video data generated by panoramic surveillance cameras poses great challenges to transmission, storage, and processing. Traditional surveillance systems usually adopt local storage and processing methods, which not only face the problems of limited storage capacity and scarce computing resources, but also are difficult to support cross-device access and real-time data analysis requirements. In addition, the distortion correction of panoramic images and the accuracy of multi-camera data fusion are insufficient, resulting in problems such as stitching gaps and inconsistent brightness in the surveillance images, seriously affecting the surveillance effect and user experience.
[0003] Cloud storage technology provides a new idea for solving the problem of surveillance data management. However, most existing cloud surveillance platforms adopt simple data storage strategies and fail to perform hierarchical management according to data access frequency and importance, resulting in waste of storage resources and system response delay. At the same time, when processing panoramic video data, these systems lack effective algorithm support for camera sensor calibration and image distortion correction, affecting the accuracy and reliability of data fusion. In addition, traditional surveillance analysis methods are difficult to effectively process the spatial dependence of distant targets in panoramic videos, ignore the hierarchical structure and activity pattern differences of different surveillance areas, and do not fully utilize the inherent statistical interdependence in video time series data, resulting in low behavior prediction accuracy. Summary of the Invention
[0004] The present invention provides a method and apparatus for video processing of a panoramic surveillance camera based on cloud storage. The present invention realizes differential modeling and behavior prediction for different regions in panoramic surveillance, provides a smooth user experience, and reduces video access latency and operation complexity.
[0005] In a first aspect, the present invention provides a method for video processing of a panoramic surveillance camera based on cloud storage. The method for video processing of a panoramic surveillance camera based on cloud storage includes: Performing encoding compression and hierarchical processing on the original video data collected by the panoramic surveillance camera to obtain an initial data set, and performing calibration frame parameter optimization calculation to obtain camera calibration parameters; Allocating the initial data set and the camera calibration parameters to a three-layer cloud storage structure of a hot data layer, a warm data layer, and a cold data layer to obtain processed data stored in layers; Performing multi-band fusion and triangulation calculation on the calibrated panoramic video data in the hot data layer to obtain a three-dimensional scene digital model; Analyze the processed data stored in layers and the 3D scene digital model to obtain the prediction result of the monitored object's behavior; Create an adaptive data preloading strategy and a multi-modal interaction interface control logic based on the prediction result of the monitored object's behavior and the user operation sequence.
[0006] In a second aspect, the present invention provides a panoramic surveillance camera video processing device based on cloud storage. The panoramic surveillance camera video processing device based on cloud storage includes: A hierarchical processing module, configured to perform encoding compression and hierarchical processing on the original video data collected by the panoramic surveillance camera to obtain an initial data set, and perform calibration frame parameter optimization calculation to obtain camera calibration parameters; An allocation module, configured to allocate the initial data set and the camera calibration parameters to a three-layer cloud storage structure of a hot data layer, a warm data layer, and a cold data layer to obtain processed data stored in layers; A calculation module, configured to perform multi-band fusion and triangulation calculation on the calibrated panoramic video data in the hot data layer to obtain a 3D scene digital model; An analysis module, configured to analyze the processed data stored in layers and the 3D scene digital model to obtain the prediction result of the monitored object's behavior; A creation module, configured to create an adaptive data preloading strategy and a multi-modal interaction interface control logic based on the prediction result of the monitored object's behavior and the user operation sequence.
[0007] In the technical solution provided by the present invention, through an efficient hierarchical data processing mechanism, the compression and secure transmission of video data are realized, significantly reducing the bandwidth requirement and ensuring data security; the precise camera calibration technology is adopted to solve the panoramic image stitching problem, eliminating the stitching gap and brightness inconsistency problems; the innovative three-layer cloud storage architecture performs intelligent hierarchical management according to the data access frequency, greatly improving the storage efficiency and system response speed; the high-precision 3D scene reconstruction technology realizes the hierarchical scene construction of static structures, semi-static objects, and dynamic objects, enhancing the realism of scene reproduction; the behavior prediction analysis method effectively solves the spatial dependence problem of distant targets in panoramic videos through a region decoupled graph convolutional network combined with a scene-aware spatio-temporal attention mechanism, realizing differential modeling and behavior prediction for different regions; the intelligent adaptive interaction system provides a smooth user experience based on user operation sequence prediction and multi-modal interaction mechanism, reducing video access latency and operation complexity.
[0008] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. The objectives and other advantages of the present invention are realized and attained by the structure particularly pointed out in the specification, claims as well as the drawings.
[0009] To make the above objectives, features and advantages of the present invention more comprehensible, the following specific preferred embodiments are given and detailed descriptions are made in conjunction with the accompanying drawings as follows. Description of the Drawings
[0010] Figure 1 It is a schematic diagram of an embodiment of the video processing method for a panoramic surveillance camera based on cloud storage in an embodiment of the present invention; Figure 2 It is a schematic diagram of an embodiment of the video processing device for a panoramic surveillance camera based on cloud storage in an embodiment of the present invention. Detailed Embodiments
[0011] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0012] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but optionally further includes other unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0013] For ease of understanding of this embodiment, first, a video processing method for a panoramic surveillance camera based on cloud storage disclosed in the embodiments of the present invention will be introduced in detail. As Figure 1 shown, the method includes the following steps: 101. Encode, compress and hierarchically process the original video data collected by the panoramic surveillance camera to obtain an initial data set, and perform calibration frame parameter optimization calculation to obtain camera calibration parameters; It can be understood that the execution subject of the present invention may be a video processing device for a panoramic surveillance camera based on cloud storage, or may also be a terminal or a server, and specific limitations are not made here. In the embodiments of the present invention, the server is taken as an example of the execution subject for illustration.
[0014] Specifically, extract and separate the original video data collected by the camera, isolate the key information therein, and obtain metadata containing the camera's unique identification code, geographical location information, timestamp, lens parameter identifier, and device status information. Perform encoding and compression processing on the original video data to optimize storage and transmission efficiency. Compress the video data using the H.265 encoding method and divide it into different levels according to the importance of the data. The base layer is used to store the low-resolution panoramic video content. The goal of this layer is to ensure that a basic surveillance picture can still be provided in a low-bandwidth environment. It uses a 1080P resolution and the frame rate is set at 10fps to ensure the integrity of panoramic coverage. At the same time, to improve the accuracy and detail performance of surveillance, perform high-resolution extraction on the target areas in the original video data and perform additional encoding on these key areas to obtain the enhanced layer video content. At this level, for detected dynamic targets, such as pedestrians, vehicles, or other abnormal activity areas, sample at a 4K or higher resolution and store at a frame rate of 25fps to ensure that the image quality of the key areas is clear enough to meet the requirements of intelligent analysis and recognition. Calculate the adaptive transmission strategy based on the current network bandwidth value and determine the transmission priority order of the base layer video content and the enhanced layer video content accordingly. Real-time monitor the network bandwidth status and set different transmission modes. For example, when the network bandwidth is below 5Mbps, only transmit the base layer data to ensure that the basic video picture is available. When the network bandwidth is between 5Mbps and 20Mbps, give priority to transmitting the metadata and the base layer, and delay the transmission of the enhanced layer data if the bandwidth permits. When the network bandwidth is above 20Mbps, transmit all data layers simultaneously to ensure the best video quality. To improve the stability of transmission and the integrity of data, optimize the data packets during the transmission process, including encrypting the metadata, the base layer video content, and the enhanced layer video content using AES-256 encryption to ensure the security of the data during transmission. At the same time, introduce forward error correction codes to enhance the data's ability to resist packet loss and effectively reduce the problem of video picture loss or blurring caused by network fluctuations. To ensure the reliability of data transmission, after the cloud server receives the encrypted transmission data packets, perform integrity verification and use the forward error correction mechanism to repair the lost data during the transmission process to form an initial data set. Use the Gauss-Helmert model to perform parameter optimization calculations on the calibration frames in the initial data set to obtain accurate camera calibration parameters.
[0015] Extract video frames containing calibration boards from the initial dataset. Calibration frames are video images captured at regular time intervals with a calibration board placed within the camera's field of view. Therefore, key frames that meet the calibration requirements are selected from a large amount of video data, and the corner detection algorithm is used to identify the feature points of the calibration board in the calibration frames, obtaining a set of feature point coordinates. At least 30 calibration frames from different perspectives are extracted for each camera to increase data redundancy and improve calibration accuracy. A joint optimization framework of internal parameters including focal length, principal point coordinates, radial distortion coefficients, and tangential distortion coefficients, and external parameters of rotation matrix and translation vector is established for the set of feature point coordinates. The Gauss-Helmert model is used to model this optimization problem, where the model is represented in implicit function form, i.e., F(X,L)=0, where X is the vector of camera parameters to be estimated and L is the observation vector. Since the calculation of this model involves non-linear relationships, the observation equation is linearized. By calculating the partial derivative matrix B of the observation equation and the partial derivative matrix A of the parameter equation, a linearized system of equations BV + AΔX + W = 0 is obtained, where B represents the partial derivative matrix of the observation equation with respect to the observed quantity, A represents the partial derivative matrix of the observation equation with respect to the parameter to be estimated, V is the correction vector of the observed value, ΔX is the parameter increment vector, and W is the constant term vector. Since the Gauss-Helmert model can simultaneously consider the uncertainties of both observed quantities and parameters, this method is more suitable for high-precision calibration of panoramic cameras than the traditional Zhang calibration method. Especially when multiple fish-eye lenses are involved in calibration simultaneously, this method can ensure global optimization and reduce the impact of error accumulation. During the calculation process, to improve the stability of feature point data, the Random Sample Consensus (RANSAC) algorithm is introduced to filter out outliers from the set of feature point coordinates. Due to illumination changes, reflections, or noise interference during actual shooting, some feature points may deviate or be mis-identified in the calibration frames. The RANSAC algorithm iteratively selects feature points that conform to the model hypothesis and eliminates mis-matched points, obtaining an optimized set of feature points. The algorithm randomly selects a set of feature points as the initial model, calculates the transformation relationship of these points, then searches for feature points that conform to this transformation relationship in the entire dataset, and selects the model with the largest number of conforming points as the final result. Least squares iterative calculation is performed based on the optimized set of feature points and the linearized system of equations to solve for the internal and external camera parameter values. The least squares method minimizes the sum of the squares of the errors between the observed values and the theoretically calculated values by continuously adjusting the parameters, effectively reducing the deviation caused by noise or measurement errors. In each iteration, the parameter increment ΔX is calculated and applied to the current estimated value, then the error is updated and recalculated until the norm of the parameter increment is less than a set threshold, such as 10⁻ 8, thus ensuring that the final calibration parameters converge to the optimal solution. While solving the camera parameters, in order to evaluate the stability and reliability of the calculation results, the parameter covariance matrix is calculated to quantify the uncertainty of the calibration parameters. The specific calculation formula is C x =N⁻¹σ 0 ², where N is the coefficient matrix of the normal equation, and σ 0 ² is the variance of unit weight. This covariance matrix can characterize the correlation between parameters and is used for subsequent error analysis. To achieve the goal of camera calibration, the pixel correspondence is calculated, that is, the mapping relationship between the original image and the corrected image is established. This mapping relationship is stored in the form of a lookup table, where the coordinates of each pixel in the original image correspond one-to-one with the coordinates in the corrected image, thus realizing efficient query for real-time correction. The calibration parameters of the camera and the pixel correspondence are stored in the cloud database and used for subsequent tasks such as video correction and 3D reconstruction.
[0016] 102. Assign the initial data set and the camera calibration parameters to the three-layer cloud storage structure of the hot data layer, warm data layer, and cold data layer to obtain the processed data stored in layers; Specifically, the access frequencies of the initial data set and camera calibration parameters are statistically analyzed, and their time attributes are analyzed to determine the storage priorities of different data blocks and perform hierarchical marking. Since there are significant differences in the access frequencies of the video data and calibration parameters generated by the panoramic monitoring system at different time periods, the access logs are used to record the number of data calls, and the data is classified in combination with time attributes. The data with high access frequency is marked as hot data, and as time goes by, the data with gradually decreasing access frequency is assigned to different storage layers, thereby optimizing the utilization rate of storage resources and improving the efficiency of data retrieval. After data classification is completed, for the newly generated data blocks within the first preset time period, due to their high access frequency, a high-performance SSD storage device is used to construct a hot data layer, and the memory mapping technology is used to directly map this data to the server memory to ensure fast data reading and low-latency access. In this process, according to the characteristics of panoramic video data, the storage layout of the data is optimized so that the data related to the same monitoring task can be stored continuously, reducing the overhead of random access, and these data are marked as hot data so that they can be quickly obtained by the intelligent analysis system or user requests in a short time. To improve the data access and storage efficiency, the storage structure of the hot data layer adopts an index-based fast search mechanism, enabling the system to complete data extraction and transmission within milliseconds. When the access frequency of the data drops and exceeds the first preset time period but is still within the second preset time period, these data are segmented in terms of spatial regions and time dimensions for storage in the warm data layer. Using the time window partitioning algorithm, the data in different time periods are grouped, and combined with the spatial location information of the camera, the panoramic video is segmented regionally to generate smaller data blocks. These sharded data blocks will be stored in the warm data layer composed of a hybrid storage device of SSD and HDD, where the SSD is used to store indexes and active data, while the HDD is used to store relatively long-term monitoring data to balance storage costs and access efficiency. At the same time, for the data containing abnormal behaviors or marked as important by users, according to the results of the intelligent analysis module, its retention time is automatically extended to the third preset time period to ensure that key data will not be wrongly cleared due to too long time, thus ensuring the integrity of the monitoring system and the traceability of data. When the storage time of the data exceeds the fourth preset time period and the access frequency further decreases, these historical data are compressed to reduce storage occupancy and optimize storage costs. An efficient lossless compression algorithm is used to compress the video data, and combined with the method based on key frames and incremental storage, to ensure that the storage requirements are minimized without affecting the quality of the video data. For structured data such as calibration parameters, the system adopts a special compression strategy to reduce storage overhead through redundant storage removal and data archiving, and the compressed data is stored in the object storage system to construct a cold data layer.The cold data layer is mainly used to store long-term archived data and supports on-demand retrieval. When the system or user needs to access this data, the data can be quickly restored through the indexing mechanism without affecting the overall storage performance of the system. To optimize the data management of hierarchical storage, a multi-dimensional index structure is constructed, and a consistency management mechanism is implemented for hot data, sharded data blocks, warm data, and long-term archived data to ensure that the migration of data between different storage layers does not affect its availability and integrity. In this process, multi-dimensional retrieval optimization is performed on the data based on time index, space index, and event index. Among them, the time index adopts the B+ tree structure to support fast time range queries, the space index adopts the R* tree structure to support spatial data lookup based on the camera perspective, and the event index adopts the inverted index to support the fast retrieval of key events. During the data migration process, the system maintains the unique identifier of the data and automatically updates the index when the data is transferred between different storage layers to ensure that users can always obtain the required data through a unified query interface without being affected by changes in the data hierarchy. In addition, regular integrity checks are performed on the stored data to prevent data corruption or loss, and combined with the role-based access control model, access permissions are set for the data in different storage layers to ensure data security and compliance.
[0017] 103. Perform multi-band fusion and triangulation calculations on the calibrated panoramic video data in the hot data layer to obtain a three-dimensional scene digital model; Specifically, the calibrated panoramic video data is retrieved from the hot data layer and undergoes multi-band fusion processing to eliminate the brightness differences and stitching gaps between different cameras or adjacent fields of view. This fusion technology decomposes the overlapping area images in different frequency domains and independently processes each frequency component, thereby retaining the high-frequency detail information while smoothing the low-frequency brightness differences, making the reconstructed panoramic image more natural visually. The fused image forms a complete 360-degree panoramic image. An equidistant cylindrical projection model is applied to the 360-degree panoramic image to establish the mapping relationship from panoramic coordinates to perspective projection coordinates, generating a virtual view image that can be used for stereo matching. The equidistant cylindrical projection ensures that regardless of the viewing angle of the user, a view conforming to the perspective characteristics can be obtained by constructing the conversion relationship from spherical coordinates to two-dimensional plane coordinates. This conversion relationship is represented by a projection equation, where the input is the pixel coordinates in the panoramic image, and the output is the image coordinates under the standard perspective projection. Through this mapping, the panoramic video data obtained by different cameras is compared under the standard perspective view. After completing the perspective projection conversion, the common visible points are extracted from the virtual view images of multiple panoramic surveillance cameras, and the point correspondence relationship is established using a feature matching algorithm to construct a cross-view feature point correspondence set. This process uses algorithms such as SIFT (Scale-Invariant Feature Transform) or ORB (Oriented FAST and Rotated BRIEF) to detect the feature points in each image, calculate their descriptors, and then uses the method of nearest neighbor search and geometric consistency constraint to screen out the truly matching point pairs. Due to the viewing angle differences between different cameras, the effects of scale, rotation, and perspective transformation are considered during the matching process, and the Random Sample Consensus algorithm is used to eliminate the abnormal matching points to ensure that the finally generated cross-view feature point correspondence set has high accuracy. Based on the extracted cross-view feature point correspondence set and combined with the internal and external parameters in the camera calibration parameters, the three-dimensional positions of these feature points in the world coordinate system are calculated through the principle of triangulation to generate a scene feature point cloud. The projection rays of each matching point in different camera coordinate systems are traced back through the internal and external parameters of the camera, and their intersection points are found in the three-dimensional space, and the coordinates of the intersection points are the three-dimensional positions of the feature points. Due to the presence of noise and errors in actual measurements, the least squares optimization method is used to minimize the error equation to improve the accuracy of the point cloud data. Through this step, a high-density three-dimensional feature point cloud is obtained, and each point in it corresponds to an actual physical position in the monitoring scene. The Poisson surface reconstruction algorithm is applied to the scene feature point cloud to generate a three-dimensional mesh model. This algorithm calculates a continuous surface function based on the point cloud data and generates a smooth three-dimensional mesh by solving the Poisson equation.On this basis, in order to improve the temporal adaptability of the 3D scene, a multi-level 3D scene structure is established according to the update frequencies of different objects. The static structure layer contains scene elements that remain unchanged for a long time, such as buildings, roads, and fixed facilities, and their data update cycle is relatively long; the semi-static object layer includes parked vehicles, placed items, etc., and its update cycle is between dozens of minutes and several hours; the dynamic object layer contains real-time changing targets such as pedestrians and moving vehicles, and the data of this layer needs to be updated frequently to ensure that the digital twin model can timely reflect the real-time state of the monitored area. In order to improve the realism of the 3D scene model, the multi-level 3D scene structure is combined with the material information extracted from the 360-degree panoramic image, and the physically based rendering technology (PBR) is applied for visual enhancement. The extraction of material information relies on image segmentation and texture mapping technologies. Using the color consistency detection algorithm, the color and reflection attributes of different surfaces are extracted from the panoramic image and projected onto the corresponding 3D mesh surface. Then, the PBR technology is used to simulate the light reflection relationship in the real world to generate a realistic visual effect. In the PBR model, the system comprehensively considers factors such as ambient light, diffuse reflection, specular reflection, and surface roughness, so that the finally generated 3D scene digital model can present a real visual experience under different lighting conditions and provide a more immersive interaction ability for the monitoring system.
[0018] 104. Analyze the processed data stored in layers and the 3D scene digital model to obtain the prediction results of the monitored object's behavior; Specifically, based on the data stored in layers, construct the trajectory sequence of the monitored object, and divide the 3D scene digital model into multiple semantic regions to form regional classification marker data. In this process, extract the video data in the most recent time period from the hot data layer, and combine it with the historical records in the warm data layer. Perform object detection and tracking frame by frame. Identify the monitored object through the object detection algorithm, and use the data association method to match the same object in consecutive frames to generate a complete motion trajectory. At the same time, based on the 3D scene digital model, conduct spatial division of the monitored area. According to the building structure, functional layout, and pedestrian flow distribution, divide the monitored scene into multiple semantic regions with different functions, such as entrances and exits, corridors, stairways, parking areas, etc., and assign a unique classification marker to each region for subsequent analysis of flow patterns and behavior prediction. After completing the regional division, calculate the transition probability between regions to obtain the flow pattern data between regions, and at the same time analyze the interaction relationship of the monitored objects within each semantic region to obtain the object interaction data within the region. Statistically analyze the movement frequency of the monitored object between different semantic regions, and construct a regional transition matrix, where the matrix elements represent the probability of flowing from one region to another region, thereby reflecting the flow pattern between different regions. At the same time, analyze the object interaction relationship within the region. Based on the object detection and trajectory analysis methods, calculate features such as the relative position, contact frequency, and relative speed between objects, and construct an interaction graph within the region to represent the behavior pattern of objects within the same region. Input the object interaction data within the region into a regional decoupled graph convolutional network with m parallel branches for behavior analysis and extraction of regional target behavior features. This network decouples the behavior patterns in different semantic regions to ensure that the behavior modeling within each region can be optimized according to different environmental characteristics. In the actual implementation process, the monitoring data of each semantic region will be sent into an independent graph convolution branch. This branch constructs the interaction relationship graph of this region, where the nodes represent the monitored objects and the edges represent the interaction relationship between them, and extracts the behavior features of the targets within the region through convolutional operations. These features will be aggregated in the fusion layer of the network to form regional target behavior features. Based on the regional target behavior features, construct a scene-aware spatio-temporal attention matrix, and design a time aggregation module containing four groups of weighting coefficients to fuse features at different time scales to generate spatio-temporal feature vectors. The scene-aware spatio-temporal attention matrix is used to model the spatial dependence relationship of the monitored objects at different time points. Its calculation method is based on the self-attention mechanism, where the features of each monitored object will be calculated for similarity with the features of other objects to generate a weighted matrix, and the elements of this matrix represent the degree of attention of a certain object to other objects at a certain moment. At the same time, the time aggregation module fuses information at multiple time scales through a gating mechanism, calculates the weighted features of the current moment, short-term history, medium-term history, and long-term history respectively, and dynamically adjusts their contribution weights using learnable parameters to obtain the final spatio-temporal feature vector.The spatio-temporal feature vectors are input into the encoder-decoder to generate the prediction results of the monitored object's behavior. Among them, the encoder uses a bidirectional long short-term memory network layer to ensure the bidirectional modeling ability of trajectory information. In this process, the trajectory sequence of each monitored object is encoded into a hidden state vector, which contains the motion features, spatial information, and environmental interaction patterns of the object at different time steps. The decoder uses this hidden state vector, combines the historical trajectory features, and uses a Gaussian mixture model for prediction to generate a multi-modal behavior prediction distribution. The advantage of the Gaussian mixture model is that it can capture multiple possible motion patterns. For example, in some scenarios, there are multiple potential action directions for the monitored object, and the Gaussian mixture model can model these different motion trends through multiple Gaussian components respectively, and make the final prediction output according to the probability weights. The decoder outputs a predicted trajectory sequence, where the trajectory points at each future time step not only contain the most likely coordinate positions but also contain the corresponding uncertainty distributions to provide more reliable behavior prediction results, enabling the monitoring system to identify potential abnormal behaviors in advance.
[0019] 105. Create an adaptive data preloading strategy and a multi-modal interaction interface control logic based on the monitored object behavior prediction results and the user operation sequence.
[0020] Specifically, the interaction behaviors of users in the monitoring system are recorded and analyzed to construct a user operation behavior model. Since the monitoring system supports multiple interaction methods, such as touch, mouse and keyboard operations, and gesture recognition, during the recording of user operation behaviors, the operation data of users are collected in real time, including the change trajectory of the user's perspective, the selection frequency of monitoring targets, the number of zoom ratio adjustments, and the area switching mode, and these data are stored in a time series format for analyzing the operation habits and preferences of users. To improve the accuracy of operation behavior modeling, the historical operation logs of users are combined, and pattern recognition algorithms are used to analyze the behavior patterns of different users in different monitoring tasks, and then high-frequency operation sequences are extracted, and clustering methods are used to classify similar operation patterns to form a user operation behavior model. The operation sequence of the user is input into a sequence prediction model with a long short-term memory network (LSTM) structure to obtain the probability distribution of the next operation. LSTM has the ability to capture long-term dependence relationships. When predicting user operations, the system not only considers the recent interaction records of the user, but also combines earlier operation habits to calculate the possible next operation of the user. The LSTM network receives the operation sequence of the past n steps as input, and through the calculation of multiple memory units, outputs the probability distribution of the next operation. This probability distribution includes the most likely operation categories, such as perspective movement, target tracking, zoom in and out, etc., and provides the corresponding confidence level to assist in decision-making. To improve the accuracy of prediction, the attention mechanism is used to assign different weights to the operations at different time steps to ensure higher attention is given to critical operations, thereby optimizing the prediction effect. Based on the probability distribution of the next operation and the prediction result of the monitored object's behavior, a three-layer differential preloading strategy is designed to generate hierarchical data prefetch instructions. Among them, the preloading strategy makes decisions based on two aspects of information: on the one hand, the operation prediction result of the user determines the data range that may need to be accessed in the future; on the other hand, the behavior prediction result of the monitored object provides the possible movement direction and position change of the target object, thereby guiding the system to preload the data in the relevant area in advance. To optimize the storage and transmission efficiency, the preloaded data is divided into three layers, including the basic layer, the detail layer, and the metadata layer. The basic layer is used to store low-resolution panoramic images to ensure that even in a low-bandwidth environment, users can still obtain basic monitoring information; the detail layer stores high-resolution data of key areas for providing clearer detailed images; the metadata layer includes target detection information, behavior analysis results, and three-dimensional scene parameters, etc., for assisting intelligent analysis and visualization. In the specific implementation process, according to the priority of the predicted operations of the user, the data at different levels are sorted to ensure that critical data is loaded first, while low-priority data can be loaded asynchronously in the background to reduce unnecessary bandwidth consumption. After generating the prefetch instructions, a hierarchical transmission protocol is implemented to ensure the stability and efficiency of data transmission.In this process, according to the current network conditions and the priority of user operations, the transmission order of data is dynamically adjusted to achieve the optimal preloading effect. For example, when the network bandwidth is low, the basic layer data is preferentially transmitted, and the data volume of the detail layer and the metadata layer is reduced. While when the bandwidth is sufficient, the system will load all levels of data simultaneously to provide the best visual experience. To improve the reliability of transmission, an adaptive bitrate control algorithm is adopted to adjust the encoding parameters of the video stream according to network fluctuations, thus avoiding frame freezing or loss caused by insufficient bandwidth. At the same time, forward error correction technology is used to enhance the data's ability to resist packet loss, ensuring better data quality in an environment with a high packet loss rate. Through the adaptive data preloading strategy, the system prepares the required data in advance before user operations occur, thus greatly reducing the data loading latency and improving the smoothness of the monitoring experience. After optimizing the data preloading strategy, next, a multi-modal interaction mechanism that supports touch operations, mouse and keyboard operations, and gesture recognition operations is constructed based on this strategy, and the monitoring screen is divided into a central operation area and an edge control area to form an interface layout with area-based responses. In this process, the central operation area is used for view adjustment, target selection, and video playback, while the edge control area is used to provide auxiliary functions such as parameter adjustment, quick operation buttons, and intelligent recommendation options. In the touch interaction mode, the user adjusts the view by dragging with a single finger, rotates the azimuth angle by pinching and rotating with two fingers, and adjusts the focal length by pinching and zooming with two fingers; in the mouse and keyboard mode, the system supports mouse dragging, scroll wheel zooming, and shortcut key combination operations to provide a more efficient control method; while in the gesture recognition mode, the system captures the user's gestures through the camera and recognizes specific instructions, such as gesture swiping for view adjustment and palm spreading for pausing or resuming video playback, thus providing a more natural interaction method. To optimize the interaction experience, according to the user's historical operation records, the interface layout is dynamically adjusted. For example, quick operation buttons are automatically added to the function areas frequently visited by the user to reduce unnecessary operation steps and improve the convenience of monitoring. Integrate the area-based response interface layout with the 3D scene digital model to achieve an intelligent view recommendation function. When the system detects abnormal behavior or a target of user concern, the system automatically calculates the best viewing angle and switches the view to the optimal position to ensure that the user can clearly observe the occurrence of key events. Specifically, based on the behavior prediction results of the monitored object, calculate its possible trajectory within a certain period in the future, and combine the occlusion information of the 3D scene to select a view that can provide the best visible range. To improve the adaptability of the recommended view, an adaptive adjustment strategy is adopted. Based on the initial recommended view, the user is allowed to make fine-tuning through interaction, and the next recommended calculation is automatically optimized after the user's adjustment to provide a monitoring experience that better meets the user's needs.Through the control logic of the multi-modal interaction interface, the system improves the intelligence level of monitoring and enhances the user experience, enabling the panoramic monitoring system to more efficiently handle complex monitoring scenarios and quickly provide the optimal monitoring perspective when critical events occur, so as to achieve more efficient security management and event response.
[0021] In the embodiments of the present invention, through an efficient hierarchical data processing mechanism, the compression and secure transmission of video data are realized, significantly reducing the bandwidth requirement and ensuring data security; the precise camera calibration technology is adopted to solve the panoramic image stitching problem, eliminating the stitching gap and brightness inconsistency problems; the innovative three-layer cloud storage architecture performs intelligent hierarchical management according to the data access frequency, greatly improving the storage efficiency and system response speed; the high-precision three-dimensional scene reconstruction technology realizes the hierarchical scene construction of static structures, semi-static objects and dynamic objects, enhancing the realism of scene reproduction; the behavior prediction analysis method effectively solves the spatial dependence problem of distant targets in panoramic videos through the region decoupled graph convolutional network combined with the scene-aware spatio-temporal attention mechanism, realizing the differential modeling and behavior prediction of different regions; the intelligent adaptive interaction system provides a smooth user experience based on user operation sequence prediction and multi-modal interaction mechanism, reducing the video access latency and operation complexity.
[0022] In a specific embodiment, the process of executing step 101 may specifically include the following steps: Extract and separate the original video data collected by the panoramic monitoring camera to obtain metadata including the unique identification code of the camera, geographical location information, time stamp, lens parameter identifier and device status information; Perform encoding compression on the original video data to obtain the base layer video content, and perform high-resolution extraction and encoding on the target area in the original video data to obtain the enhancement layer video content; Calculate the adaptive transmission strategy for the base layer video content and the enhancement layer video content based on the current network bandwidth value to obtain the transmission priority ranking; Add AES-256 encryption and forward error correction code to the metadata, the base layer video content and the enhancement layer video content to obtain the encrypted transmission data packet, and perform integrity verification and loss data repair on the encrypted transmission data packet to obtain the initial data set; Adopt the Gauss-Helmert model to perform parameter optimization calculation on the calibration frames in the initial data set to obtain the camera calibration parameters.
[0023] Specifically, parse the original data stream to obtain metadata including the unique identification code of the camera, geographical location information, time stamp, lens parameter identifier and device status information. Embed the unique identifier of the camera in each frame of the video data stream , which represents the device number of the camera and combines the geographical location information provided by the GPS sensor , where represents the latitude, represents the longitude, represents the altitude. At the same time, each frame of data is attached with a timestamp , whose unit is milliseconds, to ensure the synchronization of multi-camera data. The lens parameter identifier needs to be extracted from the internal parameters of the camera, where represents the focal length, , , is the radial distortion coefficient, , is the tangential distortion coefficient, and these parameters are used for subsequent distortion correction. The device status information includes the battery voltage , the device temperature and the network signal strength , which is used to monitor the device status in real time. After extraction, the size of the metadata is controlled within 5% of the video data to reduce the transmission burden. After separating the metadata, the original video data is encoded and compressed to generate the base layer video content, and high-resolution extraction is performed on the target area in the video to obtain the enhancement layer video content. During the video compression process, the H.265 encoding method is adopted to perform hierarchical encoding on the input video data . Among them, the base layer data adopts a resolution of 1080P and a frame rate of 10fps to ensure that the basic video monitoring ability can still be maintained under low-bandwidth conditions, while the enhancement layer data adopts a resolution of 4K and a frame rate of 25fps, and is encoded only for the detected target area. The extraction of the target area is achieved through a motion detection method based on the optical flow method, where the inter-frame optical flow is calculated as follows: ; where represents the pixel intensity of the video frame, is the optical flow vector of the target area, represents the pixel change rate in the time dimension. By solving the optical flow equation, the contour of the moving target is obtained, and on this basis, the target area is extracted and high-resolution encoded to form the enhancement layer data. After obtaining the base layer and enhancement layer data, an adaptive transmission strategy is calculated based on the current network bandwidth to determine the transmission priority of the data. In this strategy, the bandwidth thresholds and are defined to divide the transmission mode. When When it is, only the base layer data is transmitted ; When it is, the metadata and the base layer are preferentially transmitted, and the enhancement layer data is transmitted with a delay; When it is, all data layers are transmitted simultaneously. The optimization goal of the transmission priority is to maximize data integrity, and the optimization function is defined as follows: ; Among them represents the importance of the th layer data, represents the transmission delay of the data, and are the adjustment weights respectively. By dynamically adjusting the transmission order, while ensuring data integrity, the smoothness of video playback is improved. During the data transmission process, in order to ensure data security, the metadata, the base layer video content, and the enhancement layer video content are encrypted with AES256, and a forward error correction code is added to improve the packet loss resistance. In the encryption stage, the data D is divided into blocks, and block encryption is performed using the key K: ; Among them, C is the encrypted data, is the AES-256 encryption function. In order to improve the data loss resistance ability, Reed-Solomon code is used for forward error correction to construct redundant data blocks R, and its generation process is as follows: ; Among them, G is the generating matrix, and redundant data is generated through matrix transformation for data recovery in the case of network packet loss. The encrypted and error-corrected data packets are transmitted to the cloud server through HTTPS, and integrity verification and lost data repair are performed at the receiving end to ensure that a complete initial data set is finally formed. After obtaining the initial data set, the Gauss-Helmert model is used to perform parameter optimization calculation on the calibration frame to obtain high-precision camera calibration parameters. The video frames containing the calibration board are extracted from the initial data set, and the corner detection algorithm is used to identify the calibration points to form a set of feature points . A Gauss-Helmert error optimization model is established to minimize the error between the observed value and the theoretical calculated value. The basic equation of this model is: ; Among them, X represents the internal and external parameters of the camera to be optimized, including focal length, principal point coordinates, and distortion coefficients, and L represents the observed point coordinates. This equation is solved by least squares iteration, and the linearized equation set is as follows: ; where B is the partial derivative matrix of the observation equation with respect to the observed values, A is the partial derivative matrix of the observation equation with respect to the parameters, V is the correction of the observed values, is the parameter increment vector, and W is the constant term vector. The system solves by iterative optimization until its norm is less than the set threshold At this point. After optimization, calculate the covariance matrix of the calibration error: ; where is the variance of unit weight, and this matrix is used to evaluate the uncertainty of the calibration parameters. The optimized camera calibration parameters are stored in the cloud database and used for subsequent video processing and 3D reconstruction.
[0024] In a specific embodiment, the process of performing the steps to optimize the parameters of the calibration frames in the initial dataset using the Gauss-Helmert model to obtain the camera calibration parameters may specifically include the following steps: Extract video frames containing the calibration board from the initial dataset to obtain calibration frames, and perform feature point recognition on the calibration frames through a corner detection algorithm to obtain a set of feature point coordinates; Establish a joint optimization framework of internal parameters including focal length, principal point coordinates, radial distortion coefficients, and tangential distortion coefficients and external parameters of rotation matrix and translation vector for the set of feature point coordinates to obtain the Gauss-Helmert model; Calculate the partial derivative matrix of the observation equation and the partial derivative matrix of the parameter equation in the Gauss-Helmert model to obtain a linearized equation system; Input the set of feature point coordinates into the random sample consensus algorithm for outlier filtering to obtain an optimized set of feature points; Perform least squares iterative calculation based on the optimized set of feature points and the linearized equation system to obtain the internal and external parameter values of the camera, and calculate the parameter covariance matrix and pixel correspondence for the internal and external parameter values of the camera to obtain the camera calibration parameters.
[0025] Specifically, extract video frames containing the calibration board from the initial dataset, and use the corner detection algorithm to perform feature point recognition on these calibration frames to obtain a set of feature point coordinates. Since the calibration board adopts a checkerboard or dot array structure, the detection of feature points depends on image gradient changes and geometric feature matching. In the actual processing process, key frames containing the calibration board are screened from the panoramic video stream, and the response function of each pixel is calculated through the Harris corner detection algorithm: ; where R represents the corner response value, M is the image gradient matrix, is an empirical parameter. The matrix M is composed of gradient information and is defined as follows: ; where is the weight function, and are the gradients of the image in the x and y directions respectively. By calculating the local extreme points of R, the corner points on the calibration board are extracted and the set of feature point coordinates is formed for subsequent camera parameter estimation. After obtaining the feature point coordinates, a joint optimization framework including the internal and external parameters of the camera is established, namely the Gauss-Helmert error model. The internal parameters of the camera include the focal length f, the principal point coordinates , the radial distortion coefficients and the tangential distortion coefficients , while the external parameters consist of the rotation matrix R and the translation vector T, which are used to describe the position and orientation of the camera in the world coordinate system. To establish the joint optimization framework, an imaging equation based on perspective projection is constructed, and its mathematical expression is as follows: ; where represents the camera coordinates of the spatial point after being transformed by the rotation matrix R and the translation vector T, and the calculation formula is as follows: ; The distortion term is composed of the radial and tangential distortion terms, and its calculation formula is: ; where represents the squared distance from the pixel point to the principal point. By constructing the Gauss-Helmert model, the observation equation is written in the implicit function form , and then the partial derivative matrix B and the partial derivative matrix A of the parametric equation are calculated to construct a linearized equation system: ; where V is the correction of the observed value, is the parameter increment vector, and W is the constant term vector. This equation system is solved by least squares iteration to obtain the optimal internal and external parameters of the camera. To improve the robustness of the optimization, the set of feature point coordinates is input into the random sample consensus algorithm for outlier filtering. The minimum subset is randomly selected to calculate the model parameters, and the remaining data points are verified. Set the error threshold , and calculate the reprojection error of each feature point: ; where is the observed coordinate, are the projected coordinates calculated according to the current model. When this is the case, the point is considered an outlier and is removed. After filtering by the Random Sample Consensus algorithm, an optimized set of feature points is obtained for subsequent precise optimization. Based on the optimized set of feature points and the linearized equations, the internal and external parameters of the camera are iteratively calculated using the least squares method, and the parameter covariance matrix is calculated to quantify the uncertainty of the parameters: ; where is the variance of unit weight. By calculating the pixel correspondence, a lookup table is established to store the coordinate mapping relationship of each pixel point before and after distortion correction, forming the final camera calibration parameters.
[0026] In a specific embodiment, the process of performing step 102 may specifically include the following steps: Perform access frequency statistics and time attribute analysis on the initial data set and the camera calibration parameters to obtain data classification marks; Mark the data classification as a newly generated data block within the first preset time period and store it in the hot data layer composed of high-performance SSDs through memory mapping technology to obtain hot data; Execute a spatial region and time dimension segmentation algorithm on the data blocks marked as the second preset time period to obtain sharded data blocks; Store the sharded data blocks in the warm data layer composed of a hybrid storage architecture, and extend the retention time of data containing abnormal behaviors or marked as important to the third preset time period to obtain warm data; Perform compression on the historical data marked as exceeding the fourth preset time period to obtain compressed data, and store the compressed data in the cold data layer composed of an object storage system to obtain long-term archived data; Construct a multi-dimensional index structure and implement a consistency management mechanism for hot data, sharded data blocks, warm data, and long-term archived data to obtain processed data stored in layers.
[0027] Specifically, access frequency statistics and time attribute analysis are performed on the initial data set and the camera calibration parameters to determine the importance and access priority of the data, thereby providing a basis for the allocation of different storage layers. In this process, the access times of each data block are monitored, and its storage duration is analyzed in combination with the time attribute. Suppose a certain data block has an access frequency of within time t, then the weight of the data is expressed as: ; where represents the importance weight of the data, and is the adjustment coefficient, reflecting the characteristic of data decay over time, while is the time decay coefficient. When the access frequency of data is high and the storage time is short, takes a larger value, meaning that the data still has a high usage value. When the data has not been accessed for a long time, will drop rapidly, indicating that the data activity decreases and it needs to be migrated to a storage layer with a lower priority. According to the results of access frequency statistics and time attribute analysis, the data is classified and marked for different storage stages. For the newly generated data blocks within the first preset time period, that is, those data that have just been collected and have a high access frequency, these data need to be efficiently accessed and stored. Therefore, the memory mapping technology is used to directly store them into the hot data layer composed of high-performance SSDs. In this process, the memory mapping technology can map the data to the physical memory of the server, reducing the access latency to the millisecond level, thus significantly improving the read and write performance of the data. Suppose the video data generated by a certain surveillance camera ; wherein, is the access latency of the data, is the data size, is the read and write rate of the SSD. Since the random read speed of the SSD storage device is higher than that of the traditional HDD, this storage method effectively supports the efficient reading of real-time video streams and meets the requirements of low-latency data processing. As the data storage time increases, the access frequency of some data gradually decreases, but there is still a certain query demand. Therefore, these data are migrated to the warm data layer. In this process, the data is segmented in the spatial region and time dimension to optimize the storage management and retrieval efficiency. The spatial region segmentation is based on the physical layout of the panoramic surveillance scene. Using the R* tree indexing technology, the data is divided by geographical location, so that each storage unit corresponds to a specific area of the surveillance scene. The time dimension segmentation is based on the time window division algorithm, and the data is grouped by hour, day or week for subsequent query and analysis. Suppose a certain video data is divided into N sub-data blocks , then its storage structure in the warm data layer is expressed as: ; This block storage method can reduce unnecessary data redundancy and improve query efficiency, enabling users to quickly retrieve relevant data based on specific time periods or spatial locations. In the warm data layer, special processing is performed on data containing abnormal behaviors or marked as important by users to extend their retention time to the third preset period. This process is completed by the intelligent analysis module. The system identifies potential abnormal events through behavior analysis algorithms and dynamically adjusts the data storage strategy according to the severity of the events. When the storage time of data exceeds the fourth preset period and the access frequency continues to decline, the system compresses these historical data to reduce storage occupancy and optimize storage costs. Lossless compression algorithms, such as the H.265 video compression technology, are used for data compression to reduce the data volume while retaining necessary information. After data compression, it is stored in the object storage system to form the cold data layer to ensure long-term data availability. The cold data layer is mainly used to store archived data, with a relatively slow access rate but a high storage density, thus enabling the long-term storage of large-scale data. To ensure the efficient management of the entire storage system, a multi-dimensional index structure is constructed to uniformly manage hot data, sharded data blocks, warm data, and long-term archived data. In this process, the time index is based on the B+ tree structure to support efficient time range queries, the spatial index uses the R* tree to support location-based retrieval, and the event index uses the inverted index to support fast queries based on abnormal events. Through the multi-dimensional index mechanism, data consistency is maintained between different storage layers, and the index structure is automatically updated during data migration to ensure that users can access the required data through a unified query interface without being affected by changes in the storage hierarchy in terms of retrieval efficiency.
[0028] Among them, before performing multi - band fusion and triangulation calculation on the calibrated video data in the hot data layer after allocating the initial data set and camera calibration parameters to the three - layer cloud storage structure of the hot data layer, warm data layer, and cold data layer according to the access frequency, the following steps are also included: Analyze the cloud - distributed resource scheduling of the processed data stored in layers, allocate computing nodes based on data scale, processing complexity, and priority, and obtain a cloud - collaborative computing task graph; Decompose the cloud - collaborative computing task graph into a directed acyclic graph structure, perform dependency relationship sorting and parallelism analysis on each sub - task node, and obtain a task scheduling optimization plan; According to the task scheduling optimization plan, call the computing resources in the distributed computing cluster, allocate GPU or TPU acceleration units to each computing node and set the computing priority, and obtain a heterogeneous acceleration computing environment; Extract the video data streams of multiple panoramic cameras from the hot data layer, group them according to the principle of spatial adjacency and allocate them to different computing nodes in the heterogeneous acceleration computing environment, and obtain a parallel processing unit group; Perform data consistency checks on each parallel processing unit group, establish an incremental synchronization mechanism based on version vectors, solve the data conflict problem in the distributed environment, and obtain a data set with strong consistency guarantee; Establish a heartbeat monitoring mechanism between computing nodes, trigger a task automatic migration algorithm for nodes with network latency exceeding the threshold of 200ms, and migrate the computing tasks and data to standby nodes, and obtain a fault - tolerant processing guarantee mechanism; Verify and merge the results of each sub - task completed under the fault - tolerant processing guarantee mechanism, check the integrity of the results, and compensate and reconstruct missing or incorrect data, and obtain a highly reliable processing result; Construct a cloud data dependency map for the highly reliable processing result, record the data flow path and computing resource occupancy in the processing chain, and establish a data lineage information library, and obtain an intermediate data set that supports tracing and recomputation; Use the intermediate data set to construct a real - time statistical analysis model, dynamically monitor the node computing load, network transmission status, and data access hotspots, generate resource optimization suggestions, and obtain an adaptive resource scheduling strategy; Based on the adaptive resource scheduling strategy, perform pre - processing operations on the processed data stored in layers, including image denoising, color correction, contrast optimization, and resolution adjustment, and associate and store the pre - processed video data with the camera calibration parameters, and obtain a normalized video data set after collaborative processing.
[0029] In a specific embodiment, the process of executing step 103 may specifically include the following steps: Retrieve the calibrated panoramic video data from the hot data layer, apply multi - band fusion technology to decompose and fuse - reconstruct the overlapping area images in different frequency domains, and obtain a 360 - degree panoramic view; Apply an equidistant cylindrical projection model to the 360 - degree panoramic view to establish a mapping relationship from panoramic coordinates to perspective projection coordinates, and obtain a virtual - view image; Extract common visible points from the virtual perspective images of multiple panoramic surveillance cameras, establish point correspondence relationships through feature matching algorithms, and obtain a cross-perspective feature point correspondence set; Based on the cross-perspective feature point correspondence set and the internal and external parameters in the camera calibration parameters, calculate the three-dimensional positions of the feature points in the world coordinate system through the principle of triangulation to obtain a scene feature point cloud; Apply the Poisson surface reconstruction algorithm to the scene feature point cloud to generate a three-dimensional mesh model, and establish a multi-level three-dimensional scene structure according to the update frequencies of static structures, semi-static objects, and dynamic objects; Combine the multi-level three-dimensional scene structure with the material information extracted from the 360-degree panoramic image, and apply physically based rendering technology to generate a three-dimensional scene digital model.
[0030] Specifically, retrieve the calibrated panoramic video data from the thermal data layer, and apply multi-band fusion technology to decompose and reconstruct the images in the overlapping areas to eliminate the brightness unevenness, geometric distortion, and discontinuity in the overlapping areas between different camera perspectives. The key to multi-band fusion lies in performing frequency domain analysis on the images, separating different frequency components and fusing them within an appropriate frequency range to avoid obvious boundary effects caused by direct stitching. Adopt the method of multi-scale wavelet transform to and decompose the input overlapping area image to obtain image representations at different frequency levels: ; Among them, represents the weight synthesis of the fused image at different scales j, is the fusion weight of the images from different perspectives, and this weight is calculated based on factors such as local contrast and adaptive brightness adjustment. In this way, smooth transitions are made for the pixel information in the overlapping areas at different resolution levels to obtain a seamless stitched 360-degree panoramic image. Convert the 360-degree panoramic image to the perspective projection coordinate system to generate virtual perspective images. The equidistant cylindrical projection model is used to establish the mapping relationship between panoramic coordinates and perspective projection coordinates, and the mathematical description of this process is as follows: ; Among them, represents the perspective projection coordinates, f is the virtual focal length, and They are the azimuth angle and elevation angle in spherical coordinates respectively. Through transformation, any perspective in the panoramic image is mapped into a standard perspective image, so that subsequent matching calculations are based on the feature points after perspective transformation. Common visible points are extracted from the virtual perspective images of multiple panoramic surveillance cameras, and point correspondences are established using a feature matching algorithm to form a cross-perspective feature point correspondence set. To ensure the accuracy of matching, the SIFT (Scale-Invariant Feature Transform) algorithm is used to extract the feature points in each perspective image and calculate their descriptors. The nearest neighbor matching method is used to find similar feature point pairs in the perspective images of different cameras and filter them based on geometric constraints. During this process, in order to eliminate mis-matched points, the RANSAC (Random Sample Consensus) algorithm is applied to calculate the homography transformation matrix between feature point pairs and remove the abnormal matching points with large errors, obtaining a high-precision cross-perspective feature point correspondence set. After obtaining the cross-perspective feature point correspondence set, the three-dimensional positions of these feature points in the world coordinate system are calculated based on the principle of triangulation to construct a scene feature point cloud. Suppose a feature point P's imaging points in camera and are respectively and , according to the perspective projection equation, its projection relationship in the camera coordinate system is expressed as: ; where, is the camera internal parameter matrix, and respectively represent the rotation matrix and translation vector of the camera, and are scaling factors. By solving the linear equations, the three-dimensional positions of the feature points in the world coordinate system are solved to generate a dense three-dimensional feature point cloud. After obtaining the feature point cloud, it is structured to generate a complete three-dimensional mesh model. The Poisson surface reconstruction algorithm is used to construct a continuous three-dimensional surface. This algorithm reconstructs a smooth three-dimensional surface by calculating the gradient field of the scene and solving the Poisson equation. Let the gradient field of the scene be V, and the solution form of the Poisson equation is as follows: ; where, S is the three-dimensional surface function to be solved, is the divergence of the gradient field. By solving this equation, a smooth three-dimensional mesh structure is obtained, noise in the original point cloud is removed, and a complete scene model is generated. A multi-level three-dimensional scene structure is established according to the update frequency of the scene, and static structures, semi-static objects, and dynamic objects are stored in different data layers respectively. Among them, the static structure layer includes buildings, roads, and fixed facilities, and its data update period is long. The semi-static object layer includes vehicles in the parking lot, fixed items, etc., which are updated every few hours. The dynamic object layer stores real-time changing objects such as pedestrians and moving vehicles, and its data needs to be updated frequently to ensure the real-time nature of the monitoring system. To improve the visual realism of the three-dimensional scene, the multi-level three-dimensional scene structure is combined with the material information extracted from the 360-degree panoramic image, and the physically based rendering (PBR) technology is applied for visual enhancement. In this process, the color and reflection information of the object surface are extracted from the panoramic image and projected onto the three-dimensional mesh surface through texture mapping technology. The PBR rendering technology takes into account lighting, material properties, and environmental reflection factors, and its rendering equation is as follows: ; where, represents the outgoing light intensity at point P on the object surface, is the incident light intensity, is a physically based reflection model, and represent the incident direction and the outgoing direction respectively, and n is the surface normal. Through this rendering model, realistic lighting and material effects are generated, making the three-dimensional scene more immersive and consistent with the actual monitoring screen.
[0031] In a specific embodiment, the process of executing step 104 may specifically include the following steps: Construct a trajectory sequence of the monitoring object based on the processed data stored in layers, and divide the three-dimensional scene digital model into multiple semantic regions to obtain region classification marker data; Calculate the transfer probability between regions according to the region classification marker data to obtain inter-region flow pattern data, and calculate the interaction relationship of the monitoring objects within each semantic region to obtain intra-region object interaction data; Input the intra-region object interaction data into a region decoupling graph convolutional network with m parallel branches for behavior analysis to obtain region target behavior features; Construct a scene perception spatio-temporal attention matrix based on the region target behavior features, and design a time aggregation module containing four groups of weighting coefficients to fuse features at different time scales to obtain spatio-temporal feature vectors; Input the spatio-temporal feature vector into the encoder-decoder. Encode the trajectory sequence of the monitored object into a hidden state vector through the bidirectional LSTM layer in the encoder, and output the multi-modal prediction distribution by using the Gaussian mixture model in the decoder to obtain the behavior prediction result of the monitored object.
[0032] Specifically, construct the trajectory sequence of the monitored object based on the processed data stored in a hierarchical manner, and divide multiple semantic regions in combination with the three-dimensional scene digital model to generate regional classification marker data. The construction of the trajectory sequence depends on the object detection and tracking algorithms. The system extracts the latest video data from the hot data layer and performs time series analysis on each monitored object in combination with the historical records in the warm data layer. Assume that the position of the monitored object in the video frame at a certain moment is , where , are the horizontal and vertical coordinates in the image coordinate system respectively, then its trajectory sequence is expressed as: ; where, T is the trajectory sequence and n is the total number of time steps of the trajectory. To enhance the robustness of the trajectory, combine the three-dimensional scene information to perform world coordinate transformation on the two-dimensional image coordinates. The specific transformation formula is as follows: ; where, is the target position in the world coordinate system, R is the rotation matrix, and T is the translation vector. After the transformation, based on the three-dimensional scene structure, divide the monitored area into multiple semantic regions. Each region is marked according to geographical location, functional attributes, and pedestrian flow density, such as entrances and exits, stairs, corridors, parking lots, etc., and store the marks of these regions as regional classification marker data. After completing the regional classification, calculate the transition probability between regions to obtain the inter-regional flow pattern data, and calculate the interaction relationship of the monitored objects within each semantic region to form the intra-regional object interaction data. Set the transition probability between different regions and to be , and its calculation method is based on the statistical analysis of the trajectory data: ; where, represents the number of historical trajectories from region to region , is the total number of trajectories occurring in region . This probability matrix is used to analyze the flow trend between different regions. To calculate the interaction relationship of the objects within the region, based on the trajectory analysis method with time synchronization, define the interaction intensity of objects A and B within a certain region: ; Among them, represents the Euclidean distance between objects A and B at time t, is the distance influence factor. When the interaction intensity of objects within the same region is high, it means there is a strong social or behavioral correlation. The interaction data of objects within the region is input into a region decoupling graph convolutional network with M parallel branches for behavior analysis and extraction of regional target behavior features. This network models the behavior patterns in different semantic regions through independent graph convolution branches. The monitored objects in each semantic region are modeled as a graph , where is the set of objects within the region, is the interaction relationship between objects. The graph convolution operation is defined as follows: ; where is the node feature matrix of the th layer, A is the adjacency matrix, D is the degree matrix, is the trainable weight, is the activation function. Through multi-layer graph convolution operations, the high-order interaction relationships of the monitored objects are extracted, and region-level behavior features are obtained through pooling operations. Based on the extracted regional target behavior features, a scene-aware spatio-temporal attention matrix is constructed, and a time aggregation module containing four groups of weighted coefficients is designed to fuse features at different time scales to generate spatio-temporal feature vectors. The spatio-temporal attention matrix is used to model the dependencies between different time steps, and its calculation method is as follows: ; where and are the query vector and key vector at time steps and respectively, represents the attention weight of time step to time step , represents all possible time steps used in the normalization calculation. When calculating, the numerator part is used to calculate the similarity between and , and the denominator part sums over all possible time steps to achieve normalization and ensure that the weights sum to 1. To enhance the time modeling ability, a time aggregation module is adopted to perform multi-scale feature fusion through the learnable weighted coefficient : ; where represents the feature at the current time step, Represent the features of the past K frames. In this way, short-term and long-term behavior patterns are effectively captured to generate more robust spatio-temporal feature vectors. The spatio-temporal feature vectors are input into an encoder-decoder structure to generate the behavior prediction results of the monitored object. In the encoder part, a bidirectional long short-term memory network (LSTM) layer is adopted to capture the forward and backward dependencies of the time series simultaneously. The hidden state update formula of the trajectory sequence T in the LSTM network is as follows: ; where, is the hidden state at the current time step, is the input feature, are the trainable parameters. In the decoder part, a Gaussian mixture model is adopted for multi-modal trajectory prediction. The probability distribution of the predicted trajectory is expressed as: ; where, is the weight of the j-th Gaussian component, and are the mean and covariance matrix respectively. Through the Gaussian mixture model, multiple possible trajectory branches are generated to more accurately characterize the behavior pattern of the monitored object.
[0033] In a specific embodiment, the process of executing step 105 may specifically include the following steps: Record and analyze the interaction behavior of the user in the monitoring system to obtain the user operation behavior model; Input the operation sequence in the user operation behavior model into the sequence prediction model of the long short-term memory network structure to obtain the probability distribution of the next operation; Design a three-layer differential preloading strategy based on the probability distribution of the next operation and the behavior prediction result of the monitored object to obtain a hierarchical data prefetch instruction; Implement a hierarchical transmission protocol for the hierarchical data prefetch instruction, divide the video data into a base layer, a detail layer and a metadata layer, and dynamically adjust the data transmission order of each layer according to the network condition and the predicted operation priority to obtain an adaptive data preloading strategy; Construct a multi-modal interaction mechanism that supports touch operations, mouse and keyboard operations, and gesture recognition operations according to the adaptive data preloading strategy, divide the monitoring screen into a central operation area and an edge control area to obtain a sub-region response interface layout; Integrate the sub-region response interface layout with the three-dimensional scene digital model, and automatically calculate the best viewing angle when detecting abnormal behavior or the target of user attention to obtain the multi-modal interaction interface control logic.
[0034] Specifically, record and analyze the interaction behaviors of users in the monitoring system to construct a user operation behavior model. The operations of users in the monitoring system include perspective adjustment, zooming, target locking, timeline backtracking, etc., and these operations have certain regularities and patterns. Therefore, record the operation sequence of each user , where represents the specific interaction operation of the user at time t, such as camera switching, video dragging, target tracking, etc. To analyze the user operation mode, adopt Markov chain or graph model-based clustering methods, calculate the transition probabilities between different operations, and train the user operation behavior model based on historical interaction data to extract common operation patterns and abnormal operation behaviors. After obtaining the user operation behavior model, input the operation sequence into a sequence prediction model with a long short-term memory network (LSTM) structure to obtain the probability distribution of the possible next operations of the user. The input of the LSTM network is the operation sequence of the user in the past n time steps, and its hidden state update equation is as follows: ; where is the hidden state at the current time step, is the current user operation, are trainable parameters, is a non-linear activation function. After the LSTM encodes the operation sequence of the user, it outputs a probability distribution , indicating the possible operations and their probabilities that the user may execute at the next moment, enabling the system to anticipate the behavior of the user. After obtaining the probability distribution of the next operation, combine it with the behavior prediction result of the monitored object to design a three-layer differentiated preloading strategy and generate hierarchical data prefetch instructions. The core goal of the preloading strategy is to ensure that the required data has been cached in advance when the user executes an operation to reduce the loading delay. For this purpose, divide the video data into three layers, namely the base layer, the detail layer, and the metadata layer. The base layer contains low-resolution panoramic video data for providing a basic scene overview, the detail layer contains high-definition data of the target area to ensure the detail clarity of the key area, and the metadata layer contains object detection results, behavior prediction data, and three-dimensional scene information for intelligent analysis and auxiliary decision-making. Based on the user operation prediction and the monitored object behavior prediction, the system calculates the preloading priority: ; where U represents the preloading priority of the data, is the predicted probability of the user operation, is the predicted probability of the movement of the monitored object, and is the weight factor. Through this calculation, the system can dynamically adjust the order of data preloading and generate prefetch instructions for different levels of data. After generating the prefetch instructions, a hierarchical transmission protocol is implemented to optimize the loading efficiency of video data. Dynamically adjust the transmission order of different data layers. For example, in the case of low bandwidth, the base layer data is preferentially transmitted to ensure that users can smoothly view the basic picture, while when the bandwidth is sufficient, the detail layer and metadata layer are loaded simultaneously to provide complete information. Let the current network bandwidth be , then the transmission rate R of the hierarchical data is expressed as: ; where N is the number of data layers currently being loaded. When decreases, the system reduces the value of N to reduce bandwidth occupancy, while when increases, N is increased to provide higher-quality data. Through this dynamic adjustment mechanism, the system optimizes the data transmission efficiency under different network conditions, thereby improving the user experience. Based on the adaptive data preloading strategy, a multi-modal interaction mechanism that supports touch operations, mouse and keyboard operations, and gesture recognition operations is constructed, and the monitoring screen is divided into a central operation area and an edge control area to achieve a more intuitive interaction method. In the touch mode, the user adjusts the viewing angle by dragging with a single finger and adjusts the focal length by pinching to zoom, while in the mouse and keyboard mode, mouse dragging and shortcut key operations are supported. For the gesture recognition mode, the user's gestures are captured by a depth camera and specific instructions are recognized, such as swiping up to zoom in on the view and swiping left to switch the camera. The central operation area is used for displaying the main monitoring screen, while the edge control area contains quick operation buttons, event alerts, and intelligent recommendation options to ensure that users can quickly adjust the monitoring interface. While constructing the layout of the interaction interface, it is integrated with the three-dimensional scene digital model to achieve the function of automatically calculating the best viewing angle. When the system detects abnormal behavior or a target of user concern, the system calculates the best monitoring angle based on the three-dimensional scene information and automatically adjusts the viewing angle of the camera or provides a recommended view. Assume that the monitoring object is located at the three-dimensional coordinate , and the user's current viewing direction is , then the optimal viewing angle is calculated by an optimization function that maximizes the target visibility and minimizes the occlusion: ; where is the orientation vector of the target object, is the best viewing angle. According to the calculation results, the angle of the camera is automatically adjusted to ensure that the target object is within the best field of view. In addition, combined with the user's interaction records, the interface layout is dynamically adjusted, such as adding quick buttons in the areas that the user often pays attention to, to improve the operation efficiency.
[0035] The above describes the video processing method of the panoramic monitoring camera based on cloud storage in the embodiments of the present invention. Next, the video processing device of the panoramic monitoring camera based on cloud storage in the embodiments of the present invention will be described. Please refer to Figure 2 An embodiment of the video processing device of the panoramic monitoring camera based on cloud storage in the embodiments of the present invention includes: A hierarchical processing module 201, configured to perform encoding compression and hierarchical processing on the original video data collected by the panoramic monitoring camera to obtain an initial data set, and perform calibration frame parameter optimization calculation to obtain camera calibration parameters; An allocation module 202, configured to allocate the initial data set and the camera calibration parameters to a three-layer cloud storage structure of a hot data layer, a warm data layer, and a cold data layer to obtain processed data stored in layers; A calculation module 203, configured to perform multi-band fusion and triangulation calculation on the calibrated panoramic video data in the hot data layer to obtain a three-dimensional scene digital model; An analysis module 204, configured to analyze the processed data stored in layers and the three-dimensional scene digital model to obtain a monitoring object behavior prediction result; A creation module 205, configured to create an adaptive data preloading strategy and a multi-modal interaction interface control logic based on the monitoring object behavior prediction result and the user operation sequence.
[0036] Through the collaborative cooperation of the above-mentioned various components, through an efficient hierarchical data processing mechanism, the compression and secure transmission of video data are realized, significantly reducing the bandwidth requirement and ensuring data security; the precise camera calibration technology is adopted to solve the panoramic image stitching problem, eliminating the stitching gap and brightness inconsistency problems; the innovative three-layer cloud storage architecture performs intelligent hierarchical management according to the data access frequency, greatly improving the storage efficiency and system response speed; the high-precision three-dimensional scene reconstruction technology realizes the hierarchical scene construction of static structures, semi-static objects, and dynamic objects, enhancing the realism of scene reproduction; the behavior prediction analysis method effectively solves the spatial dependence problem of distant targets in panoramic videos through a region decoupled graph convolutional network combined with a scene-aware spatio-temporal attention mechanism, realizing differential modeling and behavior prediction of different regions; the intelligent adaptive interaction system provides a smooth user experience based on user operation sequence prediction and a multi-modal interaction mechanism, reducing the video access latency and operation complexity.
[0037] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0038] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0039] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.
Claims
1. A panoramic surveillance camera video processing method based on cloud storage, characterized in that: include: The original video data collected by the panoramic surveillance camera is encoded, compressed and layered to obtain the initial data set, and the calibration frame parameter optimization calculation is performed to obtain the camera calibration parameters; Allocating the initial data set and the camera calibration parameters to a three-layer cloud storage structure of a hot data layer, a warm data layer, and a cold data layer to obtain hierarchically stored processed data; Performing multi-band fusion and triangulation calculation on the calibrated panoramic video data in the thermal data layer to obtain a three-dimensional scene digital model; Analyzing the hierarchically stored processed data and the three-dimensional scene digital model to obtain a behavior prediction result of the monitored object; An adaptive data preloading strategy and a multimodal interactive interface control logic are created based on the monitored object behavior prediction results and the user operation sequence.
2. The cloud storage-based panoramic surveillance camera video processing method according to claim 1, characterized in that: The encoding, compression and layering of the raw video data collected by the panoramic surveillance camera are performed to obtain an initial data set, and calibration frame parameter optimization calculation is performed to obtain camera calibration parameters, including: Extract and separate the raw video data collected by the panoramic surveillance camera to obtain metadata including the camera's unique identification code, geographic location information, timestamp, lens parameter identifier, and device status information; Performing coding compression on the original video data to obtain base layer video content, and performing high-resolution extraction and coding on the target area in the original video data to obtain enhancement layer video content; Based on the current network bandwidth value, adaptive transmission strategy calculation is performed on the base layer video content and the enhancement layer video content to obtain transmission priority ranking; Add AES-256 encryption and forward error correction code to the metadata, the base layer video content, and the enhancement layer video content to obtain an encrypted transmission data packet, and perform integrity check and loss data repair on the encrypted transmission data packet to obtain an initial data set; The Gauss-Helmert model is used to perform parameter optimization calculation on the calibration frames in the initial data set to obtain camera calibration parameters.
3. The cloud storage-based panoramic surveillance camera video processing method according to claim 2, characterized in that: The Gauss-Helmert model is used to perform parameter optimization calculation on the calibration frames in the initial data set to obtain camera calibration parameters, including: Extracting a video frame containing a calibration plate from the initial data set to obtain a calibration frame, and performing feature point recognition on the calibration frame using a corner point detection algorithm to obtain a feature point coordinate set; A joint optimization framework of internal parameters including focal length, principal point coordinates, radial distortion coefficient and tangential distortion coefficient and external parameters including rotation matrix and translation vector is established for the feature point coordinate set to obtain a Gauss-Helmert model; Calculating the partial derivative matrix of the observation equation and the partial derivative matrix of the parameter equation in the Gauss-Helmert model to obtain a linearized equation group; Inputting the feature point coordinate set into a random sample consistency algorithm to perform outlier filtering processing to obtain an optimized feature point set; Based on the optimized feature point set and the linearized equation group, a least squares iterative calculation is performed to obtain the camera internal and external parameter values, and a parameter covariance matrix and a pixel correspondence relationship are calculated for the camera internal and external parameter values to obtain the camera calibration parameters.
4. The cloud storage-based panoramic surveillance camera video processing method according to claim 1, characterized in that: The method of allocating the initial data set and the camera calibration parameters to a three-layer cloud storage structure of a hot data layer, a warm data layer, and a cold data layer to obtain hierarchically stored processed data includes: Performing access frequency statistics and time attribute analysis on the initial data set and the camera calibration parameters to obtain data classification labels; The data is graded and marked as a newly generated data block within a first preset period, and stored in a hot data layer composed of a high-performance SSD through a memory mapping technology to obtain hot data; Execute a spatial region and time dimension segmentation algorithm on the data block marked as the second preset time period by the data classification to obtain a fragmented data block; Storing the sharded data blocks in a warm data layer formed by a hybrid storage architecture, and extending the retention time of data containing abnormal behavior or marked as important to a third preset period of time to obtain warm data; Compressing the historical data that is marked as exceeding a fourth preset period of time to obtain compressed data, and storing the compressed data in a cold data layer formed by an object storage system to obtain long-term archive data; A multi-dimensional index structure is constructed, and a consistency management mechanism is implemented for the hot data, the shard data blocks, the warm data, and the long-term archived data to obtain hierarchically stored processed data.
5. The cloud storage-based panoramic surveillance camera video processing method according to claim 1, characterized in that: The method of performing multi-band fusion and triangulation calculation on the calibrated panoramic video data in the thermal data layer to obtain a three-dimensional scene digital model includes: Retrieving the calibrated panoramic video data from the thermal data layer, applying multi-band fusion technology to decompose and fuse the overlapping area images in different frequency domains to obtain a 360-degree panoramic picture; Applying an equidistant cylindrical projection model to the 360-degree panoramic picture to establish a mapping relationship from panoramic coordinates to perspective projection coordinates to obtain a virtual perspective image; Extracting common visible points from the virtual view images of the plurality of panoramic surveillance cameras, establishing point correspondences through a feature matching algorithm, and obtaining a cross-view feature point correspondence set; Based on the cross-view feature point correspondence set and the internal and external parameters in the camera calibration parameters, the three-dimensional position of the feature points in the world coordinate system is calculated by triangulation principle to obtain a scene feature point cloud; Applying a Poisson surface reconstruction algorithm to the scene feature point cloud to generate a three-dimensional mesh model, and establishing a multi-level three-dimensional scene structure according to the update frequency of static structures, semi-static objects and dynamic objects; The multi-level three-dimensional scene structure is combined with the material information extracted from the 360-degree panoramic picture, and a three-dimensional scene digital model is generated using a physically based rendering technology.
6. The cloud storage-based panoramic surveillance camera video processing method according to claim 1, characterized in that: The step of analyzing the hierarchically stored processed data and the three-dimensional scene digital model to obtain a behavior prediction result of the monitored object includes: Building a trajectory sequence of the monitored object based on the hierarchically stored processed data, and dividing the three-dimensional scene digital model into a plurality of semantic regions to obtain region classification label data; Calculating the inter-regional transfer probability according to the regional classification label data to obtain inter-regional flow pattern data, and calculating the interaction relationship of the monitored objects in each semantic region to obtain the object interaction data in the region; Inputting the object interaction data in the region into a regional decoupled graph accumulation network containing m parallel branches to perform behavior analysis, and obtaining regional target behavior characteristics; Based on the target behavior characteristics of the region, a scene perception spatiotemporal attention matrix is constructed, and a time aggregation module including four sets of weighted coefficients is designed to fuse features of different time scales to obtain a spatiotemporal feature vector; The spatiotemporal feature vector is input into the encoder-decoder, the trajectory sequence of the monitored object is encoded into a hidden state vector through the bidirectional LSTM layer in the encoder, and the multimodal prediction distribution is output by the decoder using a Gaussian mixture model to obtain the behavior prediction result of the monitored object.
7. The cloud storage-based panoramic surveillance camera video processing method according to claim 1, characterized in that: The method of creating an adaptive data preloading strategy and a multi-modal interactive interface control logic based on the monitored object behavior prediction result and the user operation sequence includes: Record and analyze the user's interactive behavior in the monitoring system to obtain the user operation behavior model; Inputting the operation sequence in the user operation behavior model into the sequence prediction model of the long short-term memory network structure to obtain the probability distribution of the next operation; Designing a three-layer differentiated preloading strategy based on the probability distribution of the next operation and the behavior prediction result of the monitored object to obtain a hierarchical data prefetch instruction; Implementing a layered transmission protocol on the layered data prefetch instruction, dividing the video data into a basic layer, a detail layer and a metadata layer, dynamically adjusting the data transmission order of each layer according to the network status and the predicted operation priority, and obtaining an adaptive data preloading strategy; According to the adaptive data preloading strategy, a multimodal interaction mechanism supporting touch operation, mouse and keyboard operation and gesture recognition operation is constructed, and the monitoring screen is divided into a central operation area and an edge control area to obtain a regional response interface layout; The sub-region response interface layout is integrated with the three-dimensional scene digital model, and when abnormal behavior or a target of user concern is detected, the optimal viewing angle is automatically calculated to obtain a multi-modal interactive interface control logic.
8. A panoramic surveillance camera video processing device based on cloud storage, characterized in that: Used to execute the panoramic surveillance camera video processing method based on cloud storage according to any one of claims 1 to 7, the device comprises: The layered processing module is used to encode, compress and layer the raw video data collected by the panoramic surveillance camera to obtain an initial data set, and perform calibration frame parameter optimization calculation to obtain camera calibration parameters; An allocation module, used to allocate the initial data set and the camera calibration parameters to a three-layer cloud storage structure of a hot data layer, a warm data layer and a cold data layer, to obtain hierarchically stored processed data; A calculation module, used for performing multi-band fusion and triangulation calculation on the calibrated panoramic video data in the thermal data layer to obtain a three-dimensional scene digital model; An analysis module, used for analyzing the hierarchically stored processed data and the three-dimensional scene digital model to obtain a behavior prediction result of the monitored object; A creation module is used to create an adaptive data preloading strategy and a multimodal interactive interface control logic based on the monitoring object behavior prediction results and the user operation sequence.
Citation Information
Cited By
Multi-defense-area intelligent linkage alarm method based on AIoT gateway and related equipment
CN120431699A
Panoramic image three-dimensional reconstruction method based on 3DGS
CN120526066A
Object monitoring method and device based on multi-source data and storage medium
CN121462725A
Sampling frequency adjusting method and system for campus security and protection system
CN122316990A