Surveillance camera control method and system

CN122802781APending Publication Date: 2026-09-22ZHEJIANG YONGXIU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611264978.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-20
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]现有技术中,固定轮巡策略难以适应场景中的实时变化,例如在重点区域突然出现异常事件时,摄像头可能正处于其他预置位,导致关键信息的错过,多摄像头之间的覆盖区域可能存在大量冗余,而某些区域则可能长期处于监控盲区,无法自动优化覆盖效率,基于运动检测的触发控制虽然能响应局部动态,但往往只关注单个摄像头视野内的运动,忽略了跨摄像头的协同关联以及场景中的语义信息,导致运动预测精度不足,容易产生误触发或切换延迟

Benefits of technology

[0015]本发明中,通过对图像数据进行时空分块,能够同时捕获运动模式与语义信息,构建的时空异构图利用节点间注意力权重精确反映分块间的关联程度,从而在复杂场景下准确识别出真正需要关注的目标区域。自适应图神经网络传播进一步增强了全局特征的表达能力,生成的显著性评分能够动态剔除冗余背景区域,将计算资源聚焦于高价值分块,有效降低后续处理的数据量并提升响应速度。基于关注区域的预测轨迹,结合各摄像头的视场边界进行时空交互代价评估,能够精准判断目标在未来时刻的切换需求与协作潜力。动态覆盖关系图通过时序拓扑分析自动识别冗余覆盖区域与覆盖盲区,消除重复监控的同时填补视觉空白,实现摄像头视场之间的无缝衔接与互补覆盖。帕累托最优调整序列的求解平衡了覆盖质量与切换代价,确保在多目标多约束条件下仍能获得全局最优的资源配置方案。云台转动参数与焦距调整参数的联合计算保证了每个动作均与目标预测位置精确匹配,避免因机械响应滞后导致的跟踪丢失或画面模糊。下发的同步控制序列使多个摄像头能够协调联动,形成自适应的动态监控网络,显著提升了对移动目标的持续追踪能力与场景覆盖的完整性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802781A_ABST
    Figure CN122802781A_ABST
Patent Text Reader

Abstract

The application provides a kind of monitoring camera control method and system, it is related to intelligent monitoring technical field, including: collecting image data and extracting space-time block motion and semantic features, construct space-time heterogeneous graph to obtain block correlation matrix, after global correlation characteristics are obtained by adaptive graph neural network propagation, calculate the saliency score screening attention area and predict motion trajectory, based on the prediction trajectory and field boundary calculation space-time interaction cost determines time sequence identification, calculate the future time field coverage of each camera and construct dynamic coverage relationship graph to carry out time sequence topological analysis to obtain redundancy and blind area, combined with time sequence identification and field switching cost solve pareto optimal adjustment sequence, calculate the rotation and focal length parameters of pan-tilt and time sequence alignment generate synchronous control strategy issue execution, adaptive collaborative adjustment is realized, improve the comprehensiveness and real-time of monitoring coverage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent surveillance technology, and in particular to a method and system for controlling surveillance cameras. Background Technology

[0002] In the current field of surveillance camera control technology, most surveillance systems set fixed preset positions or patrol paths for each camera, switching the viewing angle sequentially according to time to cover a preset area. Some solutions use inter-frame difference or optical flow to detect moving objects, triggering the camera to turn or zoom to track motion when motion is detected, or adjusting the viewing angle after prioritizing based on the area or speed of the moving region.

[0003] In existing technologies, fixed-circuit strategies struggle to adapt to real-time changes in scenarios. For instance, when an abnormal event suddenly occurs in a key area, the camera might be in another preset position, leading to the loss of crucial information. Coverage areas between multiple cameras may have significant redundancy, while some areas may remain long-term blind spots, making it impossible to automatically optimize coverage efficiency. While motion detection-based trigger control can respond to local dynamics, it often only focuses on motion within the field of view of a single camera, ignoring cross-camera collaboration and semantic information within the scene. This results in insufficient motion prediction accuracy and a tendency for false triggers or switching delays. Furthermore, existing technologies typically do not consider the timing costs of field-of-view overlap, focus adjustment time, and gimbal rotation paths during camera switching, making it difficult to achieve an optimal balance between resource utilization and real-time performance in the overall control strategy. Summary of the Invention

[0004] This invention provides a method and system for controlling surveillance cameras, which can at least solve some of the problems existing in the prior art.

[0005] A first aspect of the present invention provides a method for controlling a surveillance camera, comprising: Image data of the target area is collected and spatiotemporally segmented to extract motion and semantic features of each segment. Based on the motion and semantic features, a spatiotemporal heterogeneous graph is constructed and the attention weights between nodes are calculated to obtain the segmented association matrix. The block association matrix is ​​propagated using an adaptive graph neural network to obtain global association features. The saliency score of each block is calculated based on the global association features, and the region of interest is selected based on the saliency score. Motion prediction is performed on the region of interest to obtain the predicted trajectory. Based on the predicted trajectory and the field of view boundary corresponding to each surveillance camera, the spatiotemporal interaction cost is calculated and the temporal identifier of the region of interest is determined. Based on the region of interest and the predicted trajectory, the field of view coverage of each surveillance camera at future times is calculated. Based on the field of view coverage, a dynamic coverage relationship graph is constructed and temporal topology analysis is performed to obtain redundant coverage areas and coverage blind spots. Combining the temporal identifier and the field of view switching cost corresponding to the surveillance camera, the Pareto optimal adjustment sequence is obtained. Based on the Pareto optimal adjustment sequence, the gimbal rotation parameters and focal length adjustment parameters are calculated, and the timing identifier of the predicted trajectory is combined to perform timing alignment to generate a synchronization control strategy. The synchronization control strategy is then sent to the corresponding monitoring camera to perform the adjustment action.

[0006] In one alternative implementation, Image data of the target region is acquired and spatiotemporally segmented to extract motion and semantic features of each segment. Based on the motion and semantic features, a spatiotemporal heterogeneous graph is constructed, and the attention weights between nodes are calculated to obtain the segmented association matrix, including: Image data of the target area is collected, the image data is divided into spatial blocks according to the spatial dimension by grid, and continuous frame images are extracted along the time dimension to construct temporal blocks. Optical flow estimation is performed on the inter-frame difference images in the temporal blocks to obtain motion vector fields, and spatial pooling is performed on the motion vector fields to obtain motion features. Deep semantic encoding is performed on the temporal blocks to obtain semantic features. A set of motion nodes is constructed based on the motion features, a set of semantic nodes is constructed based on the semantic features, and a hybrid node set is heterogeneously fused with the set of motion nodes. The spatial block position coordinates and time frame index of each node in the hybrid node set are recorded and the spatiotemporal distance is calculated. Heterogeneous edge connections are established between the nodes in the hybrid node set according to spatial adjacency and temporal continuity, and a spatiotemporal heterogeneous graph is constructed. Calculate the feature similarity between each node and its corresponding neighboring nodes in the spatiotemporal heterogeneous graph. Calculate an initial attention score based on the feature similarity and the spatiotemporal distance between different nodes. Normalize the initial attention score using softmax to obtain normalized attention weights. Calculate a block association matrix by weighting the adjacency relationships of the spatiotemporal heterogeneous graph based on the normalized attention weights.

[0007] In one alternative implementation, An adaptive graph neural network is used to propagate the block association matrix to obtain global association features. Based on these global association features, a saliency score for each block is calculated, and regions of interest are selected based on these saliency scores. Motion prediction is then performed on these regions of interest to obtain predicted trajectories, including: The motion features and semantic features are concatenated to obtain the node feature vectors of each temporal block. The neighborhood connection relationship and connection weight between nodes are determined based on the block association matrix. The node feature vectors are then subjected to multi-layer graph convolution propagation based on the neighborhood connection relationship and the connection weight to obtain the node embedding vector. The node embedding vectors are then subjected to global pooling to obtain the global association features. The global correlation feature is concatenated with the node embedding vector corresponding to each temporal block and an initial saliency score is obtained through a fully connected mapping. An adaptive screening threshold is determined based on the distribution statistics of the initial saliency score, and blocks whose initial saliency scores exceed the adaptive screening threshold are regarded as regions of interest. The position change sequence of the region of interest within consecutive time frames is extracted and temporally encoded to obtain motion state features. The motion state features and the global correlation features are fused and recursively calculated through a preset temporal prediction network to obtain the predicted trajectory.

[0008] In one alternative implementation, Extracting the position change sequence of the region of interest within consecutive time frames and performing temporal encoding to obtain motion state features, fusing the motion state features and the global correlation features, and recursively calculating the predicted trajectory using a preset temporal prediction network includes: The spatial block position coordinates of the region of interest within consecutive time frames are extracted to form a position change sequence. Inter-frame displacement vector and velocity vector are calculated for the position change sequence. The acceleration vector is obtained by calculating the temporal difference of the velocity vector and then temporally concatenated and causal convolutionally encoded with the displacement vector and the velocity vector. The motion state features are obtained by combining the preset temporal position embedding. The motion state features are concatenated with the global association features and cross-attention weights are calculated. Based on the cross-attention weights, feature weighted fusion is performed to obtain an enhanced feature representation. The hidden state vector is initialized, and the enhanced feature representation is concatenated with the hidden state vector and nonlinearly mapped through a temporal prediction network to obtain state transition features. The hidden state vector is updated based on the state transition features to obtain an updated hidden state vector. The updated hidden state vector is decoded to obtain a position prediction value and an uncertainty estimate value. The position prediction value is accumulated temporally to generate a preliminary trajectory sequence. The uncertainty threshold is solved based on the statistical distribution of the uncertainty estimate value, and adaptive smoothing is performed on the position points in the preliminary trajectory sequence whose uncertainty estimate value exceeds the uncertainty threshold value. The predicted trajectory is calculated by combining the initial position coordinates of the region of interest.

[0009] In one alternative implementation, The spatiotemporal interaction cost is calculated based on the predicted trajectory and the field of view boundary corresponding to each surveillance camera, and the temporal identifier of the region of interest is determined. The field of view coverage of each surveillance camera at future times is calculated based on the region of interest and the predicted trajectory, including: The predicted position sequence of the region of interest is extracted from the predicted trajectory. The minimum Euclidean distance between the predicted position sequence and the field of view boundary of the monitoring camera is calculated and the boundary distance cost is obtained by exponential decay weighted accumulation according to the time step. The switching time cost is calculated based on the pan-tilt rotation speed constraint of the monitoring camera. The boundary distance cost and the switching time cost are subjected to hyperbolic tangent nonlinear transformation and fused to obtain the spatiotemporal interaction cost matrix. The optimal switching time is determined based on the spatiotemporal interaction cost matrix and a time sequence identifier is generated. The spatial extent of the region of interest is extracted from the predicted trajectory to generate a bounding box. The overlap area between the bounding box and the field of view projection area of ​​the surveillance camera is determined, and the area ratio of the overlap area to the bounding box is calculated to obtain the initial coverage. The motion direction vector of the predicted trajectory and the orientation direction vector of the field of view projection area are extracted, and the cosine value of the included angle is calculated. The initial coverage is corrected by direction compensation based on the cosine value of the included angle to obtain the direction-corrected coverage. The interaction cost value between the region of interest and the corresponding surveillance camera is extracted from the spatiotemporal interaction cost matrix and subjected to exponential decay transformation. The direction-corrected coverage is weighted and adjusted based on the exponential decay transformation result to obtain the field of view coverage.

[0010] In one alternative implementation, Based on the field of view coverage, a dynamic coverage relationship graph is constructed and temporal topology analysis is performed to obtain redundant coverage areas and coverage blind spots. Combining the temporal identifiers and the field of view switching costs corresponding to the surveillance cameras, the Pareto optimal adjustment sequence is obtained, including: Extract the coverage space range of each surveillance camera from the field of view coverage and calculate the spatial overlap area. Construct a coverage relationship graph with the surveillance camera as the node and the spatial overlap area as the edge weight. Extract the tracking time period from the time sequence identifier and expand the coverage relationship graph based on the tracking time period to obtain a dynamic coverage relationship graph. Identify height nodes in the dynamic coverage graph whose degree exceeds a preset degree threshold, extract the coverage space range corresponding to the neighboring nodes of the height nodes and determine the neighboring coverage joint region, calculate the intersection of the neighboring coverage joint region and the region of interest to obtain the redundant coverage region, determine the global coverage joint region based on the coverage space range and calculate the difference with the region of interest to obtain the coverage blind zone, and construct coverage constraints based on the coverage blind zone and the redundant coverage region. The switching time sequence is extracted from the time sequence identifier. Based on the switching time sequence, the pan-tilt rotation angle difference and focal length adjustment between the continuously switching surveillance cameras are determined. The field of view switching cost is obtained by combining the preset response delay coefficient and energy consumption coefficient. The coverage rate is obtained based on the global coverage joint region and the region of interest. With the goal of minimizing the field of view switching cost and maximizing the coverage rate, a multi-objective solution is performed in combination with the coverage constraint to obtain the Pareto optimal adjustment sequence.

[0011] In one alternative implementation, Based on the Pareto optimal adjustment sequence, the gimbal rotation parameters and focal length adjustment parameters are calculated, and time-series alignment is performed using the time sequence identifier of the predicted trajectory to generate a synchronization control strategy. The synchronization control strategy is then sent to the corresponding monitoring camera to execute adjustment actions, including: The target field-of-view parameters of each surveillance camera are extracted from the Pareto optimal adjustment sequence. The gimbal rotation angle increment and focal length adjustment increment are calculated based on the target field-of-view parameters and the current field-of-view parameters. The gimbal rotation angle increment and focal length adjustment increment are modified for feasibility based on the mechanical constraint parameters of the surveillance cameras to obtain the gimbal rotation parameters and focal length adjustment parameters. The execution time corresponding to the gimbal rotation parameters and focal length adjustment parameters is calculated to obtain the adjustment execution time. Extract the tracking start time of the region of interest from the temporal identifier of the predicted trajectory at each surveillance camera. Calculate the adjustment start time of each surveillance camera based on the tracking start time and the adjustment execution duration. Sort the adjustment start times and construct a temporal adjustment queue. Iterate through the temporal adjustment queue to calculate the adjustment start time difference between adjacent surveillance cameras. Mark surveillance cameras whose adjustment start time difference is less than a preset synchronization window as a synchronization adjustment group. Time-align the PTZ rotation parameters and focal length adjustment parameters within the synchronization adjustment group according to the adjustment start time to generate a synchronization control strategy. Send the synchronization control strategy to the corresponding surveillance camera to execute the adjustment action.

[0012] A second aspect of the present invention provides a surveillance camera control system, comprising: The mapping association unit is used to collect image data of the target area and perform spatiotemporal block processing to extract motion features and semantic features of each block. Based on the motion features and semantic features, a spatiotemporal heterogeneous graph is constructed and the attention weights between nodes are calculated to obtain the block association matrix. The focus prediction unit is used to perform adaptive graph neural network propagation on the block association matrix to obtain global association features, calculate the saliency score of each block based on the global association features, filter the focus area based on the saliency score, and perform motion prediction on the focus area to obtain the predicted trajectory. The coverage adjustment unit is used to calculate the spatiotemporal interaction cost based on the predicted trajectory and the field of view boundary corresponding to each surveillance camera, and determine the temporal identifier of the area of ​​interest. Based on the area of ​​interest and the predicted trajectory, it calculates the field of view coverage of each surveillance camera at future times. Based on the field of view coverage, it constructs a dynamic coverage relationship graph and performs temporal topology analysis to obtain redundant coverage areas and coverage blind spots. Combining the temporal identifier and the field of view switching cost corresponding to the surveillance camera, it solves for the Pareto optimal adjustment sequence. The synchronization control unit is used to calculate the gimbal rotation parameters and focal length adjustment parameters based on the Pareto optimal adjustment sequence, and combine the timing identifier of the predicted trajectory to perform timing alignment to generate a synchronization control strategy, and then send the synchronization control strategy to the corresponding monitoring camera to perform adjustment actions.

[0013] A third aspect of the present invention provides an electronic device, comprising: A processor and a memory for storing processor-executable instructions, wherein the processor is configured to invoke instructions stored in the memory to perform the aforementioned method.

[0014] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0015] In this invention, by spatiotemporally segmenting image data, motion patterns and semantic information can be captured simultaneously. The constructed spatiotemporally heterogeneous graph uses attention weights between nodes to accurately reflect the correlation between segments, thereby accurately identifying the target areas that truly require attention in complex scenes. Adaptive graph neural network propagation further enhances the expressive power of global features, and the generated saliency score can dynamically eliminate redundant background areas, focusing computational resources on high-value segments, effectively reducing the amount of data processed subsequently and improving response speed. Based on the predicted trajectory of the area of ​​interest, combined with the field of view boundaries of each camera, spatiotemporal interaction cost assessment can accurately determine the target's switching needs and collaboration potential in the future. The dynamic coverage relationship graph automatically identifies redundant coverage areas and coverage blind spots through temporal topology analysis, eliminating duplicate monitoring while filling visual gaps, achieving seamless connection and complementary coverage between camera fields of view. Solving the Pareto optimal adjustment sequence balances coverage quality and switching cost, ensuring that a globally optimal resource allocation scheme can still be obtained under multi-objective and multi-constraint conditions. The joint calculation of gimbal rotation parameters and focus adjustment parameters ensures that each movement is precisely matched with the predicted target position, avoiding tracking loss or image blurring due to mechanical response lag. The issued synchronous control sequence enables multiple cameras to coordinate and form an adaptive dynamic monitoring network, significantly improving the continuous tracking capability of moving targets and the integrity of scene coverage. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the surveillance camera control method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the predicted trajectory of the area of ​​interest in the surveillance camera control method according to an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the temporal variation of the field of view coverage in the surveillance camera control method according to an embodiment of the present invention; Figure 4 This is a flowchart illustrating the coverage optimization and Pareto adjustment process of the surveillance camera control method according to an embodiment of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0019] Figure 1 This is a flowchart illustrating the surveillance camera control method according to an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes: Image data of the target area is collected and spatiotemporally segmented to extract motion and semantic features of each segment. Based on the motion and semantic features, a spatiotemporal heterogeneous graph is constructed and the attention weights between nodes are calculated to obtain the segmented association matrix. The block association matrix is ​​propagated using an adaptive graph neural network to obtain global association features. The saliency score of each block is calculated based on the global association features, and the region of interest is selected based on the saliency score. Motion prediction is performed on the region of interest to obtain the predicted trajectory. Based on the predicted trajectory and the field of view boundary corresponding to each surveillance camera, the spatiotemporal interaction cost is calculated and the temporal identifier of the region of interest is determined. Based on the region of interest and the predicted trajectory, the field of view coverage of each surveillance camera at future times is calculated. Based on the field of view coverage, a dynamic coverage relationship graph is constructed and temporal topology analysis is performed to obtain redundant coverage areas and coverage blind spots. Combining the temporal identifier and the field of view switching cost corresponding to the surveillance camera, the Pareto optimal adjustment sequence is obtained. Based on the Pareto optimal adjustment sequence, the gimbal rotation parameters and focal length adjustment parameters are calculated, and the timing identifier of the predicted trajectory is combined to perform timing alignment to generate a synchronization control strategy. The synchronization control strategy is then sent to the corresponding monitoring camera to perform the adjustment action.

[0020] In one alternative implementation, Image data of the target region is acquired and spatiotemporally segmented to extract motion and semantic features of each segment. Based on the motion and semantic features, a spatiotemporal heterogeneous graph is constructed, and the attention weights between nodes are calculated to obtain the segmented association matrix, including: Image data of the target area is collected, the image data is divided into spatial blocks according to the spatial dimension by grid, and continuous frame images are extracted along the time dimension to construct temporal blocks. Optical flow estimation is performed on the inter-frame difference images in the temporal blocks to obtain motion vector fields, and spatial pooling is performed on the motion vector fields to obtain motion features. Deep semantic encoding is performed on the temporal blocks to obtain semantic features. A set of motion nodes is constructed based on the motion features, a set of semantic nodes is constructed based on the semantic features, and a hybrid node set is heterogeneously fused with the set of motion nodes. The spatial block position coordinates and time frame index of each node in the hybrid node set are recorded and the spatiotemporal distance is calculated. Heterogeneous edge connections are established between the nodes in the hybrid node set according to spatial adjacency and temporal continuity, and a spatiotemporal heterogeneous graph is constructed. Calculate the feature similarity between each node and its corresponding neighboring nodes in the spatiotemporal heterogeneous graph. Calculate an initial attention score based on the feature similarity and the spatiotemporal distance between different nodes. Normalize the initial attention score using softmax to obtain normalized attention weights. Calculate a block association matrix by weighting the adjacency relationships of the spatiotemporal heterogeneous graph based on the normalized attention weights.

[0021] When acquiring image data of the target area, a camera array deployed at the monitoring site continuously acquires video streams. The raw video stream is decoded at a fixed frame rate to obtain a set of consecutively arranged frame images in a time series. For the spatial dimension, each frame is divided into several rectangular spatial blocks using a uniform grid partitioning strategy. The number of rows and columns of the grid is adaptively determined based on the image resolution and the complexity of the target area. Each spatial block corresponds to a local region on the image plane, preserving its pixel-level texture and structural information. For the temporal dimension, multiple consecutive frames are captured along the time axis using a sliding window approach. Image sequences of the same spatial block location on consecutive frames are combined into temporal blocks, thus simultaneously preserving the spatial structural information and temporal dynamic information of the local region.

[0022] For each temporal block, inter-frame differencing is performed, subtracting adjacent frames pixel-by-pixel to obtain a differencing image. Regions with significant grayscale changes in the differencing image correspond to the locations of motion occurrences. Based on the differencing image, a gradient-based optical flow estimation algorithm (such as Lucas-Kanade or Farnebäck dense optical flow) is used to calculate pixel-level motion vectors, resulting in a motion vector field covering the entire temporal block. This motion vector field describes the displacement direction and amplitude of each pixel between adjacent frames. To reduce feature dimensionality and enhance robustness, max pooling or average pooling operations are performed on the motion vector field in the spatial dimension, aggregating local motion vectors into fixed-length motion feature vectors. These motion feature vectors compactly encode the main motion patterns within the block, including information such as motion direction distribution, mean motion amplitude, and motion consistency.

[0023] Deep semantic encoding is performed on each temporal block. The temporal blocks are then input into a pre-trained convolutional neural network backbone (such as ResNet or MobileNet), and feature maps from intermediate layers are extracted. These feature maps are then compressed into fixed-dimensional semantic feature vectors through global average pooling. The semantic feature vectors capture high-level semantic information within the block, including target category attributes, scene structure features, and target appearance descriptions. This complements the motion features—motion features focus on dynamic changes, while semantic features focus on static structure and category information.

[0024] After obtaining motion and semantic features, motion node sets and semantic node sets are constructed respectively. Each node in the motion node set uses a motion feature vector as its node attribute, and each node in the semantic node set uses a semantic feature vector as its node attribute. Both node sets maintain a one-to-one correspondence with the original spatial blocks. When heterogeneously fusing motion and semantic nodes, feature concatenation or weighted summation is used to merge the motion and semantic features corresponding to the same spatial block into a unified hybrid node attribute vector, thus obtaining a hybrid node set. Each node in the hybrid node set carries both motion and semantic information, enabling a more comprehensive description of the spatiotemporal state of the corresponding block.

[0025] Record the spatial block center coordinates (x) of each node in the hybrid node set. i y i ) and time frame index f i This is used for subsequent calculations of spatiotemporal distance. For node i and node j, their spatiotemporal distance d... ij Taking into account both spatial Euclidean distance and time interval, the calculation method is as follows: Where λ is the scale balance coefficient between the time and space dimensions, used to adjust the relative weights of time interval and spatial distance in spatiotemporal distance, x j Let y be the x-coordinate of node j. jLet be the ordinate of node j. The introduction of spatiotemporal distance enables the graph structure to simultaneously reflect the spatial proximity and temporal correlation between nodes.

[0026] Based on the hybrid node set, heterogeneous edge connections are established according to two types of relationships. Spatial adjacency relationships are determined by the grid position of spatial blocks, establishing spatial edges between corresponding nodes of adjacent blocks (including 4-neighborhoods or 8-neighborhoods) on the image plane to reflect the spatial continuity of local regions. Temporal continuity relationships are determined by the time frame index, establishing temporal edges between nodes of the same spatial block position in adjacent time frames to reflect the dynamic evolution of the same local region in the time dimension. By incorporating both types of edges into the graph structure, a spatiotemporally heterogeneous graph is constructed, in which both node and edge types are heterogeneous, enabling the simultaneous modeling of spatial structural relationships and temporal dynamic relationships.

[0027] The feature similarity between each node in the spatiotemporal heterogeneous graph and its neighboring nodes is calculated. Dot product similarity or cosine similarity is used to measure the correlation between the mixed feature vectors of two nodes. For node i and its neighboring node j, the feature similarity s... ij It is calculated through the inner product of the feature vectors of the two nodes. Based on feature similarity, it is combined with the spatiotemporal distance d. ij Calculate the initial attention score a ij Specifically, a ij =s ij -μ·d ij , where μ is the penalty coefficient for spatiotemporal distance, which is used to suppress the attention intensity between nodes that are far apart in spatiotemporal distance, so that the attention mechanism can give priority to spatiotemporally neighboring nodes while paying attention to nodes with similar features.

[0028] The initial attention scores of all neighboring nodes of node i are normalized using softmax to obtain the normalized attention weight α. ij The calculation method is as follows ,in Let represent the set of neighboring nodes of node i. The normalized attention weights are non-negative and sum over the neighborhood of each node to 1, ensuring the probabilistic interpretability of the weight allocation while avoiding numerical instability caused by differences in the size of the node's neighborhood.

[0029] Based on normalized attention weight α ijBy weighting the adjacency relationships of the spatiotemporally heterogeneous graph and replacing the weights of each edge in the original adjacency matrix with corresponding normalized attention weights, a block association matrix is ​​obtained. Each element in the block association matrix quantitatively describes the association strength between two corresponding block nodes, taking into account both semantic and motion similarity at the feature level and proximity at the spatiotemporal level. This provides a physically meaningful weighted graph structure foundation for subsequent graph neural network propagation, enabling information to flow preferentially along paths with high association in the graph, thereby more accurately capturing the spatiotemporal dependencies between local blocks within the target region.

[0030] In one alternative implementation, An adaptive graph neural network is used to propagate the block association matrix to obtain global association features. Based on these global association features, a saliency score for each block is calculated, and regions of interest are selected based on these saliency scores. Motion prediction is then performed on these regions of interest to obtain predicted trajectories, including: The motion features and semantic features are concatenated to obtain the node feature vectors of each temporal block. The neighborhood connection relationship and connection weight between nodes are determined based on the block association matrix. The node feature vectors are then subjected to multi-layer graph convolution propagation based on the neighborhood connection relationship and the connection weight to obtain the node embedding vector. The node embedding vectors are then subjected to global pooling to obtain the global association features. The global correlation feature is concatenated with the node embedding vector corresponding to each temporal block and an initial saliency score is obtained through a fully connected mapping. An adaptive screening threshold is determined based on the distribution statistics of the initial saliency score, and blocks whose initial saliency scores exceed the adaptive screening threshold are regarded as regions of interest. The position change sequence of the region of interest within consecutive time frames is extracted and temporally encoded to obtain motion state features. The motion state features and the global correlation features are fused and recursively calculated through a preset temporal prediction network to obtain the predicted trajectory.

[0031] Motion features and semantic features are concatenated to form node feature vectors for each temporal block. For each temporal block, motion features are extracted using optical flow estimation or frame difference methods, including information such as motion direction, velocity amplitude, and motion consistency. Semantic features are extracted by a pre-trained convolutional neural network, including high-level semantic information such as target category confidence, texture descriptors, and scene semantic labels. These two types of features are then concatenated along the channel dimension to obtain a vector vector of dimension d. v Node feature vectors , where d v This is the sum of the motion feature dimension and the semantic feature dimension. This concatenation operation preserves the independent semantics of the two types of features, while providing rich initial node representations for subsequent graph convolution propagation.

[0032] The neighborhood connectivity and connection weights between nodes are determined based on the block-based association matrix. Each element in the block-based association matrix reflects the association strength between two temporal blocks. By applying a threshold filter to the elements in the matrix, node pairs with association strength exceeding a preset lower bound are retained, forming a sparse adjacency structure, i.e., neighborhood connectivity. The connection weights are directly taken from the values ​​of the corresponding elements in the block-based association matrix, and after normalization, they are used as edge weights in graph convolution propagation. The sparsity processing of neighborhood connectivity effectively reduces the computational complexity of the graph structure while ensuring that the information transmission path between strongly associated nodes is not truncated.

[0033] The node feature vectors are propagated through multi-layer graph convolution based on neighborhood connectivity and connection weights to obtain the node embedding vectors. In the l-th layer of graph convolution, the hidden layer representation of node i... It is determined by the weighted aggregation result of its own upper-level representation and the neighboring nodes, and the calculation method is as follows: W (l) Let w be the learnable weight matrix of the l-th layer. ik Let σ(·) be the normalized connection weight between node i and node k, and let σ(·) be a non-linear activation function (such as ReLU). After L layers of graph convolution stacking, the node embedding vectors... It integrates the motion and semantic information of all blocks within the L-hop neighborhood centered on node I, possessing strong global perception capabilities. The number of layers L needs to be balanced between the receptive field coverage and the risk of oversmoothing; typically, 2 to 4 layers are appropriate.

[0034] Global pooling is performed on the node embedding vectors to obtain global association features. The global pooling operation aggregates the embedding vectors of all nodes in the graph into a fixed-dimensional global representation vector g. This can be implemented using mean pooling, max pooling, or attention-weighted pooling. In attention-weighted pooling, a scalar weight β is calculated for each node. i ,pass We obtain the global association features, where β i A single-layer fully connected network The result is obtained after mapping and softmax normalization. Attention-weighted pooling can adaptively highlight nodes that contribute more to the global semantics, making the global association features more discriminative, and is suitable for monitoring scenarios with uneven target distribution.

[0035] The global correlation feature g is embedded with the node embedding vector corresponding to each temporal block. By splicing, a dimension of The concatenated vector, where d represents the dimension of the node embedding vector. gThe dimension of the global association features is then used. Subsequently, a fully connected mapping layer compresses the concatenated vector into a scalar, yielding the initial significance score r for each temporal block. i The fully connected mapping layer consists of a linear transformation and a sigmoid activation function, constraining the initial saliency score between 0 and 1, facilitating uniform processing in subsequent thresholding operations. The initial saliency score comprehensively reflects the local feature importance of a single block and its relative saliency in the global context, effectively distinguishing high-activity target regions from background noise regions.

[0036] An adaptive screening threshold is determined based on the distribution statistic of the initial significance score. The initial significance scores {r} of all time series blocks are calculated. i}, calculate the corresponding mean With standard deviation σ r Adaptive filtering threshold Set as γ is a hyperparameter used to control the stringency of the screening, typically ranging from 0.5 to 1.5. This adaptive thresholding mechanism dynamically adjusts the screening criteria based on the overall activity level of the current frame: when the scene is moving rapidly, the overall score distribution is higher, and the threshold rises accordingly, preventing a large number of background blocks from being misclassified as regions of interest; when the scene is relatively still, the threshold decreases accordingly, ensuring that sparse moving targets can still be effectively captured. The initial saliency score exceeding the adaptive screening threshold is used to... The blocks are marked as regions of interest, forming the input set for subsequent motion prediction.

[0037] Extract the position change sequence of the region of interest within consecutive time frames and perform temporal encoding to obtain motion state features. For each block marked as a region of interest, trace its spatial position in historical frames along the timeline to construct a feature of length T. h The position sequence is given, where each element at any given time step contains the two-dimensional coordinates of the block center and its corresponding frame index. The position change sequence is differentially processed to obtain frame-by-frame displacement vectors, which are then concatenated with the original coordinate sequence to form the input sequence. This input sequence is then fed into a temporal coding network (such as a multi-layer LSTM or Transformer encoder) for encoding, resulting in an output dimension of d. m Motion state feature vector m i Position coding is introduced during the temporal coding process to preserve the temporal order information between frames and avoid errors in motion trend judgment due to sequence disruption.

[0038] For motion state characteristics m i Feature fusion is performed with the globally associated feature g. The fusion method adopts a hybrid strategy combining element-wise addition and concatenation: first, g is mapped to m through linear projection. i Same dimension d mThen, the projected global features and motion state features are added element-wise to obtain the fused feature vector c. i =m i +Pg, where P is a learnable linear projection matrix. The fused features c i It includes both local motion trend information and scene-level global context information, providing a more complete input representation for subsequent trajectory prediction.

[0039] fusing feature c i The input is a preset temporal prediction network that performs recursive calculations to obtain the predicted trajectory. The temporal prediction network employs an encoder-decoder structure, where the encoder further compresses the fused features, and the decoder progressively outputs the future trajectory T in an autoregressive manner. f The predicted location coordinates at each time step are used as one of the inputs to the decoder at the next time step, forming a recursive chain. The predicted trajectory is output as a temporal coordinate sequence, including the region of interest in the future time step T. f Intra-frame spatial location estimation provides a temporal basis for subsequent field-of-view coverage calculation and camera scheduling strategy generation. During the training phase, mean squared error loss is used to optimize the deviation between predicted and true coordinates. At the same time, a trajectory smoothing regularization term is introduced to constrain the continuity of the predicted trajectory and prevent jumps or jitter.

[0040] In one alternative implementation, Extracting the position change sequence of the region of interest within consecutive time frames and performing temporal encoding to obtain motion state features, fusing the motion state features and the global correlation features, and recursively calculating the predicted trajectory using a preset temporal prediction network includes: The spatial block position coordinates of the region of interest within consecutive time frames are extracted to form a position change sequence. Inter-frame displacement vector and velocity vector are calculated for the position change sequence. The acceleration vector is obtained by calculating the temporal difference of the velocity vector and then temporally concatenated and causal convolutionally encoded with the displacement vector and the velocity vector. The motion state features are obtained by combining the preset temporal position embedding. The motion state features are concatenated with the global association features and cross-attention weights are calculated. Based on the cross-attention weights, feature weighted fusion is performed to obtain an enhanced feature representation. The hidden state vector is initialized, and the enhanced feature representation is concatenated with the hidden state vector and nonlinearly mapped through a temporal prediction network to obtain state transition features. The hidden state vector is updated based on the state transition features to obtain an updated hidden state vector. The updated hidden state vector is decoded to obtain a position prediction value and an uncertainty estimate value. The position prediction value is accumulated temporally to generate a preliminary trajectory sequence. The uncertainty threshold is solved based on the statistical distribution of the uncertainty estimate value, and adaptive smoothing is performed on the position points in the preliminary trajectory sequence whose uncertainty estimate value exceeds the uncertainty threshold value. The predicted trajectory is calculated by combining the initial position coordinates of the region of interest.

[0041] For motion trajectory prediction of the region of interest, it is necessary to extract a refined motion state description from consecutive time frames. For each region of interest, the center coordinates of the spatial block in which the region is located are recorded sequentially in the corresponding historical time frame sequence, forming a sequence of length T. h The position change sequence is obtained. The coordinate difference between adjacent frames in this sequence is calculated to obtain a frame-by-frame displacement vector sequence. Then, the difference quotient is calculated on the displacement vectors of adjacent frames to obtain a velocity vector sequence, reflecting the speed and direction of the target's movement between frames. Further, the velocity vector sequence is subjected to temporal first-order difference to obtain an acceleration vector sequence, which is used to capture the dynamic change trend of the target's motion, including behavior patterns such as acceleration, deceleration, or turning.

[0042] The displacement vector sequence, velocity vector sequence, and acceleration vector sequence are concatenated frame by frame along the feature dimension to form a multidimensional motion description sequence. Causal convolutional encoding is applied to this multidimensional sequence. Each convolutional kernel of the causal convolution only perceives information from the current frame and previous frames, strictly adhering to temporal causality constraints to avoid information leakage from future frames. To enhance the model's ability to perceive relative positions between frames, a pre-defined temporal position embedding is introduced. The frame index is mapped to a fixed-dimensional embedding vector and added element-wise with the causal convolution output to finally obtain the motion state feature vector m. i , dimension d m It fully encodes the motion state information of the region of interest i within a historical time period.

[0043] In the feature fusion stage, the motion state feature vector m i The global associated feature vector g is directly concatenated along the feature dimension to obtain a dimension d. m +d g The concatenated vector. Based on this concatenated vector, m is processed using a learnable query matrix and a key matrix respectively. iA linear transformation is performed on g, and the cross-attention weights between them are calculated. These cross-attention weights reflect which dimensions of the globally associated features are more valuable for predicting the current motion state. Based on these weights, the globally associated features are weighted, and the weighted result is fused with the motion state feature vector using residual fusion to obtain the enhanced feature representation c. i Cross-attention fusion can adaptively guide the contribution of global semantic information to local motion prediction, avoiding indiscriminate interference of global features on local motion features.

[0044] During the recursive prediction phase, the hidden state vector is initialized. The dimension is consistent with the hidden layer dimension of the temporal prediction network, and the initial value is set to an all-zero vector. At each prediction step t... p (t) p =1,2,…,T f In ), the enhanced feature representation c i With the current hidden state vector The features are concatenated along the feature dimension and input into a temporal prediction network for nonlinear mapping to obtain state transition features. The temporal prediction network employs a gated recurrent structure, internally using update and reset gates to selectively fuse historical states with the current input, ensuring the effective transmission of long-range dependency information. This is based on state transition features. Update the hidden state vector to obtain the updated hidden state vector. This vector combines information from historical movement trends with current enhanced features.

[0045] Update the hidden state vector Decoding is performed, and the predicted position values ​​are output through two independent linear layers. (Including two components: x-axis and y-axis) and uncertainty estimate (A scalar, representing the reciprocal of the prediction confidence at the current prediction step). The predicted position value represents the position offset relative to the previous prediction step. This is progressively accumulated over time, that is, the position offsets of each prediction step are sequentially superimposed onto the initial position coordinates of the region of interest to obtain a preliminary trajectory sequence. The initial position coordinates are taken from the center coordinates of the spatial block corresponding to the region of interest in the last historical frame.

[0046] Based on uncertainty estimate sequence Calculate the statistical distribution and its mean. and standard deviation σ u,i The uncertainty threshold u th,i Set as ,in The hyperparameter used to control the degree of uncertainty tolerance is typically taken in the range of 1 to 3, with larger values ​​being less desirable. The value implies that smoothing is only applied to prediction points with extremely high uncertainty. This is the estimate of the uncertainty in the initial trajectory sequence. Exceeding the threshold u th,i For each location point, an adaptive smoothing process is performed using a time-series moving weighted average centered at that point. The weight of each point within the smoothing window is inversely proportional to its uncertainty estimate; that is, the lower the uncertainty of a neighboring prediction point, the greater its contribution to the smoothing result. After adaptive smoothing, the corrected trajectory sequence is combined with the initial location coordinates of the region of interest to finally obtain the predicted trajectory of region i, which contains a total of T... f The spatial coordinates of a future moment.

[0047] Figure 2 This diagram illustrates the predicted trajectory of the region of interest in the surveillance camera control method according to an embodiment of the present invention, showing the predicted trajectory of salient blocks within the target area over the next eight time steps. The predicted trajectory is generated based on the global correlation features extracted by the spatiotemporal heterogeneous block correlation matrix and adaptive graph neural network propagation, and incorporates a cross-attention fusion mechanism combining motion state features and global semantic information. In the diagram, the horizontal axis represents the normalized coordinates (in meters) of the image plane, the vertical axis represents the normalized coordinates (in meters), and the two curves represent the predicted position changes of the center point of the region of interest in the X and Y directions, respectively.

[0048] As observed in the figure, the predicted trajectory in the X direction of the region of interest gradually moves to the right from the initial position of 2.5m, reaching 5.5m at time t+4, and finally stabilizing at around 10.0m at time t+8. Overall, it shows a continuous rightward movement trend. The movement speed is relatively slow from t+1 to t+4 (approximately 0.75m / step), accelerates slightly from t+4 to t+6 (approximately 1.35m / step), and then stabilizes from t+6 to t+8. The predicted trajectory in the Y direction starts from 0.8m, slowly rises to 1.8m in the first three time steps, then accelerates to 3.2m from t+3 to t+6, then levels off and reaches 3.4m at time t+8. The overall movement direction shows a trend of moving to the upper right. Both prediction curves have undergone adaptive smoothing based on uncertainty estimates, effectively suppressing trajectory jumps caused by insufficient model confidence. This predicted trajectory provides crucial input for determining the timing of subsequent field-of-view switching for surveillance cameras and for aligning the timing of pan-tilt rotation parameters with the focal length adjustment parameters. This enables the camera to complete pre-adjustment actions before the target enters the field-of-view boundary, achieving smooth tracking.

[0049] In one alternative implementation, The spatiotemporal interaction cost is calculated based on the predicted trajectory and the field of view boundary corresponding to each surveillance camera, and the temporal identifier of the region of interest is determined. The field of view coverage of each surveillance camera at future times is calculated based on the region of interest and the predicted trajectory, including: The predicted position sequence of the region of interest is extracted from the predicted trajectory. The minimum Euclidean distance between the predicted position sequence and the field of view boundary of the monitoring camera is calculated and the boundary distance cost is obtained by exponential decay weighted accumulation according to the time step. The switching time cost is calculated based on the pan-tilt rotation speed constraint of the monitoring camera. The boundary distance cost and the switching time cost are subjected to hyperbolic tangent nonlinear transformation and fused to obtain the spatiotemporal interaction cost matrix. The optimal switching time is determined based on the spatiotemporal interaction cost matrix and a time sequence identifier is generated. The spatial extent of the region of interest is extracted from the predicted trajectory to generate a bounding box. The overlap area between the bounding box and the field of view projection area of ​​the surveillance camera is determined, and the area ratio of the overlap area to the bounding box is calculated to obtain the initial coverage. The motion direction vector of the predicted trajectory and the orientation direction vector of the field of view projection area are extracted, and the cosine value of the included angle is calculated. The initial coverage is corrected by direction compensation based on the cosine value of the included angle to obtain the direction-corrected coverage. The interaction cost value between the region of interest and the corresponding surveillance camera is extracted from the spatiotemporal interaction cost matrix and subjected to exponential decay transformation. The direction-corrected coverage is weighted and adjusted based on the exponential decay transformation result to obtain the field of view coverage.

[0050] The predicted position coordinates of each region of interest are extracted from the predicted trajectory at each future time step, forming a predicted position sequence. For each surveillance camera, its field of view boundary is represented by a spatial polygonal profile determined by the camera's current attitude parameters (including horizontal rotation angle, pitch angle, and focal length). For each predicted position point in the predicted position sequence, the shortest Euclidean distance from that point to each side of the field of view boundary polygon is calculated, and the minimum value is taken as the boundary distance value for that time step. When accumulating the boundary distance values ​​for each time step, an exponential decay weighting mechanism is introduced: the further the prediction step is from the current time, the smaller its boundary distance contribution weight; the decay coefficient is determined by both the time step index and the decay rate. Let the t-th... s The minimum Euclidean distance at each time step is If the attenuation rate is ρ, then the boundary distance cost C e The calculation method is as follows T f This represents the number of future moments in the predicted trajectory. This weighting method makes the impact of the recent predicted location on the cost more significant, meeting the timeliness requirements of camera scheduling.

[0051] The gimbal rotation speed constraint determines the shortest time required for the camera to switch from its current orientation to the target orientation. Let the maximum angular velocities of the camera gimbal in the horizontal and pitch directions be ω and ω, respectively. h and ω v The horizontal angle difference between the current orientation and the target orientation is Δθ. h The pitch angle difference is Δθ v Then the conversion time cost C t Defined as This means that the lower bound of the actual switching time is estimated by taking the longer of the two degrees of freedom, horizontal and pitch, which can objectively reflect the constraint of the camera hardware's motion capability on the switching efficiency.

[0052] The boundary distance cost C e With conversion time cost C t After undergoing nonlinear transformations using the hyperbolic tangent function and then weighted fusion, the cost C corresponding to the region of interest and camera combination in the spatiotemporal interaction cost matrix is ​​obtained. st The calculation method is C. st =η e ·tanh(C e ) + η t ·tanh(C t ), where η e and η t These are the fusion weight coefficients for boundary distance cost and conversion time cost, respectively, and they satisfy η. e +η t =1. The hyperbolic tangent transform compresses the original cost value to the (0, 1) interval, effectively suppressing the distorting effect of extreme cost values ​​on the fusion result, so that each element of the cost matrix has a uniform dimension and numerical range. By traversing all combinations of regions of interest and all surveillance cameras, a complete spatiotemporal interaction cost matrix can be constructed.

[0053] Based on the spatiotemporal interaction cost matrix, the optimal switching time is determined for each region of interest: In each time step of the predicted trajectory, the moment with the minimum spatiotemporal interaction cost that satisfies the transition time constraint is found, and this moment is taken as the optimal trigger point for camera switching. A corresponding time sequence identifier is generated for this region of interest. The time sequence identifier records the association between the region of interest and each camera, as well as the switching trigger time, for use in subsequent scheduling optimization stages.

[0054] When calculating the field of view coverage, the spatial extent of the region of interest (ROI) at each prediction time is extracted from the predicted trajectory. Specifically, based on the detection bounding box of the ROI in the image coordinate system, depth estimation information is combined and projected onto the world coordinate system to obtain the 3D bounding box of the ROI. This 3D bounding box is then projected 2D onto the camera's field of view projection plane to obtain the corresponding 2D bounding box, whose area is denoted as A. bboxCalculate the intersection area between the 2D bounding box and the camera's field of view projection area (i.e., the polygonal range of the camera's field of view under its current or predicted pose on the projection plane), and denote the intersection area as A. inter The initial coverage Defined as Its value ranges from [0, 1], when the region of interest is completely within the field of view. When the area of ​​interest and the field of view do not intersect .

[0055] The consistency between the direction of motion and the orientation of the field of view has a significant impact on coverage quality. Extract the motion direction vector v of the predicted trajectory in the current prediction step. motion This vector is obtained by normalizing the difference between adjacent predicted positions; simultaneously, the orientation direction vector v of the camera's field of view projection area is extracted. fov This vector is determined by the projection direction of the field-of-view center axis onto the projection plane. Calculate the cosine of the angle between the two vectors. (Both vectors have been normalized). When the direction of motion is consistent with the orientation of the field of view, the cosine value is close to 1, indicating that the target tends to stay within the field of view for a longer period of time; when the two directions are opposite, the cosine value is close to -1, indicating that the target is about to leave the field of view. The initial coverage is corrected by direction compensation based on the cosine value of the angle, thus correcting the coverage. The calculation method is as follows ,in The directional compensation coefficient, with a value range of (0, 1), is used to control the impact of directional consistency on coverage correction. This correction mechanism ensures that coverage assessment not only reflects the current degree of spatial overlap but also incorporates the prediction of future coverage quality based on the target's movement trend.

[0056] Extract the interaction cost C between the current area of ​​interest and the corresponding surveillance camera combination from the spatiotemporal interaction cost matrix. st The cost adjustment factor is obtained by performing an exponential decay transformation. ,in This is the attenuation intensity coefficient. When the interaction cost is low, When the value is close to 1, the adjustment range for direction correction coverage is relatively small; when the interaction cost is high, A value significantly less than 1 results in a substantial penalty discount on coverage, reflecting a decline in actual effective coverage quality under high-cost handover scenarios. Final field-of-view coverage. The calculation method is as follows The direction correction coverage is multiplied by the cost adjustment factor to obtain a field coverage evaluation value that comprehensively considers three factors: spatial overlap, motion direction consistency, and switching cost.

[0057] The field-of-view coverage of all surveillance cameras at all predicted future times is aggregated to form a coverage time-series matrix. Rows in the matrix correspond to different surveillance cameras, and columns correspond to different future times. Each element represents the average field-of-view coverage of that camera over all areas of interest at the corresponding time. This matrix provides a quantitative basis for subsequently constructing a dynamic coverage relationship map, identifying redundant coverage areas and coverage blind spots, and supports the solution process for the Pareto optimal adjustment sequence.

[0058] Figure 3 This diagram illustrates the temporal change of the field of view coverage of the surveillance camera control method according to an embodiment of the present invention, showing the dynamic evolution of the field of view coverage of the three surveillance cameras over the next five time steps. The coverage is calculated by integrating three dimensions: spatial overlap area, consistency compensation between movement direction and field of view orientation, and weighted adjustment of spatiotemporal interaction cost, which can comprehensively reflect the actual coverage capability of each camera during the target's movement.

[0059] As can be observed from the graph, the coverage of cam-01 starts at 0.73 at time t1, rises to 0.80 at time t2, peaks at 0.83 at time t3, and then falls back to 0.80 and 0.78 at times t4 and t5 respectively. Overall, the coverage remains at a high level, indicating that the camera can stably cover the area of ​​interest under its current posture. The coverage of cam-03 steadily increases from 0.44 at time t1 to 0.70 at time t5, showing a steady upward trend, indicating that the camera's field of view is gradually approaching the target's direction of movement, suggesting good tracking potential in the future. The coverage of cam-05 gradually increases from 0.27 at time t1 to 0.62 at time t5. Although its overall level is lower than the previous two, the upward trend is obvious, indicating that its field of view is converging towards the target area.

[0060] The differentiated trends of the three curves reflect the spatial positional relationship between different cameras and the target's trajectory: cam-01 shows high consistency with the target's direction of motion and stable coverage; while cam-03 and cam-05 show a continuous increase in coverage as the target gradually enters their field of view. This temporal coverage assessment provides a quantitative basis for constructing dynamic coverage relationship maps and identifying redundant coverage areas and blind spots, and is an important input parameter for solving the Pareto optimal adjustment sequence.

[0061] Figure 4 This is a flowchart illustrating the coverage optimization and Pareto adjustment process of the surveillance camera control method according to an embodiment of the present invention. In one alternative implementation, Based on the field of view coverage, a dynamic coverage relationship graph is constructed and temporal topology analysis is performed to obtain redundant coverage areas and coverage blind spots. Combining the temporal identifiers and the field of view switching costs corresponding to the surveillance cameras, the Pareto optimal adjustment sequence is obtained, including: Extract the coverage space range of each surveillance camera from the field of view coverage and calculate the spatial overlap area. Construct a coverage relationship graph with the surveillance camera as the node and the spatial overlap area as the edge weight. Extract the tracking time period from the time sequence identifier and expand the coverage relationship graph based on the tracking time period to obtain a dynamic coverage relationship graph. Identify height nodes in the dynamic coverage graph whose degree exceeds a preset degree threshold, extract the coverage space range corresponding to the neighboring nodes of the height nodes and determine the neighboring coverage joint region, calculate the intersection of the neighboring coverage joint region and the region of interest to obtain the redundant coverage region, determine the global coverage joint region based on the coverage space range and calculate the difference with the region of interest to obtain the coverage blind zone, and construct coverage constraints based on the coverage blind zone and the redundant coverage region. The switching time sequence is extracted from the time sequence identifier. Based on the switching time sequence, the pan-tilt rotation angle difference and focal length adjustment between the continuously switching surveillance cameras are determined. The field of view switching cost is obtained by combining the preset response delay coefficient and energy consumption coefficient. The coverage rate is obtained based on the global coverage joint region and the region of interest. With the goal of minimizing the field of view switching cost and maximizing the coverage rate, a multi-objective solution is performed in combination with the coverage constraint to obtain the Pareto optimal adjustment sequence.

[0062] Starting from the field-of-view coverage of each surveillance camera at future times, the projected coverage area of ​​each camera on the ground plane (or the target's active plane) is extracted. Specifically, based on the current pan-tilt-zoom (PTZ) orientation, focal length, and installation height of the camera, the field-of-view cone is projected onto the target plane, resulting in a coverage area represented by a polygon. For any two cameras c p With c q Find the intersection of the covered spatial ranges, and calculate the area A of the intersection. pq The edge weight is used as the connection weight between the two. Taking all surveillance cameras as nodes and A as the edge weight... pq Construct a static covering relationship graph with edge weights. . When A pq When the area is below the preset minimum overlap area threshold, the corresponding edges are not retained, thus ensuring the sparsity of the graph and computational efficiency.

[0063] In static coverage diagram Based on this, a temporal dimension is introduced to expand it into a dynamic coverage relationship graph. The tracking time periods for each region of interest are extracted from the time sequence identifiers, and these tracking time periods are discretized into several time slices t. k (k = 1, 2, ..., K, where K is the total number of time slices). In each time slice t kAt this point, the gimbal orientation and focal length of each camera are updated based on the predicted trajectory, and the coverage area and spatial overlap area are recalculated to obtain a time-varying edge weight sequence. The coverage graphs corresponding to each time slice are stacked sequentially, and temporal connection edges are introduced between nodes in addition to spatial coverage edges. The weights of the temporal connection edges reflect the continuity of the coverage state of the same camera between adjacent time slices, ultimately forming a dynamic coverage graph with a spatiotemporal dual topology. .

[0064] right Perform temporal topology analysis to identify values ​​exceeding a preset degree threshold D. th The degree of a node is defined as the cumulative number of valid edges between that node and other nodes across all time slices. A high-degree node indicates that the camera has significant spatial overlap with multiple other cameras in time sequence, representing a major source of redundant coverage. The set of neighboring nodes for each high-degree node is extracted, and the union of the coverage spatial ranges of these neighboring nodes across all time slices is taken to obtain the joint coverage region of the neighborhood. .calculate Area of ​​interest The intersection of the two regions is the redundant coverage area. This indicates that multiple cameras are simultaneously covering the area of ​​repeated monitoring of the target of interest.

[0065] The determination of coverage blind spots is based on the complement calculation of the global coverage joint region. The union of the coverage spatial extents of all nodes across all time slices is used to obtain the global coverage joint region. Calculate the region of interest. and The difference set, the difference set region is the coverage blind zone. This represents the target's activity area that was not covered by any camera throughout the entire tracking period. Based on redundant coverage areas. and coverage blind spots Construct coverage constraints: This requires that after the final adjustment sequence is executed, The area does not exceed the preset blind zone area limit. ,at the same time The area does not exceed the preset redundancy area limit. This forms a hard constraint condition that is incorporated into subsequent multi-objective optimization.

[0066] The calculation of the field-of-view switching cost extracts the switching time sequence {t} from the timing identifier. sw,1 , t sw,2 , ..., t sw,M}, where M is the total number of switching operations. For each switching operation, determine the adjacent camera pairs (c) participating in the switching. p c q), calculate the difference in horizontal rotation angle of the gimbal. Pitch and rotation angle difference and focal length adjustment ΔF pq Combined with the preset response delay coefficient and energy consumption coefficient Field of view switching cost The calculation comprehensively considers two dimensions: mechanical motion delay and energy consumption. The mechanical motion delay component is composed of... , The energy consumption component is determined by ΔF, which is determined in conjunction with the maximum angular velocity of the gimbal. pq With focal length drive energy consumption coefficient Jointly decided, the two components The weighted summation yields the comprehensive cost of a single handover. The summation of the single handover costs corresponding to all handover moments gives the total field-of-view handover cost for the entire adjustment sequence. , which serves as the first objective function in multi-objective optimization.

[0067] Coverage Defined as a globally covered joint region Area of ​​interest The intersection area divided by the area of ​​interest A roi ,Right now Where A(·) represents area calculation. Coverage The larger the value, the higher the proportion of the area of ​​interest covered by surveillance cameras, which serves as the second objective function in multi-objective optimization.

[0068] To minimize the total field-of-view switching cost With maximizing coverage To address the dual objectives, a multi-objective optimization problem is constructed by combining constraints on the coverage blind zone area and redundant coverage area. Since there is an inherent conflict between the two objectives—reducing handover costs often requires minimizing camera adjustments, while increasing coverage may necessitate more adjustments—a Pareto multi-objective solution method is employed. Specifically, the gimbal target orientation and focal length target values ​​of each camera at each handover moment constitute the decision variable vector, and a non-dominated sorting genetic algorithm is used to iteratively evolve the population. In each generation, a calculation is performed on each individual in the population (i.e., the candidate adjustment sequence). and Individuals on the Pareto front are retained based on non-dominated ranking and crowding distance selection mechanisms. After iterative convergence, individuals from the Pareto front are selected according to preset preference weights (prioritizing coverage that is not lower than a preset coverage lower limit). Under the premise of minimizing the switching cost, the optimal solution is selected to obtain the Pareto optimal adjustment sequence. This sequence explicitly gives the target gimbal orientation and target focal length of each camera at each switching moment, which is used for subsequent calculation of gimbal rotation parameters and focal length adjustment parameters.

[0069] In one alternative implementation, Based on the Pareto optimal adjustment sequence, the gimbal rotation parameters and focal length adjustment parameters are calculated, and time-series alignment is performed using the time sequence identifier of the predicted trajectory to generate a synchronization control strategy. The synchronization control strategy is then sent to the corresponding monitoring camera to execute adjustment actions, including: The target field-of-view parameters of each surveillance camera are extracted from the Pareto optimal adjustment sequence. The gimbal rotation angle increment and focal length adjustment increment are calculated based on the target field-of-view parameters and the current field-of-view parameters. The gimbal rotation angle increment and focal length adjustment increment are modified for feasibility based on the mechanical constraint parameters of the surveillance cameras to obtain the gimbal rotation parameters and focal length adjustment parameters. The execution time corresponding to the gimbal rotation parameters and focal length adjustment parameters is calculated to obtain the adjustment execution time. Extract the tracking start time of the region of interest from the temporal identifier of the predicted trajectory at each surveillance camera. Calculate the adjustment start time of each surveillance camera based on the tracking start time and the adjustment execution duration. Sort the adjustment start times and construct a temporal adjustment queue. Iterate through the temporal adjustment queue to calculate the adjustment start time difference between adjacent surveillance cameras. Mark surveillance cameras whose adjustment start time difference is less than a preset synchronization window as a synchronization adjustment group. Time-align the PTZ rotation parameters and focal length adjustment parameters within the synchronization adjustment group according to the adjustment start time to generate a synchronization control strategy. Send the synchronization control strategy to the corresponding surveillance camera to execute the adjustment action.

[0070] When extracting the target field-of-view parameters for each surveillance camera from the Pareto optimal adjustment sequence, it is necessary to analyze the target horizontal orientation angle, target pitch angle, and target focal length for each adjustment action unit in the sequence. The target field-of-view parameters describe the desired field-of-view state that the camera should achieve after adjustment, and together with the current field-of-view parameters, form the basis for parameter difference calculation. The current field-of-view parameters are read from the camera's real-time status register, including the current horizontal orientation angle, current pitch angle, and current focal length. The difference between the target horizontal orientation angle and the current horizontal orientation angle yields the pan-tilt unit's horizontal rotation angle increment Δθ. h,n The difference between the target pitch angle and the current pitch angle is used to obtain the gimbal pitch rotation angle increment Δθ. v,n The focal length adjustment increment Δf is obtained by subtracting the target focal length value from the current focal length value. n , where the subscript n represents the index of the nth camera adjustment action unit in the adjustment sequence.

[0071] When making feasibility corrections for angle and focal length increments, the mechanical constraint parameters of the surveillance camera need to be introduced. These mechanical constraint parameters include the upper and lower limits of the pan-tilt unit's horizontal rotation range [θ]. h,min θh,max ], Upper and lower limits of pitch direction rotation range [θ] v,min θ v,max ], focal length adjustment range upper and lower limits [f min f max [and the maximum single horizontal rotational angular velocity Ω of the gimbal] h and the maximum single-cycle angular velocity of pitch Ω v The logic for feasibility correction is as follows: if the target's horizontal orientation angle exceeds [θ]... h,min θ h,max If the range is not specified, the target horizontal orientation angle will be truncated to the nearest boundary value, and Δθ will be recalculated. h,n The same truncation logic applies to both the pitch and focal length directions. The truncated angle and focal length increments are the final gimbal rotation and focal length adjustment parameters. Let the corrected horizontal rotation angle increment be denoted as... The pitch rotation angle increment is Focal length adjustment increment is .

[0072] Adjust execution time T exec,n The calculation comprehensively considers the time required for two parallel sub-actions: gimbal rotation and focus adjustment. The time required for gimbal rotation is determined by the longer of the horizontal and vertical rotation times, which are respectively... and The maximum of the two values ​​is taken to obtain the gimbal rotation completion time T. ptz,n The time required for focus adjustment depends on the focus adjustment increment. With the rated adjustment rate v of the focus motor f The calculation yields the time required to complete the focus adjustment. Since gimbal rotation and focus adjustment can be performed in parallel, the adjustment execution time is taken as the greater of the two, i.e., T. exec,n =max(T) ptz,n T foc,n This execution time reflects the actual time required for the camera to complete all mechanical actions from receiving the control command, and is the core basis for subsequent timing alignment.

[0073] Extracting the region of interest from the temporal identifiers of the predicted trajectory at the tracking start time of each surveillance camera, the temporal identifier records the absolute timestamp of when each region of interest is expected to enter the field of view of each camera. Let T be the tracking start time corresponding to the nth camera. track,n Then adjust the start time T. start,n The camera needs to complete all adjustments before the area of ​​interest reaches the field of view; therefore, the calculation method is T. start,n =T track,n -T exec,n If the calculation result is negative or earlier than the current system time, then T will be...start,n The time is corrected to the current system time, and the camera is marked as having insufficient adjustment time in subsequent synchronization analysis so as to trigger a priority promotion mechanism at the execution level.

[0074] Adjustment start time T for all cameras start,n After sorting in ascending order, a timing adjustment queue is constructed. The timing adjustment queue is an ordered list, where each element contains the camera identifier, adjustment start time, gimbal rotation parameters, and focus adjustment parameters. When traversing the timing adjustment queue, the adjustment start time difference ΔT between adjacent elements is calculated for each pair. adj,n =T start,n+1 -T start,n When ΔT adj,n Smaller than the preset synchronization window W sync At that time, the two cameras in the adjacent pair are included in the same synchronization adjustment group. Preset synchronization window W sync The setting is determined based on the communication latency of the camera control bus in the actual deployment scenario and the maximum speed of the target movement in the scenario. The typical value range is between 50 milliseconds and 200 milliseconds, and the specific value is pre-written in the system configuration file.

[0075] The construction of the synchronization adjustment group adopts a connected component merging strategy: if camera A and camera B satisfy the synchronization condition, and camera B and camera C also satisfy the synchronization condition, then A, B, and C are merged into the same synchronization adjustment group, regardless of whether the time difference between A and C satisfies W. sync Within each synchronization adjustment group, the earliest adjustment start time T within the group is used. start,group As the unified trigger moment for this group, all control commands for the cameras within the group are issued at T. start,group The system is uniformly issued at all times, and each camera adjusts its own execution time to complete the action independently, thereby achieving time-sequential coordination of the adjustment actions of cameras within the group.

[0076] When aligning the gimbal rotation parameters and focal length adjustment parameters within the synchronous adjustment group according to the adjustment start time, a delay offset ΔT needs to be added to the control command of each camera. offset,n =T start,n -T start,group Cameras with zero delay offset will immediately begin adjustment at the unified trigger moment, while cameras with a delay offset greater than zero will wait for the corresponding offset duration after the unified trigger moment before starting adjustment, ensuring that the adjustment completion time of each camera is as aligned as possible with its respective tracking start time. The camera identifier, delay offset, and gimbal rotation parameters (including...) of each camera will be recorded. and ) and focal length adjustment parameters (including The control command frames are encapsulated into control command frames, and all control command frames are organized into a complete synchronization control strategy after being arranged in a synchronization adjustment group.

[0077] After the synchronization control strategy is generated, the control command frames of each synchronization adjustment group are sent in batches to the corresponding monitoring cameras via the camera control bus. During the sending process, the control bus prioritizes the synchronization adjustment group with the earliest adjustment start time. For multiple control command frames within the same synchronization adjustment group, they are simultaneously broadcast to all cameras within the group. Each camera, upon receiving the command, adjusts its response according to the delay offset ΔT. offset,n After a specified waiting period, the gimbal rotation and focus adjustment are executed in parallel. Once completed, the camera reports its completion status to the control center, which then updates the current field-of-view parameters of each camera, providing an accurate initial state reference for calculating the parameters of the next Pareto optimal adjustment sequence.

[0078] A second aspect of the present invention provides a surveillance camera control system, comprising: The mapping association unit is used to collect image data of the target area and perform spatiotemporal block processing to extract motion features and semantic features of each block. Based on the motion features and semantic features, a spatiotemporal heterogeneous graph is constructed and the attention weights between nodes are calculated to obtain the block association matrix. The focus prediction unit is used to perform adaptive graph neural network propagation on the block association matrix to obtain global association features, calculate the saliency score of each block based on the global association features, filter the focus area based on the saliency score, and perform motion prediction on the focus area to obtain the predicted trajectory. The coverage adjustment unit is used to calculate the spatiotemporal interaction cost based on the predicted trajectory and the field of view boundary corresponding to each surveillance camera, and determine the temporal identifier of the area of ​​interest. Based on the area of ​​interest and the predicted trajectory, it calculates the field of view coverage of each surveillance camera at future times. Based on the field of view coverage, it constructs a dynamic coverage relationship graph and performs temporal topology analysis to obtain redundant coverage areas and coverage blind spots. Combining the temporal identifier and the field of view switching cost corresponding to the surveillance camera, it solves for the Pareto optimal adjustment sequence. The synchronization control unit is used to calculate the gimbal rotation parameters and focal length adjustment parameters based on the Pareto optimal adjustment sequence, and combine the timing identifier of the predicted trajectory to perform timing alignment to generate a synchronization control strategy, and then send the synchronization control strategy to the corresponding monitoring camera to perform adjustment actions.

[0079] A third aspect of the present invention provides an electronic device, comprising: A processor and a memory for storing processor-executable instructions, wherein the processor is configured to invoke instructions stored in the memory to perform the aforementioned method.

[0080] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0081] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for controlling a surveillance camera, characterized in that, include: Image data of the target area is collected and spatiotemporally segmented to extract motion and semantic features of each segment. Based on the motion and semantic features, a spatiotemporal heterogeneous graph is constructed and the attention weights between nodes are calculated to obtain the segmented association matrix. The block association matrix is ​​propagated using an adaptive graph neural network to obtain global association features. The saliency score of each block is calculated based on the global association features, and the region of interest is selected based on the saliency score. Motion prediction is performed on the region of interest to obtain the predicted trajectory. Based on the predicted trajectory and the field of view boundary corresponding to each surveillance camera, the spatiotemporal interaction cost is calculated and the temporal identifier of the region of interest is determined. Based on the region of interest and the predicted trajectory, the field of view coverage of each surveillance camera at future times is calculated. Based on the field of view coverage, a dynamic coverage relationship graph is constructed and temporal topology analysis is performed to obtain redundant coverage areas and coverage blind spots. Combining the temporal identifier and the field of view switching cost corresponding to the surveillance camera, the Pareto optimal adjustment sequence is obtained. Based on the Pareto optimal adjustment sequence, the gimbal rotation parameters and focal length adjustment parameters are calculated, and the timing identifier of the predicted trajectory is combined to perform timing alignment to generate a synchronization control strategy. The synchronization control strategy is then sent to the corresponding monitoring camera to perform the adjustment action.

2. The method according to claim 1, characterized in that, Image data of the target region is acquired and spatiotemporally segmented to extract motion and semantic features of each segment. Based on the motion and semantic features, a spatiotemporal heterogeneous graph is constructed, and the attention weights between nodes are calculated to obtain the segmented association matrix, including: Image data of the target area is collected, the image data is divided into spatial blocks according to the spatial dimension by grid, and continuous frame images are extracted along the time dimension to construct temporal blocks. Optical flow estimation is performed on the inter-frame difference images in the temporal blocks to obtain motion vector fields, and spatial pooling is performed on the motion vector fields to obtain motion features. Deep semantic encoding is performed on the temporal blocks to obtain semantic features. A set of motion nodes is constructed based on the motion features, a set of semantic nodes is constructed based on the semantic features, and a hybrid node set is heterogeneously fused with the set of motion nodes. The spatial block position coordinates and time frame index of each node in the hybrid node set are recorded and the spatiotemporal distance is calculated. Heterogeneous edge connections are established between the nodes in the hybrid node set according to spatial adjacency and temporal continuity, and a spatiotemporal heterogeneous graph is constructed. Calculate the feature similarity between each node and its corresponding neighboring nodes in the spatiotemporal heterogeneous graph. Calculate an initial attention score based on the feature similarity and the spatiotemporal distance between different nodes. Normalize the initial attention score using softmax to obtain normalized attention weights. Calculate a block association matrix by weighting the adjacency relationships of the spatiotemporal heterogeneous graph based on the normalized attention weights.

3. The method according to claim 1, characterized in that, An adaptive graph neural network is used to propagate the block association matrix to obtain global association features. Based on these global association features, a saliency score for each block is calculated, and regions of interest are selected based on these saliency scores. Motion prediction is then performed on these regions of interest to obtain predicted trajectories, including: The motion features and semantic features are concatenated to obtain the node feature vectors of each temporal block. The neighborhood connection relationship and connection weight between nodes are determined based on the block association matrix. The node feature vectors are then subjected to multi-layer graph convolution propagation based on the neighborhood connection relationship and the connection weight to obtain the node embedding vector. The node embedding vectors are then subjected to global pooling to obtain the global association features. The global correlation feature is concatenated with the node embedding vector corresponding to each temporal block and an initial saliency score is obtained through a fully connected mapping. An adaptive screening threshold is determined based on the distribution statistics of the initial saliency score, and blocks whose initial saliency scores exceed the adaptive screening threshold are regarded as regions of interest. The position change sequence of the region of interest within consecutive time frames is extracted and temporally encoded to obtain motion state features. The motion state features and the global correlation features are fused and recursively calculated through a preset temporal prediction network to obtain the predicted trajectory.

4. The method according to claim 3, characterized in that, Extracting the position change sequence of the region of interest within consecutive time frames and performing temporal encoding to obtain motion state features, fusing the motion state features and the global correlation features, and recursively calculating the predicted trajectory using a preset temporal prediction network includes: The spatial block position coordinates of the region of interest within consecutive time frames are extracted to form a position change sequence. Inter-frame displacement vector and velocity vector are calculated for the position change sequence. The acceleration vector is obtained by calculating the temporal difference of the velocity vector and then temporally concatenated and causal convolutionally encoded with the displacement vector and the velocity vector. The motion state features are obtained by combining the preset temporal position embedding. The motion state features are concatenated with the global association features, and cross-attention weights are calculated. Based on the cross-attention weights, feature weighted fusion is performed to obtain an enhanced feature representation. The hidden state vector is initialized, and the enhanced feature representation is concatenated with the hidden state vector and nonlinearly mapped through a temporal prediction network to obtain state transition features. The hidden state vector is updated based on the state transition features to obtain an updated hidden state vector. The updated hidden state vector is decoded to obtain a position prediction value and an uncertainty estimate value. The position prediction value is accumulated temporally to generate a preliminary trajectory sequence. The uncertainty threshold is solved based on the statistical distribution of the uncertainty estimate value, and adaptive smoothing is performed on the position points in the preliminary trajectory sequence whose uncertainty estimate value exceeds the uncertainty threshold value. The predicted trajectory is calculated by combining the initial position coordinates of the region of interest.

5. The method according to claim 1, characterized in that, The spatiotemporal interaction cost is calculated based on the predicted trajectory and the field of view boundary corresponding to each surveillance camera, and the temporal identifier of the region of interest is determined. The field of view coverage of each surveillance camera at future times is calculated based on the region of interest and the predicted trajectory, including: The predicted position sequence of the region of interest is extracted from the predicted trajectory. The minimum Euclidean distance between the predicted position sequence and the field of view boundary of the monitoring camera is calculated and the boundary distance cost is obtained by exponential decay weighted accumulation according to the time step. The switching time cost is calculated based on the pan-tilt rotation speed constraint of the monitoring camera. The boundary distance cost and the switching time cost are subjected to hyperbolic tangent nonlinear transformation and fused to obtain the spatiotemporal interaction cost matrix. The optimal switching time is determined based on the spatiotemporal interaction cost matrix and a time sequence identifier is generated. The spatial extent of the region of interest is extracted from the predicted trajectory to generate a bounding box. The overlap area between the bounding box and the field of view projection area of ​​the surveillance camera is determined, and the area ratio of the overlap area to the bounding box is calculated to obtain the initial coverage. The motion direction vector of the predicted trajectory and the orientation direction vector of the field of view projection area are extracted, and the cosine value of the included angle is calculated. The initial coverage is corrected by direction compensation based on the cosine value of the included angle to obtain the direction-corrected coverage. The interaction cost value between the region of interest and the corresponding surveillance camera is extracted from the spatiotemporal interaction cost matrix and subjected to exponential decay transformation. The direction-corrected coverage is weighted and adjusted based on the exponential decay transformation result to obtain the field of view coverage.

6. The method according to claim 1, characterized in that, Based on the field of view coverage, a dynamic coverage relationship graph is constructed and temporal topology analysis is performed to obtain redundant coverage areas and coverage blind spots. Combining the temporal identifiers and the field of view switching costs corresponding to the surveillance cameras, the Pareto optimal adjustment sequence is obtained, including: Extract the coverage space range of each surveillance camera from the field of view coverage and calculate the spatial overlap area. Construct a coverage relationship graph with the surveillance camera as the node and the spatial overlap area as the edge weight. Extract the tracking time period from the time sequence identifier and expand the coverage relationship graph based on the tracking time period to obtain a dynamic coverage relationship graph. Identify height nodes in the dynamic coverage graph whose degree exceeds a preset degree threshold, extract the coverage space range corresponding to the neighboring nodes of the height nodes and determine the neighboring coverage joint region, calculate the intersection of the neighboring coverage joint region and the region of interest to obtain the redundant coverage region, determine the global coverage joint region based on the coverage space range and calculate the difference with the region of interest to obtain the coverage blind zone, and construct coverage constraints based on the coverage blind zone and the redundant coverage region. The switching time sequence is extracted from the time sequence identifier. Based on the switching time sequence, the pan-tilt rotation angle difference and focal length adjustment between the continuously switching surveillance cameras are determined. The field of view switching cost is obtained by combining the preset response delay coefficient and energy consumption coefficient. The coverage rate is obtained based on the global coverage joint region and the region of interest. With the goal of minimizing the field of view switching cost and maximizing the coverage rate, a multi-objective solution is performed in combination with the coverage constraint to obtain the Pareto optimal adjustment sequence.

7. The method according to claim 1, characterized in that, Based on the Pareto optimal adjustment sequence, the gimbal rotation parameters and focal length adjustment parameters are calculated, and time-series alignment is performed using the time-series identifier of the predicted trajectory to generate a synchronization control strategy. The synchronization control strategy is then sent to the corresponding monitoring camera to execute adjustment actions, including: The target field-of-view parameters of each surveillance camera are extracted from the Pareto optimal adjustment sequence. The gimbal rotation angle increment and focal length adjustment increment are calculated based on the target field-of-view parameters and the current field-of-view parameters. The gimbal rotation angle increment and focal length adjustment increment are modified for feasibility based on the mechanical constraint parameters of the surveillance cameras to obtain the gimbal rotation parameters and focal length adjustment parameters. The execution time corresponding to the gimbal rotation parameters and focal length adjustment parameters is calculated to obtain the adjustment execution time. Extract the tracking start time of the region of interest from the temporal identifier of the predicted trajectory at each surveillance camera. Calculate the adjustment start time of each surveillance camera based on the tracking start time and the adjustment execution duration. Sort the adjustment start times and construct a temporal adjustment queue. Iterate through the temporal adjustment queue to calculate the adjustment start time difference between adjacent surveillance cameras. Mark surveillance cameras whose adjustment start time difference is less than a preset synchronization window as a synchronization adjustment group. Time-align the PTZ rotation parameters and focal length adjustment parameters within the synchronization adjustment group according to the adjustment start time to generate a synchronization control strategy. Send the synchronization control strategy to the corresponding surveillance camera to execute the adjustment action.

8. A surveillance camera control system, used to implement the method of any one of claims 1-7, characterized in that, include: The mapping association unit is used to collect image data of the target area and perform spatiotemporal block processing to extract motion features and semantic features of each block. Based on the motion features and semantic features, a spatiotemporal heterogeneous graph is constructed and the attention weights between nodes are calculated to obtain the block association matrix. The focus prediction unit is used to perform adaptive graph neural network propagation on the block association matrix to obtain global association features, calculate the saliency score of each block based on the global association features, filter the focus area based on the saliency score, and perform motion prediction on the focus area to obtain the predicted trajectory. The coverage adjustment unit is used to calculate the spatiotemporal interaction cost based on the predicted trajectory and the field of view boundary corresponding to each surveillance camera, and determine the temporal identifier of the area of ​​interest. Based on the area of ​​interest and the predicted trajectory, it calculates the field of view coverage of each surveillance camera at future times. Based on the field of view coverage, it constructs a dynamic coverage relationship graph and performs temporal topology analysis to obtain redundant coverage areas and coverage blind spots. Combining the temporal identifier and the field of view switching cost corresponding to the surveillance camera, it solves for the Pareto optimal adjustment sequence. The synchronization control unit is used to calculate the gimbal rotation parameters and focal length adjustment parameters based on the Pareto optimal adjustment sequence, and combine the timing identifier of the predicted trajectory to perform timing alignment to generate a synchronization control strategy, and then send the synchronization control strategy to the corresponding monitoring camera to perform adjustment actions.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.