Air-ground collaborative tracking method and system based on privacy-enhanced federated continual learning
Patent Information
- Application Number
- CN202611376924.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-09-07
- Publication Date
- 2026-10-02
AI Technical Summary
[0006]为了解决现有技术中无法实现保持数据隐私的前提下实现巡逻车端模型的持续进化与空地高效协同追踪的问题,本发明提供一种基于隐私增强联邦持续学习的空地协同追踪方法,其可以在保障数据安全的同时,实现巡逻车与无人航空器的联合持续追踪
[0010]本申请提供的基于隐私增强联邦持续学习的空地协同追踪方法及系统,其通过Shamir秘密分享联邦学习,使得车端敏感梯度数据在传输与聚合过程中始终保持碎片化加密,即使边缘服务器被攻破也无法反推单个车端的原始数据,满足了数据的隐私合规要求;通过异常事件触发的车端紧急自适应与云端经验增量聚合机制,实现了“边用边学”的持续进化能力,车端模型能够快速适应新出现的异常场景;通过空地协同追踪指令,将巡逻车的地面视角与无人航空器的高空视角融合,解决了遮挡与视野受限问题,显著提升了对逃逸嫌疑目标的追踪成功率。
Smart Images

Figure CN122866248A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent traffic control technology, specifically to an air-ground cooperative tracking method and system based on privacy-enhanced federated continuous learning. Background Technology
[0002] With the rapid development of autonomous driving and intelligent connected vehicle technologies, patrol cars and unmanned aerial vehicles (UAVs) are already being used for routine road patrols, traffic violation capture, and suspect target identification. However, existing patrol cars and UAVs still face the following problems when tracking escaped vehicles or suspect targets: First, the vehicle-side model relies on pre-trained fixed weights, and when encountering new abnormal scenarios during patrols, it cannot share the information with other vehicles in a timely manner.
[0003] Second, model updates between multiple patrol vehicles and between vehicles and the cloud involve a large amount of raw perception data, which includes sensitive information such as vehicle location, faces, and license plates. Uploading this data directly to the central server poses a serious risk of privacy leakage.
[0004] Third, patrol vehicles operating from a single location have limited visibility and are prone to losing track of targets when they enter obscured areas or complex road conditions. They also lack an air-ground coordination mechanism that links with unmanned aerial vehicles.
[0005] Some existing research, such as the patent document with publication number CN121857686A, proposes a "learn-as-you-go" mechanism for edge models, but it does not solve the problem of data privacy protection. For example, the patent solution with application number CN202511403891.X involves air-to-ground cooperative tracking, but its solution to data privacy issues also uses existing encrypted transmission schemes. Therefore, how to achieve continuous evolution of patrol vehicle-side models and efficient air-to-ground cooperative tracking while protecting data privacy is a pressing technical challenge in this field. Summary of the Invention
[0006] To address the challenge of achieving continuous evolution of patrol vehicle models and efficient air-ground collaborative tracking while maintaining data privacy in existing technologies, this invention provides an air-ground collaborative tracking method based on privacy-enhanced federated continuous learning. This method can achieve joint continuous tracking between patrol vehicles and unmanned aerial vehicles while ensuring data security. Furthermore, this solution also discloses an air-ground collaborative tracking system based on privacy-enhanced federated continuous learning.
[0007] The technical solution of this invention is as follows: an air-ground cooperative tracking method based on privacy-enhanced federated continuous learning, characterized by comprising the following steps: S1: Obtain the local gradient update data generated by the edge perception model deployed on the patrol vehicle after identifying and tracking the suspect target, and use the Shamir secret sharing mechanism to fragment and encrypt the local gradient update data to generate encrypted gradient fragments. The edge awareness model includes a target detection model and a target tracking model; the target detection model is used to detect suspicious targets in the video stream in real time, and the target tracking model is used to perform cross-frame identity preservation and trajectory tracking on the detected suspicious targets; The fragmentation encryption method is as follows: The local gradient update data of each vehicle is split into N encrypted fragments, denoted as: encrypted gradient fragments; The original gradient update data is reconstructed using any K encrypted gradient fragments, where K is a preset threshold, N is a preset total number of fragments, and 2 ≤ K ≤ N. S2: The encrypted gradient fragments are uploaded to the edge federated learning server through the vehicle-edge collaborative channel. The edge federated learning server performs homomorphic weighted aggregation on the encrypted gradient fragments uploaded by each vehicle to obtain the aggregated gradient, and updates the regional global model based on the aggregated gradient; and synchronizes the regional global model to the vehicle and the cloud center. When the edge federated learning server performs homomorphic weighted aggregation of encrypted gradient fragments uploaded by each vehicle, it adopts a batch fragment aggregation algorithm based on Shamir secret sharing additive homomorphism. It directly completes the multi-vehicle gradient weighted aggregation in the fragment data domain without reconstruction or obtaining plaintext gradient update data from a single vehicle. After aggregation, the aggregated fragments are reconstructed to obtain the global aggregated gradient. S3: Generate a global model based on the global models of all regions in the cloud center; Meanwhile, the cloud center receives and detects abnormal event trigger signals from the vehicle in real time; when the abnormal event trigger signal triggers the incremental training condition, the cloud center uses the abnormal event data to incrementally fine-tune the benchmark model and generate an enhanced global model. S4: The cloud center feeds the enhanced global model back to the vehicle and unmanned aerial vehicle via over-the-air download technology, and generates air-ground collaborative tracking instructions based on the enhanced global model, and schedules associated unmanned aerial vehicles to conduct cross-domain continuous tracking of the suspected target.
[0008] Its further features are: The method for generating encrypted gradient fragments includes the following operations: The specific fragment generation process is as follows: Let the local gradient update data at the vehicle end be a vector g∈R dChoose a large prime number q > max(‖g‖, N), and randomly generate a K-1 degree polynomial over the finite field GF(q): f(x) = g + a1x + a2x 2 +…+a K-1 x K-1 (mod q); Where a1, a 2, … , a K-1 For random coefficients; x, x 2 ,…,x K-1 The x-coordinate is non-zero; Then the i-th encrypted fragment is: fragment i=(x i ,f(x i ), i=1,2,…,N; The target tracking model is based on the DeepSORT algorithm and introduces adaptive switching of motion modes: it automatically switches between uniform linear motion model and uniform turning nonlinear motion model according to the yaw rate of the suspected target, and uses extended Kalman filter to complete nonlinear state prediction; it supports local emergency adaptive fine-tuning triggered by abnormal events, and only performs local small-batch fine-tuning on the fully connected layer of the ReID network, and outputs the experience increment for federated incremental learning. The adaptive switching logic for motion modes includes the following operations; b1: During the actual tracking process, the cloud center monitors the estimated value of the yaw rate ω of the suspect target in real time; When |ω| exceeds the preset threshold δ for three consecutive time steps, execute step b3; Otherwise, proceed to step b2; b2: Execute the uniform linear motion model; s coord t|t-1 =F linear ·s coord t-1|t-1 ; In the formula, s coord t|t-1 Let t be the predicted location of the suspect target in the preceding frame of the video at time t; s coord t-1|t-1 Let s be the motion state vector of the suspected target at time t-1; coord t =[x t ,y t ,θ t ,v t ,ω t ] Tx and y represent planar positions, θ is the heading angle of the suspected target, v is the velocity of the suspected target, and ω is the yaw rate. F linear This represents the state transition matrix of a uniform linear motion model; ; △t is the prediction step size; b3: Execute the uniform speed turning motion model; s coord t|t-1 =F nonlinear (s coord t-1|t-1 ); In the formula, F nonlinear For representing the nonlinear function of the uniform speed turning motion model; ; Furthermore, state prediction is performed based on extended Kalman filtering: Assuming that v and ω remain constant within the prediction step size Δt, for the nonlinear model, the cloud center calculates the Jacobian matrix F. t Used for covariance prediction: ; In the formula, s t-1|t-1 Let be the motion state vector of the suspected target at time t-1; In step S2, the process of updating the global model of the region includes the following steps: Let U be the set of vehicles participating in this round of aggregation, and let ΔW be the local gradient update amount of vehicle u. u Its training sample size is n u The aggregate gradient ΔW calculated by the edge federated learning server is... agg for: △W agg =(∑ u∈U (n u ·△W u )) / (∑ u∈U n u ); The edge federated learning server updates the regional global model at time t+1 according to the asynchronous update rule, obtaining the updated regional global model W'. global (t+1) : W' global (t+1) =W' global (t) -ηw·△W agg ; Where ηw is the federated learning aggregation rate; W' global (t)Represents the global model of the region at time t; In step S3, the abnormal event triggering signals include: daily detection of suspected targets triggering events, vehicle-side model confidence continuously falling below a preset threshold, or new detection tasks triggering events; The incremental fine-tuning process includes: after the vehicle detects an abnormal event that triggers the incremental training condition, it first activates the local emergency adaptive mechanism and uses the currently cached abnormal scene data to perform small-batch rapid fine-tuning of the last two layers of the fully connected layer of the ReID network in the edge perception model; the vehicle calculates the difference in model weights before and after fine-tuning as an experience increment, which is then encrypted by Shamir's secret sharing and uploaded to the edge federated learning server and the cloud center; the cloud center performs weighted fusion of the experience increments uploaded by each vehicle to generate an enhanced global model; The enhanced global model generation process specifically includes the following steps: a1: After the vehicle detects an abnormal event, it reports it to the cloud center through the edge federated learning server; Simultaneously, the bottom convolutional layers are frozen, and the vehicle-side initiates a local emergency adaptive mechanism to perform rapid, small-batch fine-tuning of the last two fully connected layers of the ReID network in the edge perception model. Then, the difference in model weights before and after the repair is calculated, generating an empirical increment ΔW. exp ; △W exp =W car '-W car ; In the formula, W car To fix the previous model weights, W car 'These are the weights for the repaired model; a2: The vehicle will increment the experience by △W exp The data is split into encrypted fragments and uploaded to the edge federated learning server; the edge server aggregates and reconstructs the incremental fragments uploaded by each vehicle and uploads the resulting experience increments to the cloud; a3: The cloud center detects abnormal events in real time based on the abnormal event detection function ε(t); ε(t)=II(Conf(t)<θ conf ) or II(Alarm(t)=1) or II(Event(t)=1); In the formula, II() represents an indicator function, which takes the value of 1 when the condition in parentheses is true, and 0 otherwise when the condition is false; or represents a logical OR operation; Conf(t) is the confidence level of the vehicle-side model in identifying the suspect target at time t; Alarm(t) is the daily detection trigger event for the suspect target issued by the vehicle-side; Event(t) is the trigger event for adding a new detection task. a4: When the output value of ε(t) is 1, it indicates that an abnormal event occurred at time t, and step a5 is executed; Otherwise, when the output value of ε(t) is 0, continue to detect abnormal events in real time based on the abnormal event detection function ε(t); a5: The cloud center is in the global model W global Based on this, the aggregated experience increments are weighted and fused to generate an enhanced global model W. enh ; ; Among them, △W exp (m) This represents the experience increment corresponding to the m-th vehicle terminal; M is the total number of vehicle terminals that reported abnormal events, and η is the cloud aggregation weight factor; The airborne model carried in the unmanned aerial vehicle is a homologous instance of the vehicle-side global model after structural scaling to meet edge computing power constraints. The two share a unified loss gradient space when they are aggregated in federated learning. After completing inference and training batches locally, the airborne model generates local gradient updates, which are then uploaded to the edge federated learning server via Shamir's secret sharing and fragmentation encryption, and subsequently synchronized to the cloud center. The cloud center generates an enhanced global model, which is then fed back to all vehicles and all unmanned aerial vehicles simultaneously via over-the-air download technology.
[0009] An air-ground cooperative tracking system based on privacy-enhanced federated continuous learning is characterized by comprising: a vehicle-side privacy-aware tracking module, an edge federated learning server, a cloud-based incremental evolution module, an air-ground cooperative tracking scheduling module, and an unmanned aerial vehicle cooperative unit; The vehicle-side privacy awareness and tracking module is deployed on the vehicle side and is used to acquire local gradient update data generated by the edge perception model after identifying and tracking the suspect target. The local gradient update data is fragmented and encrypted using the Shamir secret sharing mechanism to generate encrypted gradient fragments. The vehicle-side privacy awareness and tracking module includes: a target detection module and a target tracking module; the target detection module embeds a target detection model, which uses current parameters to detect targets in the video stream in real time; the target tracking module embeds a target tracking model, which continuously detects and tracks suspected targets and reports the target location in real time. The edge federated learning server is deployed in the roadside equipment and communicates with the vehicle-side privacy awareness and tracking module through the vehicle-edge collaborative channel. It is used to receive encrypted gradient fragments uploaded by each vehicle, perform homomorphic weighted aggregation on each encrypted gradient fragment to obtain an aggregated gradient, update the regional global model based on the aggregated gradient, and transmit the regional global model to the vehicle and the cloud center. The cloud-based incremental evolution module and the air-to-ground collaborative tracking and scheduling module are deployed in the cloud center and communicate with the edge federated learning server. They are used to receive abnormal event trigger signals reported by vehicle-side and roadside equipment. After the abnormal event detector determines the signal, the incremental training trigger starts the incremental training process. The transfer learning fine-tuning engine performs incremental fine-tuning on the baseline model to generate an enhanced global model. The air-to-ground collaborative tracking and scheduling module feeds the global model back to each patrol vehicle and generates air-to-ground collaborative tracking control commands to schedule unmanned aerial vehicles (UAVs) to track suspected targets. The UAVs and each patrol vehicle interact with each other through a two-way communication link, transmitting high-altitude target images and location information back to the vehicle in real time. The vehicle reports the target tracking status to the UAV. The unmanned aerial vehicle (UAV) collaborative unit is used to automatically take off after receiving scheduling instructions. The UAV collaborative unit adaptively selects the access method according to the current communication status: when it is within the coverage area of the edge federated learning server, it accesses the edge federated learning server and participates in the three-level federated learning architecture; when it is outside the coverage area of the edge server, it communicates directly with the cloud center and participates in the two-level federated learning architecture. The UAV collaborative unit performs aerial tracking based on the target position and predicted trajectory reported by the vehicle and transmits back images of the suspected target from a high-altitude perspective. When working in the two-level federated learning architecture, the UAV still performs Shamir secret sharing fragment encryption locally, and the encrypted fragments are directly uploaded to the cloud center, where the cloud center completes the fragment aggregation and reconstruction.
[0010] The air-ground cooperative tracking method and system based on privacy-enhanced federated continuous learning provided in this application, through Shamir's secret sharing of federated learning, ensures that sensitive gradient data on the vehicle side remains fragmented and encrypted throughout the transmission and aggregation process. Even if the edge server is compromised, it is impossible to reverse engineer the original data of a single vehicle side, thus meeting the data privacy compliance requirements. Through the vehicle-side emergency adaptation triggered by abnormal events and the cloud-based incremental experience aggregation mechanism, it achieves the ability to continuously evolve by "learning on the go," enabling the vehicle-side model to quickly adapt to newly emerging abnormal scenarios. Through air-ground cooperative tracking commands, the ground view of the patrol vehicle and the high-altitude view of the unmanned aerial vehicle are fused, solving the problems of occlusion and limited field of view, and significantly improving the tracking success rate of escape suspects. Attached Figure Description
[0011] Figure 1 A flowchart for a privacy-enhanced federated continuous learning-based air-to-ground cooperative tracking method; Figure 2 A time sequence diagram of incremental fine-tuning and model feedback in a privacy-enhanced federated continuous learning-based air-ground cooperative tracking method; Figure 3 This is a block diagram of an air-ground collaborative tracking system based on privacy-enhanced federated continuous learning. Detailed Implementation
[0012] like Figure 3 As shown, this application constructs an air-ground cooperative tracking system based on privacy-enhanced federated continuous learning, which includes: a vehicle-side privacy-aware tracking module, an edge federated learning server, a cloud-based incremental evolution module, and an air-ground cooperative tracking scheduling module.
[0013] This application employs a three-tiered hierarchical architecture based on an edge federated learning mechanism to construct an air-ground collaborative tracking system. The system comprises a cloud center, edge federated learning servers, and terminals, including unmanned aerial vehicles (UAVs) and vehicles. The edge federated learning servers are positioned between the cloud and the terminals, forming a three-tiered architecture. Specifically, they are installed in existing roadside equipment within urban traffic systems, distributed across different urban areas, and can communicate in real-time with UAVs and patrol vehicles within their respective regions.
[0014] The vehicle-mounted privacy-aware tracking module, deployed in patrol vehicles, acquires local gradient update data generated by the edge-aware model after identifying and tracking suspected targets. It then uses the Shamir secret sharing mechanism to fragment and encrypt this local gradient update data, generating encrypted gradient fragments. In practical application, the patrol vehicles are pre-registered and, while tracking targets, establish communication connections with the edge federated learning servers in different areas of the city. All edge federated learning servers simultaneously communicate with the cloud center.
[0015] The cloud-based incremental evolution module and the air-to-ground collaborative tracking and scheduling module are located in the cloud center. These modules issue operational-level task commands to patrol vehicles and unmanned aerial vehicles via control instructions.
[0016] The vehicle-side privacy-aware tracking module includes a target detection module and a target tracking module. The target detection module embeds a lightweight target detection model. In this embodiment, the target detection model adopts a YOLOv5s and MobileNetV2 fusion architecture. Based on the lightweight target detection model, targets are identified. Detected targets include one or more of the following: motor vehicles, non-motor vehicles, pedestrians, static obstacles, dynamic obstacles, traffic signs, traffic lights, and environmental targets related to driving decisions. Specific targets are set according to actual conditions. The construction and training methods of the lightweight target detection model are based on existing technologies. The output of the lightweight target detection model includes: suspected target category, spatial location, confidence level, and motion state.
[0017] The target tracking module embeds a target tracking model, which can be implemented based on existing target tracking algorithms, such as the standard DeepSORT algorithm.
[0018] The DeepSORT (Deep Simple Online and Realtime Tracking) algorithm is a multi-target tracking algorithm proposed by Nicolai Wojke et al. in 2017. Its core idea is to integrate deep appearance features (ReID features) into the SORT algorithm to solve the ID switching problem when a target is occluded for a long time. The standard DeepSORT algorithm includes: state estimation (Kalman filter), data association (Hungarian algorithm), cascaded matching, and trajectory lifecycle management.
[0019] In this embodiment, to better suit the application scenario of air-to-ground cooperative tracking, the standard DeepSORT algorithm has been improved, and an improved DeepSORT algorithm has been designed and embedded into the target tracking model. The specific improvements include the following two points; the rest are the same as the standard DeepSORT algorithm: Improvement 1: Adaptive switching of motion models: Based on the yaw rate of the suspected target, the system automatically switches between a uniform linear motion model and a uniform-rate turning nonlinear motion model, and uses an extended Kalman filter to perform nonlinear state prediction. Specifically, the standard DeepSORT uniform linear motion model is extended to an adaptive switching mechanism. When the target yaw rate |ω| is less than a preset threshold δ for a preset continuous time period, a linear Kalman filter is used. When |ω| is greater than or equal to the preset threshold δ for a preset continuous time period, the system automatically switches to a uniform-rate turning nonlinear model and performs state prediction through an extended Kalman filter to adapt to various maneuvering modes of the suspected target vehicle.
[0020] Improvement 2: Support for incremental fine-tuning triggered by abnormal events: Local mini-batch fine-tuning is performed only on the fully connected layers of the ReID network, and the resulting experience increment is used for federated incremental learning. Specifically, the last two layers of the fully connected layers of the ReID network deployed on the vehicle are configured to support local emergency adaptive fine-tuning on the vehicle (3-8 mini-batches), and the weight difference ΔW after fine-tuning is... exp The data is secretly shared and encrypted by Shamir and uploaded to the cloud. After being aggregated in the cloud, it is distributed to all vehicles in the entire domain via over-the-air (OTA) download technology, enabling the continuous evolution of the tracking model.
[0021] When an abnormal event is detected, the vehicle-side privacy awareness and tracking module sends an abnormal event trigger signal to the cloud layer through the edge federated learning server.
[0022] The edge federated learning server comprises a gradient verification unit, an aggregation and reconstruction engine, and an asynchronous global model updater. Deployed in roadside equipment, the edge federated learning server communicates with the vehicle-side privacy-aware tracking module via a vehicle-edge collaborative channel. The gradient verification unit receives encrypted gradient fragments uploaded by each vehicle, performs integrity checks and outlier removal on each fragment before aggregation and reconstruction, and removes the removed fragments. The aggregation and reconstruction engine performs homomorphic weighted aggregation on the Shamir encrypted gradient fragments uploaded by each vehicle to obtain the aggregated gradient, and the asynchronous global model updater asynchronously updates the regional global model based on the aggregated gradient. The edge federated learning server transmits the updated regional global model to both the vehicle-side devices and the cloud center. At the cloud center, a global model is generated based on all regional global models.
[0023] The cloud-based incremental evolution module and the air-to-ground collaborative tracking and scheduling module are deployed in the cloud center and communicate with the edge federated learning server. They are used to receive abnormal event trigger signals reported by the vehicle. After the abnormal event detector determines the signal, the incremental training trigger starts the incremental training process. The transfer learning fine-tuning engine performs incremental fine-tuning on the baseline model to generate an enhanced global model. The air-to-ground collaborative tracking and scheduling module feeds the global model back to each patrol vehicle and generates air-to-ground collaborative tracking control commands to schedule unmanned aerial vehicles (UAVs) to track suspected targets. The UAVs and each patrol vehicle interact with each other through a two-way communication link, transmitting high-altitude target images and location information back to the vehicle in real time. The vehicle reports the target tracking status to the UAV.
[0024] The unmanned aerial vehicle (UAV) receives air-to-ground coordinated tracking and control commands from the cloud layer and interacts with patrol vehicles at the vehicle layer via a two-way communication link: the UAV transmits high-altitude target images and positions in real time, and the vehicle reports the target tracking status to the UAV. The UAV coordination unit, deployed within the UAV, automatically takes off after receiving dispatch commands, performs aerial tracking based on the target positions and predicted trajectories reported by the vehicle, and transmits high-altitude images of suspected targets.
[0025] Inter-layer interaction: The vehicle-side layer connects upwards to the edge layer via the C-V2X communication protocol, transmitting encrypted gradient fragments (Shamir fragments); the edge layer connects upwards to the cloud layer, transmitting reconstructed gradient update data and anomaly event information; the cloud layer connects downwards to the unmanned aerial vehicle (UAV) via network communication, transmitting air-to-ground collaborative tracking and control commands and onboard model OTA update packages; the UAV and the patrol vehicles at the vehicle-side layer exchange data via bidirectional communication links, transmitting high-altitude target images and location information, target tracking status, etc. Solid / dashed arrows indicate the flow of data, control commands, and anomaly event trigger signals between layers.
[0026] The UAV collaborative unit adaptively selects the access method based on the current communication status: when within the coverage area of the edge federated learning server, it accesses the edge federated learning server through the vehicle-edge collaborative channel to participate in the three-level federated learning architecture; when the UAV is outside the coverage area of all edge servers, it communicates directly with the cloud center through the 5G core network to participate in the two-level federated learning architecture. The UAV collaborative unit performs aerial tracking based on the target position and predicted trajectory reported by the vehicle and transmits back images of the suspected target from a high-altitude perspective. When operating in the two-level federated learning architecture, the UAV still performs Shamir secret sharing fragment encryption locally, and the encrypted fragments are directly uploaded to the cloud center for fragment aggregation and reconstruction.
[0027] The air-ground cooperative tracking method implemented based on the above-mentioned privacy-enhanced federated continuous learning-based air-ground cooperative tracking system, such as... Figure 1 As shown, it includes the following steps.
[0028] S1: Obtain the local gradient update data generated by the edge perception model deployed on the patrol vehicle after identifying and tracking the suspect target, and use the Shamir secret sharing mechanism to fragment and encrypt the local gradient update data to generate encrypted gradient fragments; the fragmentation encryption method is: split the local gradient update data of each vehicle into N encrypted fragments, denoted as: encrypted gradient fragments.
[0029] The original gradient update data is reconstructed using any K encrypted gradient fragments, where K is a preset threshold, N is a preset total number of fragments, and 2 ≤ K ≤ N.
[0030] The specific values of K and N are set and adjusted according to the actual situation. In this embodiment, N=5 and K=3.
[0031] During continuous patrols, the lightweight target detection model uses current parameters to detect and track targets such as vehicles, pedestrians, and license plates in the video stream in real time. After completing a local training batch, the gradient update amount ΔW relative to the current global model is calculated. To protect the sensitive location and identity information contained in this gradient update amount, the vehicle-side calls the Shamir secret sharing library to split ΔW into N fragments.
[0032] The lightweight object detection model in this application adopts a YOLOv5s + MobileNetV2 architecture. MobileNetV2 serves as the lightweight backbone feature extraction network, replacing the native CSPDarknet53 of YOLOv5s, and is responsible for basic feature extraction from the input video frames. YOLOv5s retains the native neck PANet multi-scale feature fusion network and head detection branch, and receives the multi-scale feature maps output by MobileNetV2 to complete cross-layer fusion enhancement of features. Finally, at three different scales (far, mid, and near), it completes the classification and bounding box regression of target types such as vehicles, pedestrians, and license plates, respectively, and outputs accurate detection results.
[0033] The object detection subnetwork adopts the YOLOv5s architecture, and its trainable parameters include: feature pyramid network weights and detector head weights; the feature extraction subnetwork adopts the MobileNetV2 architecture, and its trainable parameters include: depthwise separable convolutional layer weights and fully connected layer weights. The local gradient update data ΔW is the gradient vector of all the above trainable parameters in the current local training batch.
[0034] To achieve target tracking, the vehicle also maintains the motion state vector of the suspected target in real time. The specific motion state vector includes: the target's planar position coordinates (x, y) in the global physical coordinate system, heading angle θ, velocity v, and yaw rate ω. The motion state vector is used for trajectory prediction in air-ground cooperative tracking commands, but does not participate in the gradient update of model weights.
[0035] During target tracking, the vehicle-mounted system predicts the motion state vector of the suspected target in real time based on the extended Kalman filter algorithm. The motion state vector of the suspected target at time t is defined as s. coord t =[x t ,y t ,θ t ,v t ,ω t ] T ;coord represents the global physical coordinate space, and the motion state vector s coord t Used for trajectory prediction and path planning in air-to-ground cooperative tracking commands, with motion state vector s coord t As a real-time estimate, it is an intermediate variable in the tracking algorithm, rather than a network weight parameter in the edge-aware model, and therefore does not participate in the gradient update process of federated learning.
[0036] The vehicle-mounted system also features a camera that captures real-time video streams, which are then fed into a lightweight object detection model for object detection. The accompanying vehicle-mounted system also requires camera calibration parameters to convert image pixel coordinates into actual spatial orientation. These calibration parameters include: extrinsic parameters, intrinsic parameter matrix (IR), and lens distortion coefficients. Parameters in the intrinsic parameter matrix (IR) include: focal length (f... x ,f y ), coordinates of the optical axis center point (c x ,c y Lens distortion coefficients are used to describe the optical imaging characteristics of the camera, including: radial distortion parameters k1, k2, and k3; and tangential distortion coefficients p1 and p2. Extrinsic parameters include a rotation matrix R (3×3) and a translation vector t (3×1), used to describe the camera's mounting attitude and position in the vehicle coordinate system, and to convert the target orientation in the camera coordinate system to the target azimuth in the vehicle coordinate system. The specific methods for obtaining calibration parameters can be implemented based on existing technologies.
[0037] The camera calibration parameters are determined once during vehicle factory calibration, and the specific measurement method is based on existing technology, remaining fixed during patrol. The calibration parameters are only used in the preprocessing stage of perception data, mapping the original image coordinates to spatial orientation. They do not participate in the forward inference or backpropagation calculation of the neural network, and therefore do not fall within the scope of learnable parameters of the edge perception model, nor are they included in the gradient update data △W of federated learning.
[0038] The target tracking model uses the standard Kalman filter from the DeepSORT algorithm for inter-frame state prediction. track t-1 =[x,y,a,h,x',y',a',h'] T Where x and y are the coordinates of the center of the target detection box in the image pixel coordinate system; a is the aspect ratio of the detection box: a = width / height; h is the height of the detection box; and x', y', a', and h' are the first derivatives of the above variables, i.e., the rate of change. The expected position of the target in the current frame is predicted when a missing element is detected using a standard Kalman filter.
[0039] It needs to be clarified that the state vector s in this method track t-1 Acting on the image plane, it is used to maintain the consistency of the target's identity across consecutive frames in the target tracking model, achieving image identity locking; while the global physical coordinate state vector and motion state vector s coord t The air-ground collaborative physical path deduction used in the air-ground collaborative tracking and scheduling module belongs to different technical levels. The two establish a data transmission relationship through the camera intrinsic and extrinsic parameter matrix and work together, but their mathematical definitions and optimization objectives do not interfere with each other.
[0040] Specifically, the image coordinates (x, y) and their derivatives output by the tracking layer filter are restored to the direction vector in the camera coordinate system through the camera intrinsic parameter matrix IR, and then converted into the target orientation in the vehicle coordinate system by combining the extrinsic parameters R, t. Next, the vehicle's own pose obtained by the onboard positioning module (GPS / IMU) is fused to calculate the target's physical position (X, Y) in the global coordinate system, and the heading angle θ, velocity v, and yaw rate ω are calculated based on the rate of position change, ultimately forming the planning layer state vector s. coord .
[0041] The vehicle-side privacy-aware tracking module employs a processing architecture that decouples detection and tracking but tightly couples data. Specifically, the lightweight object detection model (YOLOv5s + MobileNetV2) performs end-to-end inference on each frame of the image, outputting the target's bounding box position Z. t The bounding box, along with the category confidence score, serves as the observation input to the target tracking model.
[0042] This method implements a target tracking model based on an improved DeepSORT, where the Kalman filter utilizes the state vector s from the previous frame. track t-1 =[x,y,a,h,x',y',a',h'] T Predict the identity marker of the current frame using the target location s track t|t-1 The covariance matrix is then calculated. Subsequently, the Hungarian algorithm in the target tracking model combines the Mahalanobis distance (motion information) and cosine distance (appearance feature information) between the detection box and the predicted trajectory to construct a cost matrix and perform optimal matching.
[0043] When a match is successful, the Kalman filter updates and corrects the predicted state using the detection box, outputting the precise state s of the current frame. track t During this process, the noise covariance matrix is dynamically adjusted to adapt to the target's maneuvering changes. When a new target detection box cannot match any existing trajectory, the target tracking module initializes a new trajectory and marks it as "temporary". When an existing trajectory fails to obtain a detection match for multiple consecutive frames, the trajectory is determined to be lost and deleted. This closed-loop architecture ensures that the vehicle can maintain stable tracking and identity preservation of suspected targets in real time in complex urban road environments.
[0044] It is important to emphasize that the entire detection and tracking loop described above is completed locally on the vehicle. The target tracking model outputs a motion state vector of the suspect target with a unique identifier (ID). This result is transformed into global physical coordinates and used for subsequent air-ground coordinated scheduling. The network parameters related to target appearance feature extraction during the tracking process only participate in cloud incremental fine-tuning and OTA feedback when abnormal events are triggered.
[0045] The specific process for generating encrypted fragments is as follows: Let the local gradient update data at the vehicle end be a vector g∈R d Choose a large prime number q > max(‖g‖, N), and randomly generate a K-1 degree polynomial over the finite field GF(q): f(x) = g + a1x + a2x 2 +…+a K-1 x K-1 (mod q); Where a1, a 2, … , a K-1 q is a random coefficient, ||g|| is the magnitude of vector g, and N is the preset total number of fragments, where N is less than or equal to q-1.
[0046] N distinct non-zero x-coordinates are pre-selected: x1, x2, ..., x N Substituting these non-zero x-coordinates into the polynomial f(x), we can calculate the corresponding function value f(x). i ), each group (x i f(x) is a cryptographic gradient fragment; In the polynomial f(x), a modulo q operation is performed on each coefficient after the polynomial operation, which restricts the values of all coefficients to the integer range of [0, q-1] to avoid numerical overflow. At the same time, the operation of the entire polynomial is carried out under the rules of the finite field GF(q), which satisfies the cryptographic operation constraints of secret sharing.
[0047] Calculate the corresponding f(x), then combine x and f(x) into a set of numbers as the generated encryption fragment; then the i-th encryption fragment is: fragment i=(x i ,f(x i )), i=1,2,…,N.
[0048] Any single encrypted fragment is a random number pair (x i ,f(x iThe data itself contains no readable information and cannot deduce the constant term g. Only after collecting at least K fragments can f(0) = g, i.e., the original gradient update ΔW, be uniquely reconstructed through Lagrange interpolation. Less than K fragments cannot obtain any information about g. This ensures that even if some fragments are intercepted during transmission, attackers cannot recover the original gradient data.
[0049] S2: The encrypted gradient fragments are uploaded to the edge federated learning server through the vehicle-edge collaborative channel. The edge federated learning server performs homomorphic weighted aggregation on the encrypted gradient fragments uploaded by each vehicle to obtain the aggregated gradient, and asynchronously updates the regional global model based on the aggregated gradient. At the same time, the edge federated learning server performs edge-cloud collaborative evaluation on the updated regional global model to achieve synchronous updates between the edge federated learning server and the cloud center.
[0050] When the edge federated learning server performs homomorphic weighted aggregation of encrypted gradient fragments uploaded by each vehicle, it adopts a batch fragment aggregation algorithm to complete global gradient aggregation without decrypting the gradient data of individual vehicles.
[0051] Specifically, the vehicle sends fragments to the roadside edge federated learning server covering the area via the C-V2X communication protocol. The edge server collects fragments from multiple patrol vehicles and, without accessing the plaintext gradients of individual vehicles, utilizes the additive homomorphism secretly shared by Shamir to directly perform weighted averaging at the fragment level (without needing to recover the individual vehicle gradient update amount ΔW), without reconstructing or obtaining the plaintext gradient update data of individual vehicles. After aggregation, the aggregated fragments are reconstructed to obtain the global aggregated gradient, and the global model maintained by the edge side is updated asynchronously.
[0052] Shamir's secret sharing has additive homomorphic properties, where the fragments of multiple secrets can be directly added to their corresponding position values. The fragments after addition are equivalent to the original secrets being added and then fragmented. Therefore, weighted summation can be performed directly at the fragment level. The aggregated fragment set is then subjected to Lagrange interpolation, which only restores the overall gradient after aggregation and does not restore any individual vehicle's local gradient.
[0053] Let U be the set of vehicles participating in this round of aggregation, and let ΔW be the local gradient update amount of vehicle u. u Its training sample size is n u Then the aggregate gradient ΔW calculated by the edge server agg for: △W agg =(∑ u∈U (n u ·△W u )) / (∑ u∈U n u ); Subsequently, the edge updates the global model according to the asynchronous update rules: W global (t+1) =W global (t) - ηw·△W agg ; Where ηw is the federated learning aggregation rate; W global (t) This represents the edge global model at time t.
[0054] △W agg It is a comprehensive gradient obtained by weighting the gradient update data reconstructed by the edge server for each vehicle according to the sample size, and is directly used to asynchronously update the global model of the region.
[0055] Throughout the process, the edge server cannot directly see the raw perception data of a single vehicle; it can only see the gradient update data, thus achieving privacy protection of "the model moves while the data remains stationary."
[0056] S3: In the cloud center, an anomaly event detector monitors in real time whether the incremental training condition is triggered; when the anomaly event trigger signal triggers the incremental training condition, the incremental training trigger uses the anomaly event data to incrementally fine-tune the baseline model and generate an enhanced global model.
[0057] In practice, the cloud center uses an anomaly detector to monitor anomaly events reported by the vehicle in real time. Once an anomaly occurs, incremental training is triggered. The anomaly events included in the incremental training conditions are set according to the actual situation. In this embodiment, the anomaly event trigger signals include: routine detection of suspected targets, vehicle model confidence falling below a preset threshold, or a new detection task.
[0058] Among them, routine detection of suspected targets triggers events, such as: a city originally had no tunnels, and the original target detection scenarios included rainy days, foggy days, and sunny days; but a tunnel was added in a certain administrative area of the city, and the patrol cars in that area added a new "inside the tunnel" detection scenario when conducting routine detection of suspected targets, which led to the parameter update of the target detection model; the parameters related to this scenario update will also be synchronized to the training vehicles in other areas. New detection task triggering events refer to new tasks added outside of the existing daily operational tasks. The daily operational tasks of patrol vehicles mainly involve automatically detecting illegal and irregular events such as illegal parking, suspicious vehicles, and suspects during patrol duties. However, patrol vehicles also frequently receive ad-hoc assistance tasks, and the specific impact of each new detection task on the detection model varies depending on the actual situation. For example, assisting in a special rectification campaign targeting large vehicles such as trucks and dump trucks will add new detection targets to the corresponding target detection model.
[0059] The incremental fine-tuning process includes: after the vehicle detects an abnormal event that triggers the incremental training condition, it first activates the local emergency adaptive mechanism and uses the currently cached abnormal scene data to perform small-batch rapid fine-tuning of the last two fully connected layers of the ReID network in the edge perception model; the vehicle calculates the difference in model weights before and after fine-tuning as the experience increment, which is then uploaded to the edge server after being encrypted by Shamir's secret sharing; the cloud performs weighted fusion of the experience increments uploaded by each vehicle to generate an enhanced global model.
[0060] For example, in a scenario where an abnormal event is detected by the vehicle and incremental training conditions are triggered, such as when the confidence level of a feature of a certain escaped vehicle drops below 0.6, the vehicle first activates the local emergency adaptive mechanism: using the currently cached abnormal scene data, such as the last 10 frames of images, the vehicle performs 3 to 8 small batches of rapid fine-tuning on the last two fully connected layers of the ReID network in the edge perception model, freezes the bottom convolutional layers, and corrects the parameters with an extremely low learning rate to ensure that the current patrol car's tracking of the suspect target is not interrupted due to network latency.
[0061] In the anomaly detector, the anomaly detection function ε(t) is defined as: ε(t)=II(Conf(t)<θ conf ) or II(Alarm(t)=1) or II(Event (t)=1); In the formula, II() represents an indicator function, which takes the value 1 when the condition within the parentheses is true, and 0 otherwise when the condition is false. In practical applications, the indicator function II() is implemented based on a judge to answer "yes" or "no". For example, II(Conf<0.6) means that if the model's recognition confidence Conf is lower than the preset threshold of 60% at this moment, this formula is equal to 1; otherwise, it is equal to 0.
[0062] "OR" represents the logical "OR" operation. That is, whenever the identification confidence level falls below a preset threshold, or a suspected target is reported from the vehicle and triggers a routine detection event or a new detection task, the anomaly detection function ε(t) outputs 1, triggering the subsequent incremental training process. The multiple indicator functions in the anomaly detection function ε(t) are independent of each other; any one of them can trigger the function without mutual exclusion. The value of ε(t) is only 0 when all indicator functions output 0.
[0063] Conf(t) represents the confidence level of the vehicle-side model in identifying the suspect target at time t; θ confThe preset threshold is set to 0.6 in this embodiment; Alarm(t) is the daily detection trigger event for suspected targets issued by the vehicle. Event(t)=1 is the trigger event for a new detection task. When ε(t)=1, the incremental training condition is triggered.
[0064] Suppose that at a certain time t, the unmanned patrol vehicle encounters the following three situations: Recognition Confidence (Conf): The current confidence in recognizing the target has dropped to Conf(t) = 0.45, corresponding to the preset threshold θ. conf =0.6.
[0065] New detection scenario (Alarm): The vehicle-side target detection model does not issue a trigger event, and the corresponding alarm is Alarm(t)=0.
[0066] New detection task trigger event (Event): No new detection task is added, and the corresponding event value is Event(t)=0.
[0067] The first step is to calculate the three indicator functions (determiners): Item 1: II (0.45 < 0.6), the condition in parentheses is true, so the output is 1.
[0068] Item 2: II (0=1), if the condition in parentheses is not true, then output 0.
[0069] Item 3: II (0=1), if the condition in parentheses is not true, then output 0.
[0070] Step 2: Perform a logical OR calculation; The formula becomes: ε(t) = 1 or 0 or 0.
[0071] Since as long as there is a "1", the result of the logical OR is "true", so ε(t) = 1.
[0072] After completing the local adaptive repair, the vehicle does not upload the repaired complete model or the original abnormal image. Instead, it calculates the difference in model weights before and after repair and generates an empirical increment △W. exp ; △W exp =W car '-W car ; In the formula, W car To fix the previous model weights, W car 'These are the weights for the repaired model; The increment vector △W exp It only includes the direction and magnitude of parameter corrections for this abnormal scene, and does not contain any sensitive information that could be used to deduce the original image.
[0073] Subsequently, the vehicle-side, based on the Shamir secret sharing mechanism in step S1, will increment the experience by △W. exp The data is split into encrypted fragments and uploaded to the edge federated learning server; the edge server aggregates and reconstructs the incremental fragments uploaded by each vehicle. The cloud center is located in the global model W. global Based on this, the aggregated experience increments are weighted and fused to generate an enhanced global model W. enh : ; Among them, △W exp (m) This represents the experience increment corresponding to the m-th vehicle terminal; M is the total number of vehicle terminals that report abnormal events. In specific applications, the reporting source device of the abnormal event is recorded and distinguished separately in the data field; η is the cloud aggregation weight factor, which is used to control the degree of influence of a single abnormal event on the global model and prevent overfitting.
[0074] Ultimately, the cloud center will use this enhanced global model W enh Over-the-air (OTA) data transfer technology feeds back to all patrol vehicles across the entire area. Through the closed-loop mechanism of "vehicle-side emergency adaptation → extraction of incremental experience → cloud aggregation and sharing," herd immunity against single-point anomalies is achieved, ensuring both the real-time nature of vehicle tracking and protecting the privacy and security of relevant sensing data.
[0075] S4: The cloud center feeds back the enhanced global model to each patrol vehicle and unmanned aerial vehicle via over-the-air download technology, and generates air-ground collaborative tracking instructions based on the enhanced global model, scheduling associated unmanned aerial vehicles to conduct cross-domain continuous tracking of suspected targets.
[0076] The air-to-ground collaborative tracking commands are set according to actual needs. In this embodiment, they include: the current spatial coordinates of the suspect target, motion trajectory prediction, patrol vehicle tracking path planning, unmanned aerial vehicle takeoff point and cruise route, and channel allocation strategy for vehicle-to-air data link.
[0077] like Figure 2 The figure shown is a time sequence diagram of incremental experience aggregation and air-ground collaborative feedback under abnormal triggering.
[0078] The participants in the system include: vehicles, edge federated learning servers, cloud centers, unmanned aerial vehicles, and other patrol vehicles. (Time from top to bottom) Daily federated learning phase: Vehicle → Edge server: Upload encrypted gradient fragments (Shamir fragments 1~N, periodically); Inside the edge server: Aggregate and reconstruct the gradients of each vehicle and update the regional global model; Edge server → Vehicle: Distribute the updated regional global model (asynchronously) and send it to the cloud center; Cloud center: Update the global model.
[0079] Anomaly Triggering and Vehicle-Side Emergency Adaptation Phase: Implemented on the vehicle side, an abnormal event is detected, triggering incremental training conditions (identification confidence below a preset threshold / network alarm / roadside event); local emergency adaptation is initiated, the last two fully connected layers of the ReID network are fine-tuned (3-8 mini-batches), the bottom convolutional layers are frozen, and parameters are corrected with an extremely low learning rate; empirical increments are calculated.
[0080] Experience increment encryption upload stage: Vehicle end → Edge server: Upload encrypted experience increment fragments (Shamir fragments 1~N); Inside the edge server: Aggregate and reconstruct the experience increment fragments of each vehicle.
[0081] Cloud aggregation phase: Edge server → Cloud: Forward the reconstructed experience increments of each vehicle; Inside the cloud: Weighted aggregation of the experience increments of each vehicle to generate an enhanced global model.
[0082] OTA Synchronous Feedback Phase: Cloud → Edge Server: Deploy enhanced global model; Edge Server → Vehicle: Deploy OTA upgrade package, hot-swap model; Edge Server → Other Patrol Vehicles: Synchronously deploy OTA upgrade package, hot-swap model, achieving "one place encounters an anomaly, the whole system evolves." Cloud → Unmanned Aerial Vehicle: Send air-to-ground collaborative tracking commands and onboard model OTA update packages.
[0083] Air-ground collaborative tracking execution phase: The unmanned aerial vehicle transmits high-altitude target images and locations back to the vehicle in real time, assisting the vehicle in completing cross-domain continuous tracking.
[0084] The unmanned aerial vehicle (UAV) is equipped with an onboard lightweight perception model, including a target detection sub-network and a motion tracking sub-network. The target detection sub-network adopts the same YOLOv5s and MobileNetV2 fusion architecture as the vehicle-mounted target detection model, and its network depth and input resolution are lightweighted and pruned according to the computing power limitations of the onboard computing platform. The motion tracking sub-network employs a detection-based lightweight multi-target tracking algorithm, which can be implemented using existing target tracking algorithms such as SORT and DeepSort. The number of parameters in the UAV's onboard model is on a lower scale than that of the ground patrol vehicle model, to accommodate the power consumption and computing power constraints of the onboard computing platform.
[0085] The airborne model carried in the unmanned aerial vehicle in this application is a homologous instance of the vehicle-side global model after structural scaling to meet edge computing power constraints. The two share a unified loss gradient space during federated learning aggregation and complete knowledge synchronization through parameter mapping algorithm during OTA feedback, ensuring the consistency of the perception capabilities of the entire system, rather than simply copying the same weight file.
[0086] The cloud center and the unmanned aerial vehicle (UAV) exchange data bidirectionally via 5G or C-V2X data links. Downlink data includes: air-to-ground collaborative tracking commands (including the suspected target's spatial coordinates, trajectory prediction, cruise route, and takeoff point), onboard model parameter update packages (issued via OTA), and gimbal and flight control parameters. Uplink data includes: real-time transmission of the suspected target's position and motion status from the UAV, de-identified target appearance feature vectors (not the original images), onboard model local gradient update data (uploaded after being fragmented and encrypted via Shamir's secret sharing), and the UAV's own status information, such as remaining battery power, position, flight attitude, and communication quality.
[0087] The vehicle-side global model and the UAV-borne model belong to a homogeneous but heterogeneous twin model system. Both use the same MobileNetV2 backbone network and adopt a unified feature alignment strategy to participate in federated learning aggregation. The UAV model inherits the initial weights from the vehicle-side model through knowledge distillation to ensure knowledge homogeneity. Structurally, the UAV model is scaled down three times based on the vehicle-side model: (1) width scaling, reducing the width multiplier of MobileNetV2 from 1.0 to 0.5 and halving the number of channels per layer; (2) depth scaling, reducing the repetition of the Bottleneck module and reducing the number of network layers; (3) head simplification, as the target scale is relatively concentrated under the high-altitude view of the UAV, the three heads are simplified to two. Through the above scaling, the number of parameters of the UAV-borne model is reduced from approximately 7.0 × 10 6 Reduced to approximately 2.0 × 10 6 The reduction is approximately 70%, enabling it to run in real-time on an airborne computing platform. Despite their heterogeneous structures, both share a unified loss gradient space during federated learning aggregation, and the gradient update amount ΔW calculated by the unmanned aerial vehicle model is... uav Gradient update amount △W of the vehicle-side model car By jointly participating in cloud-based weighted aggregation, we can ensure that the perception capabilities of vehicles and unmanned aerial vehicles evolve in sync, rather than simply copying weighted files.
[0088] The cloud-based enhanced global model is simultaneously fed back to all patrol vehicles and all unmanned aerial vehicles via OTA (Over-The-Air) download technology, enabling the synchronous evolution of the integrated air-ground perception model. The camera calibration parameters (intrinsic and extrinsic parameter matrices) of the unmanned aerial vehicles are pre-calibrated at the factory and remain unchanged during flight, and are not included in the scope of federated learning gradient updates.
[0089] Specifically, the cloud packages the enhanced global model into an OTA upgrade package and distributes it to all online patrol vehicles via the 5G core network. Upon receiving the package, the vehicle-side device hot-swaps the current inference model, achieving instant capability upgrades. Simultaneously, based on the reported location and trajectory predictions of the suspected target, the cloud generates air-ground collaborative tracking instructions: planning the drone's following path, designating the drone's takeoff and continuously locking onto the target from the air, and when the target enters a building-obscured area, the drone guides the vehicle-side device to predict the exit on the other side of the obstruction using a high-altitude perspective, achieving seamless tracking.
[0090] The target motion prediction performed by the vehicle-side target tracking model uses a Kalman filter-based model: Let the motion state vector of the suspected target at time t be: s coord t =[x t ,y t ,θ t ,v t ,ω t ] T , representing planar position, heading angle, velocity, and yaw rate, respectively. The system automatically switches prediction models based on the absolute value of the currently estimated yaw rate: When three consecutive time steps satisfy |ω t-1 When |<δ, the target position s of the current frame coord t|t-1 =F linear ·s coord t-1|t-1 ; When three consecutive time steps satisfy |ω t-1 When |≥δ, the target position s of the current frame coord t|t-1 =F nonlinear (s coord t-1|t-1 ); Where δ is a preset angular velocity threshold, used to distinguish between approximate linear motion and obvious turning motion; according to the experimental range of δ, it is set to 0.05~0.1 rad / s, and in this embodiment δ=0.08.
[0091] F linear F is the state transition matrix; nonlinear Let be a nonlinear function representing a uniform-speed turning motion model.
[0092] (1) Linear model (uniform linear motion): When three consecutive time steps satisfy |ω|<δ, a uniform linear motion model is adopted, and the state transition matrix is: ; The state vector now simplifies to: [x,y,θ,v,ω] T , where ω is considered a constant (keeping the value from the previous time step), and Δt is the prediction step size.
[0093] (2) Nonlinear model (uniform speed turning motion); When three consecutive time steps satisfy |ω|≥δ, a uniform velocity turning motion model is adopted, and the nonlinear function F nonlinear : ; In the formula, s is the motion state vector of the suspected target, x and y are the planar positions of the target, θ is the heading angle of the target, v is the velocity of the target, ω is the yaw rate of the target; Δt is the prediction step size; Assuming that v and ω remain constant within the prediction step size Δt, for nonlinear models, the cloud center needs to calculate the Jacobian matrix F. t Used for covariance prediction: ; In the formula, s t-1|t-1 Let t-1 be the motion state vector of the suspected target.
[0094] (3) Adaptive switching logic: During the actual tracking process, the cloud monitors ω in real time. t The estimated value is obtained. When |ω| crosses the threshold within three consecutive time steps, a model switch is performed, and the covariance matrix is reset to eliminate the impact of abrupt changes. This adaptive mechanism can use a linear model to reduce computational load on straightaways and automatically enable a nonlinear model to ensure tracking accuracy on curves or serpentine escape routes.
[0095] The cloud-based system obtains the target position s of the current frame based on the trajectory of the suspected target output by the aforementioned adaptive prediction model. coord t|t-1 Plan patrol vehicle tracking routes and unmanned aerial vehicle patrol routes to achieve air-ground collaborative tracking.
[0096] like Figure 2 The figure shown is a timing diagram of incremental fine-tuning and model feedback provided in this application. As another embodiment of the present invention, a practical application scenario is specifically described: A city's traffic management department deployed five unmanned patrol vehicles and two vehicle-mounted unmanned aerial vehicles (UAVs). At 3 PM one afternoon, a black sedan was identified by a roadside checkpoint as a suspected hit-and-run vehicle (an anomaly triggered). The edge perception model on the nearest patrol vehicle (vehicle ID-03) achieved a confidence level of 0.92 for the license plate, successfully locking onto the vehicle and initiating tracking. The vehicle uploaded local gradient update fragments to the edge server every 10 minutes.
[0097] When the fleeing vehicle suddenly entered a tree-lined tunnel (an area of continuous tree shade and shadows under a bridge), the ambient light abruptly decreased from bright light to low light. The vehicle-side model's confidence level plummeted to 0.45 due to this dramatic change in lighting, triggering incremental training. The vehicle immediately activated its local emergency adaptive mechanism, using 10 cached low-light images to perform five small batches of rapid fine-tuning locally, raising the confidence level to 0.81. Subsequently, the vehicle calculated the empirical increment ΔW. exp The data is encrypted using Shamir and uploaded to the edge server. The edge server calculates the aggregated gradient and updates the global model for the region. Simultaneously, it combines the updated global model with the empirical increment ΔW. exp The data is transmitted to the cloud center. The cloud aggregates incremental experience from multiple vehicles based on the global model, generating an enhanced global model, which is then distributed to all patrol vehicles via OTA. The cloud then generates air-to-ground coordination instructions: when the escaping vehicle rapidly moves through a shadowed area, the vehicle's tracking signal becomes unstable due to continuous shadow obstruction, preventing target detection and tracking. At this point, the cloud dispatches an unmanned aerial vehicle (UAV) to take off from the patrol vehicle's roof, fly over the tree-lined tunnel, and continuously lock onto the suspect vehicle's trajectory using a high-altitude wide-angle view. The UAV transmits the target vehicle's real-time coordinates and direction of travel back to the vehicle, guiding it to successfully intercept the vehicle at the tunnel's exit. Throughout the process, all gradient data from the vehicle is transmitted in fragmented form, and sensitive target location information is not exposed in plaintext, meeting data security regulations.
[0098] It should be noted that the types of abnormal events and the hyperparameters for incremental fine-tuning (such as the number of fine-tuning rounds and the learning rate) can be dynamically configured according to actual business needs. The edge federated learning server can connect to multiple roadside edge nodes simultaneously to form a distributed federated learning network, further improving the model convergence speed.
Claims
1. A privacy-enhanced federated continuous learning-based air-ground cooperative tracking method, characterized in that, It includes the following steps: S1: Obtain the local gradient update data generated by the edge perception model deployed on the patrol vehicle after identifying and tracking the suspect target, and use the Shamir secret sharing mechanism to fragment and encrypt the local gradient update data to generate encrypted gradient fragments. The edge awareness model includes a target detection model and a target tracking model; the target detection model is used to detect suspicious targets in the video stream in real time, and the target tracking model is used to perform cross-frame identity preservation and trajectory tracking on the detected suspicious targets; The fragmentation encryption method is as follows: The local gradient update data of each vehicle is split into N encrypted fragments, denoted as: encrypted gradient fragments; The original gradient update data is reconstructed using any K encrypted gradient fragments, where K is a preset threshold, N is a preset total number of fragments, and 2 ≤ K ≤ N. S2: The encrypted gradient fragments are uploaded to the edge federated learning server through the vehicle-edge collaborative channel. The edge federated learning server performs homomorphic weighted aggregation on the encrypted gradient fragments uploaded by each vehicle to obtain the aggregated gradient, and updates the regional global model based on the aggregated gradient; and synchronizes the regional global model to the vehicle and the cloud center. When the edge federated learning server performs homomorphic weighted aggregation of encrypted gradient fragments uploaded by each vehicle, it adopts a batch fragment aggregation algorithm based on Shamir secret sharing additive homomorphism. It directly completes the multi-vehicle gradient weighted aggregation in the fragment data domain without reconstruction or obtaining plaintext gradient update data from a single vehicle. After aggregation, the aggregated fragments are reconstructed to obtain the global aggregated gradient. S3: Generate a global model based on the global models of all regions in the cloud center; Meanwhile, the cloud center receives and detects abnormal event trigger signals from the vehicle in real time; when the abnormal event trigger signal triggers the incremental training condition, the cloud center uses the abnormal event data to incrementally fine-tune the benchmark model and generate an enhanced global model. S4: The cloud center feeds the enhanced global model back to the vehicle and unmanned aerial vehicle via over-the-air download technology, and generates air-ground collaborative tracking instructions based on the enhanced global model, and schedules associated unmanned aerial vehicles to conduct cross-domain continuous tracking of the suspected target.
2. The air-ground cooperative tracking method based on privacy-enhanced federated continuous learning according to claim 1, characterized in that: The method for generating encrypted gradient fragments includes the following operations: The specific fragment generation process is as follows: Let the local gradient update data at the vehicle end be a vector g∈R d Choose a large prime number q > max(‖g‖, N), and randomly generate a K-1 degree polynomial over the finite field GF(q): f(x)=g+a1x+a2x 2 +…+a K-1 x K-1 (mod q); Where a1, a 2, … , a K-1 For random coefficients; x, x 2 ,…,x K-1 The x-coordinate is non-zero; Then the i-th encrypted fragment is: fragment i=(x i ,f(x i )), i=1,2,…,N.
3. The air-ground cooperative tracking method based on privacy-enhanced federated continuous learning according to claim 1, characterized in that: The target tracking model is based on the DeepSORT algorithm and introduces adaptive switching of motion modes: it automatically switches between a uniform linear motion model and a uniform turning nonlinear motion model according to the yaw rate of the suspected target, and uses extended Kalman filtering to complete nonlinear state prediction; it supports local emergency adaptive fine-tuning triggered by abnormal events, and only performs local small-batch fine-tuning on the fully connected layer of the ReID network, and outputs the experience increment for federated incremental learning.
4. The air-ground cooperative tracking method based on privacy-enhanced federated continuous learning according to claim 1, characterized in that: In step S2, the process of updating the global model of the region includes the following steps: Let U be the set of vehicles participating in this round of aggregation, and let ΔW be the local gradient update amount of vehicle u. u Its training sample size is n u The aggregate gradient ΔW calculated by the edge federated learning server is... agg for: △W agg =(∑ u∈U (n u ·△W u )) / (∑ u∈U n u ); The edge federated learning server updates the regional global model at time t+1 according to the asynchronous update rule, obtaining the updated regional global model W'. global (t+1) : IN' global (t+1) =W' global (t) -ηw △W agg ; Where ηw is the federated learning aggregation rate; W' global (t) This represents the global model of the region at time t.
5. The air-ground cooperative tracking method based on privacy-enhanced federated continuous learning according to claim 3, characterized in that: In step S3, the abnormal event triggering signals include: daily detection of suspected targets triggering events, vehicle-side model confidence continuously falling below a preset threshold, or new detection tasks triggering events; The incremental fine-tuning process includes: after the vehicle detects an abnormal event that triggers the incremental training condition, it first activates the local emergency adaptive mechanism and uses the currently cached abnormal scene data to perform small-batch rapid fine-tuning of the last two layers of the fully connected layer of the ReID network in the edge perception model; the vehicle calculates the difference in model weights before and after fine-tuning as an experience increment, which is then uploaded to the edge federated learning server and the cloud center after being encrypted by Shamir's secret sharing; the cloud center performs weighted fusion of the experience increments uploaded by each vehicle to generate an enhanced global model.
6. The air-ground cooperative tracking method based on privacy-enhanced federated continuous learning according to claim 5, characterized in that: The enhanced global model generation process specifically includes the following steps: a1: After the vehicle detects an abnormal event, it reports it to the cloud center through the edge federated learning server; Simultaneously, the bottom convolutional layers are frozen, and the vehicle-side initiates a local emergency adaptive mechanism to perform rapid, small-batch fine-tuning of the last two fully connected layers of the ReID network in the edge perception model. Then, the difference in model weights before and after the repair is calculated, generating an empirical increment ΔW. exp ; △In exp =W car '-IN car ; In the formula, W car To fix the previous model weights, W car 'These are the weights for the repaired model; a2: The vehicle will increment the experience by △W exp The data is split into encrypted fragments and uploaded to the edge federated learning server; the edge server aggregates and reconstructs the incremental fragments uploaded by each vehicle and uploads the resulting experience increments to the cloud; a3: The cloud center detects abnormal events in real time based on the abnormal event detection function ε(t); ε(t)=II(Conf(t)<θ conf ) or II(Alarm(t)=1) or II(Event(t)=1); In the formula, II() represents an indicator function, which takes the value of 1 when the condition in parentheses is true, and 0 otherwise when the condition is false; or represents a logical OR operation; Conf(t) is the confidence level of the vehicle-side model in identifying the suspect target at time t; Alarm(t) is the daily detection trigger event for the suspect target issued by the vehicle-side; Event(t) is the trigger event for adding a new detection task. a4: When the output value of ε(t) is 1, it indicates that an abnormal event occurred at time t, and step a5 is executed; Otherwise, when the output value of ε(t) is 0, continue to detect abnormal events in real time based on the abnormal event detection function ε(t); a5: The cloud center is in the global model W global Based on this, the aggregated experience increments are weighted and fused to generate an enhanced global model W. enh ; ; Among them, △W exp (m) This represents the experience increment corresponding to the m-th vehicle terminal; M is the total number of vehicle terminals that reported abnormal events, and η is the cloud aggregation weight factor.
7. The air-ground cooperative tracking method based on privacy-enhanced federated continuous learning according to claim 3, characterized in that: The adaptive switching logic for motion modes includes the following operations; b1: During the actual tracking process, the cloud center monitors the estimated value of the yaw rate ω of the suspect target in real time; When |ω| exceeds the preset threshold δ for three consecutive time steps, execute step b3; Otherwise, proceed to step b2; b2: Execute the uniform linear motion model; s coord t|t-1 =F linear ·s coord t-1|t-1 ; In the formula, s coord t|t-1 Let t be the predicted location of the suspect target in the preceding frame of the video at time t; s coord t-1|t-1 Let s be the motion state vector of the suspected target at time t-1; coord t =[x t ,y t ,θ t ,v t ,ω t ] T x and y represent planar positions, θ is the heading angle of the suspected target, v is the velocity of the suspected target, and ω is the yaw rate. F linear This represents the state transition matrix of a uniform linear motion model; ; △t is the prediction step size; b3: Execute the uniform speed turning motion model; s coord t|t-1 =F nonlinear (s coord t-1|t-1 ); In the formula, F nonlinear For representing the nonlinear function of the uniform speed turning motion model; ; Furthermore, state prediction is performed based on extended Kalman filtering: Assuming that v and ω remain constant within the prediction step size Δt, for the nonlinear model, the cloud center calculates the Jacobian matrix F. t Used for covariance prediction: ; In the formula, s t-1|t-1 Let t-1 be the motion state vector of the suspected target.
8. The air-ground cooperative tracking method based on privacy-enhanced federated continuous learning according to claim 1, characterized in that: The airborne model carried in the unmanned aerial vehicle is a homologous instance of the vehicle-side global model after structural scaling to meet edge computing power constraints. The two share a unified loss gradient space when they are aggregated in federated learning. After completing inference and training batches locally, the airborne model generates local gradient updates, which are then uploaded to the edge federated learning server via Shamir's secret sharing and fragmentation encryption, and subsequently synchronized to the cloud center. The cloud center generates an enhanced global model, which is then fed back to all vehicles and all unmanned aerial vehicles simultaneously via over-the-air download technology.
9. An air-ground cooperative tracking system based on privacy-enhanced federated continuous learning, characterized in that, It includes: a vehicle-side privacy awareness and tracking module, an edge federated learning server, a cloud-based incremental evolution module, an air-ground collaborative tracking and scheduling module, and an unmanned aerial vehicle collaborative unit; The vehicle-side privacy awareness and tracking module is deployed on the vehicle side and is used to acquire local gradient update data generated by the edge perception model after identifying and tracking the suspect target. The local gradient update data is fragmented and encrypted using the Shamir secret sharing mechanism to generate encrypted gradient fragments. The vehicle-side privacy awareness and tracking module includes: a target detection module and a target tracking module; the target detection module embeds a target detection model, which uses current parameters to detect targets in the video stream in real time; the target tracking module embeds a target tracking model, which continuously detects and tracks suspected targets and reports the target location in real time. The edge federated learning server is deployed in the roadside equipment and communicates with the vehicle-side privacy awareness and tracking module through the vehicle-edge collaborative channel. It is used to receive encrypted gradient fragments uploaded by each vehicle, perform homomorphic weighted aggregation on each encrypted gradient fragment to obtain an aggregated gradient, update the regional global model based on the aggregated gradient, and transmit the regional global model to the vehicle and the cloud center. The cloud-based incremental evolution module and the air-to-ground collaborative tracking and scheduling module are deployed in the cloud center and communicate with the edge federated learning server. They are used to receive abnormal event trigger signals reported by vehicle-side and roadside equipment. After the abnormal event detector determines the signal, the incremental training trigger starts the incremental training process. The transfer learning fine-tuning engine performs incremental fine-tuning on the baseline model to generate an enhanced global model. The air-to-ground collaborative tracking and scheduling module feeds the global model back to each patrol vehicle and generates air-to-ground collaborative tracking control commands to schedule unmanned aerial vehicles (UAVs) to track suspected targets. The UAVs and each patrol vehicle interact with each other through a two-way communication link, transmitting high-altitude target images and location information back to the vehicle in real time. The vehicle reports the target tracking status to the UAV. The unmanned aerial vehicle (UAV) collaborative unit is used to automatically take off after receiving scheduling instructions. The UAV collaborative unit adaptively selects the access method according to the current communication status: when it is within the coverage area of the edge federated learning server, it accesses the edge federated learning server and participates in the three-level federated learning architecture; when it is outside the coverage area of the edge server, it communicates directly with the cloud center and participates in the two-level federated learning architecture. The UAV collaborative unit performs aerial tracking based on the target position and predicted trajectory reported by the vehicle and transmits back images of the suspected target from a high-altitude perspective. When working in the two-level federated learning architecture, the UAV still performs Shamir secret sharing fragment encryption locally, and the encrypted fragments are directly uploaded to the cloud center, where the cloud center completes the fragment aggregation and reconstruction.
Citation Information
Patent Citations
Traffic accident escape tracking method and system based on unmanned vehicle and air-ground integration
CN120877534A
Patrol artificial intelligence model-based patrol method, system, equipment and medium
CN121857686A