An edge-assisted drone collaborative and adaptive bird's-eye view perception method and system

By using an edge server-assisted UAV collaborative adaptive bird's-eye perception method, which combines lightweight processing on the UAV side with computing on the edge server side, the computational complexity and bandwidth consumption issues of UAV bird's-eye perception are solved, and real-time high-precision bird's-eye perception is achieved.

CN122289905BActive Publication Date: 2026-07-31UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UNIV OF SCI & TECH OF CHINA
Filing Date
2026-05-26
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing UAV bird's-eye perception methods face challenges in terms of computational complexity and bandwidth consumption, making it difficult to achieve real-time, high-precision bird's-eye perception on resource-constrained UAV platforms.

Method used

An edge-assisted UAV collaborative adaptive bird's-eye view perception method is adopted. By combining lightweight processing on the UAV end and computing on the edge server end, spatial regions of interest and semantic regions of interest are generated and fused to perform differentiated video processing. The video transmission strategy is adaptively adjusted according to network status and perception performance.

Benefits of technology

It reduces the computational burden on the drone, improves the system's real-time performance and bird's-eye view accuracy, adapts to dynamic network environments, and reduces bandwidth consumption and end-to-end latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289905B_ABST
    Figure CN122289905B_ABST
Patent Text Reader

Abstract

This invention relates to the field of autonomous driving technology and discloses an edge-assisted UAV cooperative adaptive bird's-eye view perception method and system. The method includes: acquiring video frames of a road scene and extracting semantic regions of interest (ROIs); fusing the spatial ROIs with the semantic ROIs to obtain a fused ROI; the UAV performs differential processing on the video frames based on control parameters and the fused ROIs, encodes the processed video frames into a video stream, and uploads it to an edge server; the edge server decodes the video stream, performs a bird's-eye view perception task, updates the spatial ROI based on perception performance, and determines the control parameters for the next moment to adjust the differential processing and encoding of the video frames for the next moment; this method utilizes the UAV's aerial view advantage to acquire road scene video and performs the main bird's-eye view perception calculations on the edge server, reducing the computational burden on the UAV and improving the feasibility of system deployment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving technology, and more specifically to an edge-assisted drone cooperative adaptive bird's-eye view perception method and system. Background Technology

[0002] With the development of autonomous driving (AD), environmental perception has become a crucial foundational capability in autonomous driving systems. While traditional onboard sensors can perceive the surrounding environment in real time, their perception range is limited and they are easily affected by occlusion, making it difficult to obtain complete and stable global perception performance in complex road environments. To expand the perception range and reduce the impact of occlusion, researchers have begun to introduce roadside infrastructure (RI) or unmanned aerial vehicles (UAVs) to participate in cooperative perception. Among them, UAVs have advantages such as high field of view, high maneuverability, and wide coverage, enabling them to acquire aerial views of road scenes and provide more complete global spatial information for autonomous driving systems.

[0003] Bird's-eye view (BEV) perception is a widely used environmental modeling method in the field of autonomous driving in recent years. This method maps environmental information collected from different perspectives onto a unified overhead space, allowing vehicles, pedestrians, obstacles, and road structures in the scene to be represented in a unified coordinate system. Due to its strong global perspective, convenient data fusion, and advantages for subsequent planning and decision-making, BEV perception has attracted widespread attention. Video collected by drones from the air naturally possesses an overhead advantage and is more suitable as a data source for BEV perception compared to traditional vehicle-mounted perspectives.

[0004] However, the use of drones in bird's-eye view perception still presents significant challenges. First, bird's-eye view perception models typically involve high computational complexity, while drone platforms have limited onboard computing power and battery capacity, making it difficult to complete complex perception tasks locally in real time. Second, directly transmitting high-resolution video collected by drones to edge servers for processing consumes substantial wireless bandwidth, increasing transmission latency and failing to meet the real-time requirements of autonomous driving systems. Third, existing video compression and transmission methods largely rely on semantic information or traditional coding features in two-dimensional images. While this can reduce the amount of data transmitted to some extent, these methods often neglect the spatial structural information upon which bird's-eye view perception depends, thus impacting the accuracy of subsequent bird's-eye view perception.

[0005] In existing technologies, one type of method detects target regions such as vehicles and pedestrians in images and prioritizes preserving the clarity of these regions, thereby reducing the transmission overhead of background areas. Another type of method compresses image and video data by adjusting encoding parameters to reduce the amount of data transmitted. Yet another type of method attempts to extract intermediate features on the front-end device before transmission to reduce the transmission pressure of the original video. However, the above methods are either primarily geared towards two-dimensional target recognition tasks and cannot meet the spatial structure information requirements of bird's-eye view perception; or they require high computing power from the front-end device, making them unsuitable for resource-constrained UAV platforms; or they lack the ability to adaptively adjust processing strategies based on dynamic changes in the wireless link. Therefore, existing methods struggle to simultaneously balance transmission overhead, system real-time performance, and bird's-eye view perception accuracy. Summary of the Invention

[0006] To address the aforementioned technical issues, this invention provides an edge-assisted UAV collaborative adaptive bird's-eye view perception method and system. This method performs only lightweight processing on the UAV end, while the main bird's-eye view perception calculations are completed on the edge server end. Furthermore, the perception performance of the edge server guides the video processing and transmission strategies on the UAV end, thereby improving the real-time performance and accuracy of bird's-eye view perception under bandwidth-constrained conditions.

[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: In a first aspect, the present invention provides an edge-assisted UAV cooperative adaptive bird's-eye perception method, comprising the following steps: S1, the drone captures video frames of the road scene and extracts semantic regions of interest; S2, the UAV receives the spatial region of interest and control parameters fed back from the edge server, and merges the spatial region of interest with the semantic region of interest to obtain the fused region of interest; S3, the drone performs differential processing on video frames based on control parameters and the region of interest for fusion. It maintains or increases the resolution of the region within the region of interest for fusion, and reduces the resolution of the region outside the region of interest for fusion. The differentially processed video frames are then encoded into video streams and uploaded to the edge server. S4, the edge server decodes the video stream, performs the bird's-eye view perception task, updates the spatial region of interest based on the obtained perception performance, analyzes the network status and perception performance, and decides the control parameters for the next moment. S5: The edge server feeds back the spatial region of interest and control parameters to the drone to adjust the video frame differentiation processing and encoding for the next moment. S6. Repeat steps S1 to S5 until the set termination condition is met; the termination condition can be a specific termination command, completion of a specific task, etc.

[0008] In one embodiment, the fused region of interest is jointly determined by the semantic region of interest local on the UAV and the spatial region of interest fed back by the edge server, wherein the semantic region of interest is used to indicate foreground targets and the spatial region of interest is used to indicate background spatial regions that are important for bird's-eye view perception.

[0009] In one embodiment, the UAV terminal performs differential processing on video frames based on control parameters and fused regions of interest, specifically including: At time t, the original image frame is denoted as The image frames after differentiation processing are denoted as Then we have: ; in, This represents an adaptive space processing function; This represents video processing parameters and indicates the intensity of differentiated processing. This indicates the region of interest for merging.

[0010] In one embodiment, the analysis of network state and perception performance to determine the control parameters for the next moment specifically includes: Control parameters include video processing parameters and encoding parameters; network status includes video bitrate, end-to-end latency and available bandwidth; perception performance specifically includes target detection results, spatial layout information and perception performance indicators. end-to-end Latency includes drone-side encoding latency Wireless transmission latency Edge server-side decoding latency and edge server-side bird's-eye view latency perception : .

[0011] The process of encoding the differentiated video frames into a video stream and then uploading it to the edge server specifically includes: Let the video bitrate at time t be... : ; in, Represents the encoding function. These are encoding parameters used to control the quality of video frames; These are the image frames after differential processing.

[0012] In one embodiment, it further includes: To achieve a balance between perception performance and network status, the edge server defines a reward function for each time step; the perception performance at time step t is denoted as... End-to-end delay is denoted as Then the payoff function for:

[0013] in, This represents the latency penalty weight; during continuous operation, maximizing the benefit is the optimization objective. ; in, As a discount factor, It expresses expectation.

[0014] Secondly, the present invention provides an edge-assisted UAV cooperative adaptive bird's-eye view perception system, comprising: Drones: The drone acquires video frames of road scenes, extracts semantic regions of interest (ROIs), receives spatial ROIs and control parameters from the edge server, merges the spatial ROIs and semantic ROIs to obtain a fused ROI, and performs differential processing on the video frames based on the control parameters and the fused ROIs. The resolution of the regions within the fused ROIs is maintained or increased, while the resolution of the regions outside the fused ROIs is reduced. The differentially processed video frames are then encoded into a video stream and uploaded to the edge server. Edge server: Decodes video stream, performs bird's-eye view perception task, updates spatial region of interest based on the obtained perception performance, analyzes network status and perception performance, determines control parameters for the next moment, and feeds back spatial region of interest and control parameters to the UAV to adjust the video frame differentiation processing and encoding for the next moment.

[0015] In one embodiment, the fused region of interest is jointly determined by the semantic region of interest local on the UAV and the spatial region of interest fed back by the edge server, wherein the semantic region of interest is used to indicate foreground targets and the spatial region of interest is used to indicate background spatial regions that are important for bird's-eye view perception.

[0016] In one embodiment, the UAV terminal performs differential processing on video frames based on control parameters and fused regions of interest, specifically including: At time t, the original image frame is denoted as The image frames after differentiation processing are denoted as Then we have: ; in, This represents an adaptive space processing function; This represents video processing parameters and indicates the intensity of differentiated processing. This indicates the region of interest for merging.

[0017] In one embodiment, the analysis of network state and perception performance to determine the control parameters for the next moment specifically includes: Control parameters include video processing parameters and encoding parameters; network status includes video bitrate, end-to-end latency and available bandwidth; perception performance specifically includes target detection results, spatial layout information and perception performance indicators. end-to-end Latency includes drone-side encoding latency Wireless transmission latency Edge server-side decoding latency and edge server-side bird's-eye view latency perception : .

[0018] The process of encoding the differentiated video frames into a video stream and then uploading it to the edge server specifically includes: Let the video bitrate at time t be... : ; in, Represents the encoding function. These are encoding parameters used to control the quality of video frames; These are the image frames after differential processing.

[0019] In one embodiment, it further includes: To achieve a balance between perception performance and network status, the edge server defines a reward function for each time step; the perception performance at time step t is denoted as... End-to-end delay is denoted as Then the payoff function for:

[0020] in, This represents the latency penalty weight; during continuous operation, maximizing the benefit is the optimization objective. ; in, As a discount factor, It expresses expectation.

[0021] The system and method in this invention correspond to each other; the specific technical solutions applicable to the method are also applicable to the system.

[0022] Compared with the prior art, the beneficial technical effects of the present invention are: This method leverages the aerial view advantage of UAVs to acquire road scene videos and performs the main bird's-eye view perception calculations on the edge server, reducing the computational burden on the UAV and improving the feasibility of system deployment.

[0023] This method generates a spatial region of interest on the edge server and merges it with a semantic region of interest extracted by the UAV. This can reduce video transmission redundancy while better preserving spatial structural information that is important for bird's-eye view perception tasks.

[0024] This method can adaptively adjust the video processing and encoding parameters of the UAV based on the network status and sensing performance of the wireless link, thereby improving the system's adaptability to dynamic network environments.

[0025] Compared to video transmission methods that rely solely on two-dimensional semantic information, this method takes into account the spatial structure information requirements of bird's-eye view perception tasks; compared to methods that directly transmit raw high-resolution video, this method reduces bandwidth consumption and end-to-end latency; and compared to methods that require the front end to perform complex perception calculations, this method is more suitable for resource-constrained UAV platforms.

[0026] This method achieves collaborative optimization of video transmission and perception tasks through a closed-loop approach of "lightweight processing on the UAV end - perception on the edge server end - feedback on the edge server end - update strategy on the UAV end", which can continuously improve system performance during long-term operation.

[0027] This method clearly distinguishes between the drone-side functions, edge server-side functions, and wireless link functions in its structure, which facilitates subsequent deployment and expansion under different drone platforms, different edge computing platforms, and different communication network conditions. Attached Figure Description

[0028] Figure 1 This is a diagram of the overall system architecture of the present invention.

[0029] Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation

[0030] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.

[0031] This invention combines UAV video acquisition, lightweight semantic processing, adaptive spatial processing, video encoding and transmission, edge server-based bird's-eye view perception, and feedback decision-making to achieve efficient UAV video transmission and high-precision edge-side bird's-eye view perception under conditions of limited bandwidth and dynamically changing wireless links. The method generates a spatial region of interest (ROI) related to bird's-eye view perception at the edge server and merges it with a semantic ROI extracted by the UAV, guiding the UAV to perform differentiated processing on video frames. Simultaneously, it adaptively adjusts video processing and encoding parameters based on real-time network conditions and perception performance, thereby reducing transmission redundancy and improving overall system performance.

[0032] Based on the characteristics of drone-assisted bird's-eye perception and the need for optimized video transmission, a design was developed as follows: Figure 1 The system shown consists of three main parts: the drone terminal, the edge server terminal, and the wireless communication link.

[0033] 1. Drone terminal.

[0034] The drone-based terminal directly faces real-world road scenarios, responsible for video capture and lightweight front-end processing. It includes modules for video capture, semantic information extraction, spatial processing, video encoding, and communication.

[0035] The video acquisition module captures video frames of the road scene and passes them to the semantic information extraction module and the spatial processing module. The semantic information extraction module performs lightweight analysis on the captured video frames to identify key target regions in the scene, including but not limited to vehicles, pedestrians, and other traffic participants, forming semantic regions of interest (ROIs). The spatial processing module receives the ROIs from the edge server and processes them... Semantic region of interest obtained by the semantic information extraction module The regions of interest are fused to guide video processing. : .

[0036] After obtaining the region of interest (ROI) for fusion, the spatial processing module performs differential processing on the video frames. For image regions within the ROI, the original resolution or a higher resolution is maintained; for image regions outside the ROI, downsampling is used to reduce resolution and minimize redundant information. The processed image frames are then passed to the video encoding module for encoding and compression. The video encoding module encodes the processed video frames according to the currently set encoding parameters, generating a video stream, which is then sent to the edge server via the communication module.

[0037] The drone also receives control parameters and spatial regions of interest from the edge server, and adjusts the spatial processing and encoding of subsequent video frames based on the received information to achieve adaptive video transmission.

[0038] Specifically, at the t-th decision time, the original image frame is denoted as... The image frame after spatial processing is denoted as Then we have:

[0039] in, This represents an adaptive spatial processing function, specifically the downsampling function provided by the OpenCV library. This represents video processing parameters and indicates the intensity of differentiated processing. This indicates the region of interest for merging.

[0040] 2. Edge server side.

[0041] The edge server is the core component of the entire system, primarily responsible for video decoding, bird's-eye view perception, spatial feedback generation, and adaptive decision-making. The edge server includes modules for video decoding, bird's-eye view perception, spatial feedback, state analysis, decision-making, and communication.

[0042] The video decoding module receives the video stream from the UAV, decodes it to recover the processed video frames, and passes the decoding results to the bird's-eye view (BOV) perception module. The BOV perception module performs a bird's-eye view perception task based on the decoded video frames, obtaining target detection results (target location, target category, or a binary region map generated from the target location), spatial layout information, and perception performance metrics. The spatial feedback module extracts key spatial regions that significantly impact BOV perception based on the perception performance, converts these key regions into regions of interest (ROIs) usable by the UAV, and sends them to the UAV. The BOV perception task uses a BEVDet-based model. The output spatial layout information includes the position and size of vehicles, pedestrians, obstacles, and other key objects in the bird's-eye view. The key spatial regions are generated directly from the bird's-eye view using the BOV perception model and can be mapped back to the image to obtain the UAV's ROI. The perception performance metric is the mean Average Precision (mAP) metric, the default metric for the BEVDet model.

[0043] The state analysis module is used to statistically analyze and organize the current sensing performance and network status, including but not limited to video bitrate, end-to-end latency, available bandwidth, and sensing performance metrics. Based on the information provided by the state analysis module, the decision module determines the video processing and encoding parameters for the next moment and sends these parameters to the drone via the communication module. Thus, the edge server can perform closed-loop adjustment of the drone's video transmission process based on real-time network changes and sensing performance.

[0044] Specifically, let the video bitrate at the t-th decision time be... Then it can be written as: ; in, This represents an abstract encoding function representing the encoding process based on the DCVC-RT model. These are encoding parameters in the DCVC-RT model's encoding process, used to control the quality of video frames. Correspondingly, end-to-end latency consists of encoding latency, transmission latency, decoding latency, and perception latency. ; in, This indicates the encoding latency at the drone end. Indicates wireless transmission latency. This indicates the decoding latency on the edge server side. This indicates the latency perceived from the edge server's perspective.

[0045] 3. Wireless communication link.

[0046] The wireless communication link is used for data transmission and control information exchange between the UAV and the edge server. On one hand, the wireless communication link sends the video stream acquired and processed by the UAV to the edge server; on the other hand, it sends the region of interest and control parameters generated by the edge server back to the UAV. Since the available bandwidth of the wireless communication link changes over time, this solution needs to dynamically adjust the video transmission method based on the link status.

[0047] Specifically, if the available bandwidth at the t-th decision time is Then the transmission delay This can be expressed as the relationship between the video bitrate and the currently available bandwidth, i.e.: ; After completing the current moment's perception and state analysis, the edge server updates the parameters for the next moment based on the new available bandwidth, end-to-end latency, and perception state.

[0048] 4. System module interface relationships.

[0049] The interface relationship between the UAV and the edge server: The UAV sends the processed and encoded video stream from the acquired video frames to the edge server via a communication interface. The edge server receives the video stream, performs decoding, bird's-eye view perception, and state analysis, and then returns the spatial region of interest (ROI) and control parameters to the UAV via the communication interface. The UAV adaptively processes subsequent video frames based on the returned results. Specifically, the content uploaded by the UAV to the edge server includes at least the processed video stream and basic state information corresponding to the current video frame, such as control parameters and timestamps; the content returned by the edge server to the UAV includes at least the ROI and control parameters. The ROI guides the UAV to re-determine the image areas that need to be prioritized in the next moment, and the control parameters guide the UAV to adjust the spatial processing intensity and encoding intensity.

[0050] Interface relationships among internal modules of the UAV: ​​The various modules within the UAV communicate via data interfaces. The video acquisition module sends acquired video frames to the semantic information extraction module and the spatial processing module; the semantic information extraction module sends its obtained semantic regions of interest (ROIs) to the spatial processing module; the spatial processing module also receives ROIs returned from the edge server and performs differentiated processing on the video frames after fusion; the video encoding module receives the processed video frames and control parameters and generates a video stream; the communication module is responsible for sending the video stream to the edge server and receiving ROIs and control parameters from the edge server. Specifically, after extracting key targets from the image frames, the semantic information extraction module sends the target location, target category, or a binary region map generated from the target location to the spatial processing module. The spatial processing module receives the ROI from the edge server, fuses it with its local semantic ROI, and forms the final fused ROI. Subsequently, the spatial processing module performs different spatial processing methods on key and non-key regions based on the range of the fused ROI. The processed image frames are then sent to the video encoding module. The entire internal process of the drone consists of a one-way data link of "collection-identification-fusion-processing-encoding-transmission".

[0051] Interface relationships among internal modules of the edge server: The modules within the edge server also communicate via data interfaces. The video decoding module receives the video stream sent by the UAV and sends the decoded video frames to the bird's-eye view perception module. The perception performance output by the bird's-eye view perception module is sent to the spatial feedback module and the state analysis module respectively. The state analysis module organizes the network state and perception performance and sends it to the decision module. The decision module sends the generated control parameters to the communication module, which then distributes the control parameters to the UAV. The spatial feedback module sends the generated spatial region of interest to the communication module, which returns it to the UAV. Specifically, the perception performance output by the bird's-eye view perception module includes at least target detection results, spatial layout information, and perception performance indicators. The spatial feedback module extracts the spatial region of interest (region of interest) that has a significant impact on bird's-eye view perception based on the above results. The state analysis module obtains the current bitrate, end-to-end latency, current available bandwidth, and current perception performance, and organizes the current available bandwidth, previous bitrate, previous end-to-end latency, and previous perception performance indicators into a state variable and sends it to the decision module. The decision module outputs the video processing parameters and encoding parameters for the next moment based on the state variable. The entire internal process of the edge server forms a closed-loop link of "decoding-perception-feedback / analysis-decision-deployment".

[0052] The interface relationship between the wireless communication link and the overall system: The wireless communication link undertakes uplink data transmission and downlink control feedback functions in the system. The UAV uploads the video stream to the edge server via the wireless communication link, and the edge server returns the region of interest and control parameters to the UAV via the wireless communication link. Considering that the available bandwidth of the wireless link is variable, the edge server needs to continuously monitor and estimate the current available bandwidth status, and use this as an important basis for subsequent parameter decisions. When the wireless link is in good condition, the system can appropriately increase the coding quality parameters or reduce the background sampling intensity to improve the bird's-eye view perception accuracy at the edge; when the wireless link is in poor condition, the system can increase the background sampling intensity or adjust the coding parameters to reduce the transmission bit rate and transmission latency. Therefore, the wireless communication link not only undertakes data transmission but also directly affects the closed-loop adjustment process of the entire system.

[0053] 5. Platform Function Description.

[0054] This method addresses the perception needs of drone-assisted autonomous driving and mainly provides four types of functions: video acquisition and lightweight processing, spatial feedback-guided video differentiation processing, edge-side bird's-eye view perception, and network adaptive optimization.

[0055] Video Acquisition and Lightweight Processing: This method provides video acquisition and lightweight processing capabilities on the UAV side. The UAV acquires video of the road scene and extracts semantic regions of interest (ROIs) based on a lightweight model, providing a basis for subsequent video differentiation processing without significantly increasing the onboard computational burden. Specifically, this function allows the UAV to perform preliminary identification and local semantic extraction of key target areas without running a complete bird's-eye view perception model. Since the UAV primarily performs lightweight processing, it reduces the computational and energy consumption pressure on the onboard equipment, improving the system's deployability on practical UAV platforms. Furthermore, the UAV can process the acquired video at fixed time intervals or fixed frame intervals to maintain consistency with the decision-making cycle of the edge server.

[0056] Spatial Feedback-Guided Differentiated Video Processing: This method provides differentiated video processing functionality guided by feedback from the edge server. The edge server identifies spatial regions important to the perception task based on its bird's-eye view perception performance. The UAV then fuses these spatial regions with its local semantic region of interest (GROUP) to form a fused GROUP. Background regions outside the fused GROUP are downsampled to reduce the transmission overhead of irrelevant redundant information. Compared to methods relying solely on the UAV's local semantic information, this function further utilizes the spatial structure information contained in the edge server's bird's-eye view perception performance. This allows the UAV to retain not only the target foreground region but also background spatial structures crucial for bird's-eye view perception, such as road boundaries, key intersections, or the contextual spatial region surrounding the target. This reduces the amount of data transmitted while better maintaining the accuracy of edge-side bird's-eye view perception.

[0057] Edge-side bird's-eye view perception: This method provides edge-side bird's-eye view perception. The edge server decodes the video transmitted from the drone and performs bird's-eye view perception calculations to generate scene understanding results for autonomous driving tasks. Since the main calculation process is completed on the edge server, the problem of insufficient local computing power of the drone is avoided. Specifically, the edge server can analyze the positional distribution of vehicles, pedestrians, and other traffic participants in the scene and obtain target perception performance under a unified top-down space. The perception performance of the edge server is not only used for the autonomous driving system itself, but can also be used to generate regions of interest in space and serve subsequent video processing decisions on the drone.

[0058] Network Adaptive Optimization: This method provides adaptive optimization for dynamic wireless links. The edge server adaptively adjusts the video processing and encoding parameters of the UAV based on network conditions and bird's-eye view perception performance, enabling the system to balance perception accuracy and end-to-end latency under different network conditions. Specifically, the edge server combines the current available bandwidth, the bitrate of the previous moment, the end-to-end latency of the previous moment, and the perception performance indicators of the previous moment into a state variable, and outputs the control parameters for the next moment based on this state variable.

[0059] System operation data: Operation data is generated during and after the operation of this method, including video bitrate data, end-to-end latency data, bandwidth change data, perception performance data, and control parameter recording data.

[0060] 6. Method and process description.

[0061] The workflow of this method is mainly divided into the following four parts: (1) Acquire scene video and extract semantic regions of interest from the UAV terminal: In this process, the video acquisition module on the drone captures video frames of the road scene and sends them to the semantic information extraction module. The semantic information extraction module extracts key target areas such as vehicles and pedestrians from the video frames to form semantic regions of interest.

[0062] (2) The drone processes and uploads video frames: The drone performs differentiated processing and encoding compression on video frames based on the current video processing and encoding parameters, generates a video stream, and uploads it to the edge server via a wireless communication link.

[0063] (3) The edge server performs bird's-eye view perception and generates feedback information: After receiving the video stream, the edge server decodes it and performs a bird's-eye view perception task to obtain scene perception performance. The spatial feedback module extracts key spatial regions based on the perception performance and generates regions of interest. The state analysis module calculates the current bitrate, latency, available bandwidth, and perception performance metrics.

[0064] (4) The edge server generates control parameters and sends them to the drone: The decision-making module determines the video processing and encoding parameters for the next moment based on the state variables output by the state analysis module, and returns the control parameters and spatial region of interest to the UAV. The UAV then processes and transmits the video for the next moment based on the returned results, entering the next loop.

[0065] Specific process one: System initialization and initial processing: Upon system startup, the drone first acquires initial scene video and inputs the initial video frames into the semantic information extraction module. Since the edge server has not yet returned the region of interest at the initial moment, the drone can process and upload the video frames using pre-set initial video processing and encoding parameters. After receiving the initial video stream, the edge server completes the first round of bird's-eye view perception, state analysis, and spatial feedback generation.

[0066] This initial processing procedure is used to establish the first closed loop from the drone end to the edge server end and back to the drone end, so that the adaptive processing in subsequent moments has a state basis.

[0067] The second specific process involves integrating region-of-interest driven video processing: After receiving the spatial region of interest (ROI) returned by the edge server, the drone fuses this ROI with the local semantic region of interest (SRI) to obtain a fused ROI. The spatial processing module then performs differential processing on the image frame based on the fused ROI. Key regions retain their original resolution or a higher resolution, while background regions are downsampled to reduce their resolution.

[0068] If we denote the region of interest for fusion at time t as... The spatially processed image frame can then be represented as: .

[0069] Specific process three: Edge server-side status evaluation and parameter update: After completing video decoding, bird's-eye view perception, and state statistics at each decision-making moment, the edge server organizes the current system state into a state vector. This state vector includes at least the bitrate of the previous moment, the end-to-end latency of the previous moment, the currently available bandwidth, and the perception performance metrics of the previous moment. Let's denote it as... Then it can be written as: ; in, The bitrate of the previous moment. The end-to-end delay of the previous moment. The available bandwidth at the current moment. The performance metrics perceived at the previous moment.

[0070] The decision module outputs the control parameters for the next time step based on the state vector. Let the control parameters at time t be denoted as... Then we have: ; in, Indicates video processing parameters, This represents the encoding parameters.

[0071] The fourth specific step involves closed-loop optimization based on perceived performance and latency: To achieve a balance between perceived performance and end-to-end latency, edge servers can define a reward function for each decision time. If the perceived performance index at time t is... End-to-end delay is denoted as Then the payoff function can be expressed as: ; in, This represents the delay penalty weight. During continuous operation, the system objective is to maximize long-term gains, which can be expressed as: ; in, This is the discount factor.

[0072] Example: In this embodiment, the drone cruises above the road scene, capturing video of vehicles, pedestrians, and road structures. The video capture module on the drone sends the captured video frames to the semantic information extraction module. The semantic information extraction module can use the YOLOv8s model to identify key target regions in the video frames, forming semantic regions of interest.

[0073] In the initial stage of the system, since the edge server has not yet returned the region of interest in space, the UAV can process the video frames according to the preset initial processing parameters, and after encoding by the video encoding module, send them to the edge server via the communication module.

[0074] After receiving the video stream from the drone, the edge server decodes the video stream using its video decoding module and sends the decoded video frames to the bird's-eye view perception module. The bird's-eye view perception module performs bird's-eye view perception processing on the scene to obtain target distribution and spatial layout information in the current scene.

[0075] The spatial feedback module extracts key spatial regions that significantly impact the bird's-eye view perception task based on the bird's-eye view perception performance, and sends these key spatial regions as regions of interest to the UAV. Simultaneously, the state analysis module analyzes the current network state and perception performance, and the decision-making module generates control parameters for the next moment based on the analysis results, returning these parameters to the UAV.

[0076] After receiving the spatial region of interest (ROI) and control parameters from the edge server, the UAV fuses the spatial ROI with the local semantic ROI to form a fused ROI. The spatial processing module then performs differentiated processing on the video frames based on the fused ROI, maintaining higher resolution for key areas and downsampling non-key areas to reduce redundant data transmission.

[0077] The processed video frames are re-encoded by the video encoding module and then sent back to the edge server. The edge server continues to perform decoding, bird's-eye view perception, spatial feedback generation, and control parameter updates. This cycle repeats, enabling adaptive transmission of UAV video and edge-side bird's-eye view perception even under dynamically changing wireless link conditions.

[0078] In one embodiment, when the wireless link bandwidth is high, the edge server can control the drone to improve video fidelity to obtain higher bird's-eye view accuracy; when the wireless link bandwidth is low, the edge server can control the drone to enhance the compression or downsampling of the background area to reduce transmission overhead and end-to-end latency.

[0079] In another embodiment, the system can record video bitrate, end-to-end latency, bandwidth changes, and perception performance during operation for subsequent system performance analysis and optimization.

[0080] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0081] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple steps or stages, which are not necessarily completed at the same time, but may be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but may be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0082] Based on the description of the above method embodiments, the present invention also provides a system. The system may be a system that uses software (applications), modules, components, servers, clients, etc., using the methods described in the embodiments of this specification, combined with necessary implementation hardware. Since the implementation schemes and methods for solving the problem are similar, the specific system implementations in the embodiments of this specification can be found in the implementations of the foregoing methods, and repeated details will not be described again. Although the system is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0083] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0084] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.

[0085] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. An edge-assisted UAV cooperative adaptive bird's-eye view perception method, characterized in that, Includes the following steps: S1, the drone captures video frames of the road scene and extracts semantic regions of interest; S2, the UAV receives the spatial region of interest and control parameters fed back from the edge server, and merges the spatial region of interest with the semantic region of interest to obtain the fused region of interest; S3, the drone performs differential processing on video frames based on control parameters and the region of interest for fusion. It maintains or increases the resolution of the region within the region of interest for fusion, and reduces the resolution of the region outside the region of interest for fusion. The differentially processed video frames are then encoded into video streams and uploaded to the edge server. S4, the edge server decodes the video stream, performs the bird's-eye view perception task, updates the spatial region of interest based on the obtained perception performance, analyzes the network status and perception performance, and decides the control parameters for the next moment. S5: The edge server feeds back the spatial region of interest and control parameters to the drone to adjust the video frame differentiation processing and encoding for the next moment. S6. Repeat steps S1 to S5 until the set termination condition is met.

2. The edge-assisted UAV cooperative adaptive bird's-eye view perception method according to claim 1, characterized in that, The fused region of interest is jointly determined by the semantic region of interest local on the UAV and the spatial region of interest fed back from the edge server. The semantic region of interest is used to indicate foreground targets, and the spatial region of interest is used to indicate background spatial regions that are important for bird's-eye view perception.

3. The edge-assisted UAV cooperative adaptive bird's-eye view perception method according to claim 1, characterized in that, The UAV terminal performs differentiated processing on video frames based on control parameters and fused regions of interest, specifically including: At time t, the original image frame is denoted as The image frames after differentiation processing are denoted as Then we have: ; in, This represents an adaptive space processing function; This represents video processing parameters and indicates the intensity of differentiated processing. This indicates the region of interest for merging.

4. The edge-assisted UAV cooperative adaptive bird's-eye view perception method according to claim 1, characterized in that, The analysis of network state and perception performance, and the determination of control parameters for the next moment, specifically include: Control parameters include video processing parameters and encoding parameters; network status includes video bitrate, end-to-end latency and available bandwidth; perception performance specifically includes target detection results, spatial layout information and perception performance indicators. end-to-end Latency includes drone-side encoding latency Wireless transmission latency Edge server-side decoding latency and edge server-side bird's-eye view latency perception : ; The process of encoding the differentiated video frames into a video stream and then uploading it to the edge server specifically includes: Let the video bitrate at time t be... : ; in, Represents the encoding function. These are encoding parameters used to control the quality of video frames; These are the image frames after differential processing.

5. The edge-assisted UAV cooperative adaptive bird's-eye view perception method according to claim 1, characterized in that, Also includes: To achieve a balance between perceived performance and network status, the edge server defines a revenue function for each time step. The perception performance at time t is denoted as End-to-end delay is denoted as Then the payoff function for: in, This represents the latency penalty weight; during continuous operation, maximizing the benefit is the optimization objective. ; in, As a discount factor, It expresses expectation.

6. An edge-assisted UAV cooperative adaptive bird's-eye view perception system, characterized in that, include: Drones: The drone acquires video frames of road scenes, extracts semantic regions of interest (ROIs), receives spatial ROIs and control parameters from the edge server, merges the spatial ROIs and semantic ROIs to obtain a fused ROI, and performs differential processing on the video frames based on the control parameters and the fused ROIs. The resolution of the regions within the fused ROIs is maintained or increased, while the resolution of the regions outside the fused ROIs is reduced. The differentially processed video frames are then encoded into a video stream and uploaded to the edge server. Edge server: Decodes video stream, performs bird's-eye view perception task, updates spatial region of interest based on the obtained perception performance, analyzes network status and perception performance, determines control parameters for the next moment, and feeds back spatial region of interest and control parameters to the UAV to adjust the video frame differentiation processing and encoding for the next moment.

7. The edge-assisted UAV cooperative adaptive bird's-eye view perception system according to claim 6, characterized in that, The fused region of interest is jointly determined by the semantic region of interest local on the UAV and the spatial region of interest fed back from the edge server. The semantic region of interest is used to indicate foreground targets, and the spatial region of interest is used to indicate background spatial regions that are important for bird's-eye view perception.

8. The edge-assisted UAV cooperative adaptive bird's-eye view perception system according to claim 6, characterized in that, The UAV terminal performs differentiated processing on video frames based on control parameters and fused regions of interest, specifically including: At time t, the original image frame is denoted as The image frames after differentiation processing are denoted as Then we have: ; in, This represents an adaptive space processing function; This represents video processing parameters and indicates the intensity of differentiated processing. This indicates the region of interest for merging.

9. The edge-assisted UAV cooperative adaptive bird's-eye view perception system according to claim 6, characterized in that, The analysis of network state and perception performance, and the determination of control parameters for the next moment, specifically include: Control parameters include video processing parameters and encoding parameters; network status includes video bitrate, end-to-end latency and available bandwidth; perception performance specifically includes target detection results, spatial layout information and perception performance indicators. end-to-end Latency includes drone-side encoding latency Wireless transmission latency Edge server-side decoding latency and edge server-side bird's-eye view latency perception : ; The process of encoding the differentiated video frames into a video stream and then uploading it to the edge server specifically includes: Let the video bitrate at time t be... : ; in, Represents the encoding function. These are encoding parameters used to control the quality of video frames; These are the image frames after differential processing.

10. The edge-assisted UAV cooperative adaptive bird's-eye view perception system according to claim 6, characterized in that, Also includes: To achieve a balance between perceived performance and network status, the edge server defines a revenue function for each time step. The perception performance at time t is denoted as End-to-end delay is denoted as Then the payoff function for: in, This represents the latency penalty weight; during continuous operation, maximizing the benefit is the optimization objective. ; in, As a discount factor, It expresses expectation.