EVTOL multi-camera cooperative video compression coding method based on multi-source perception and intelligent partitioning

Through the video compression and coding method of multi-source perception and intelligent partitioning, the problems of low multi-source data utilization and poor adaptability of dynamic scenes in eVTOL multi-camera video compression are solved, efficient video data transmission and stable video quality are achieved, and commercial application of eVTOL is supported.

CN120568072APending Publication Date: 2025-08-29CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510710137.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing eVTOL multi-camera video compression technology has problems such as low utilization of multi-source data, insufficient encoding efficiency, and poor adaptability to dynamic scenes, especially in mobile network environments, resulting in lag and delay in video transmission.

Method used

The video compression and coding method of multi-source perception and intelligent partition is adopted, and multi-modal sensor data is integrated through a multi-level collaborative processing architecture, combined with deep learning and dynamic partitioning technology to achieve efficient compression and transmission of video data.

Benefits of technology

Significantly reduce bandwidth requirements, improve the system's adaptability in complex scenarios, ensure the video quality in key areas, improve the system's adaptability and stability, and support the commercial application of low-altitude economy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120568072A_ABST
    Figure CN120568072A_ABST
Patent Text Reader

Abstract

The invention discloses an eVTOL multi-camera cooperative video compression coding method based on multi-source perception and intelligent partitioning. The method comprises the following steps: constructing a scene three-dimensional perception model by fusing multi-source data of visible light, infrared and depth sensors; the method comprises the following steps: realizing dynamic video partitioning based on motion vectors and semantic analysis, and dividing a picture into a core region, a secondary region and a background region; establishing a parallax compensation motion prediction model by adopting a cross-camera reference frame sharing mechanism; and high-fidelity compression of the key area is realized through layered entropy coding and a dynamic quantization parameter distribution strategy. And the decoding end reversely executes multi-source data fusion and partition reconstruction according to the coding metadata. The method is suitable for eVTOL multi-camera video real-time transmission scenes such as polling, surveying and mapping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of eVTOL video compression coding, and specifically relates to a multi-camera collaborative video processing and intelligent compression method. Background Art

[0002] With the continuous expansion of eVTOL application scenarios, multi-camera collaborative operation has become an inevitable trend in industry development. In professional fields such as power inspection, disaster monitoring, and border patrol, eVTOLs often need to simultaneously carry multimodal sensors such as visible light, infrared thermal imaging, and depth perception to achieve all-weather, multi-dimensional environmental perception. However, this multi-sensor configuration poses a huge challenge to the real-time transmission and processing of video data. Traditional single-camera independent encoding schemes cannot effectively utilize the temporal and spatial correlations between multiple perspectives, resulting in severe data redundancy and wasted bandwidth resources. Especially in mobile network environments, high-bitrate video transmission often suffers from problems such as lag and delay, seriously affecting operational efficiency.

[0003] In the existing technology, although some studies have attempted to solve the problem of multi-camera video compression, there are still many limitations. On the one hand, the solution with fixed coding parameters is difficult to adapt to the complex and changing scene requirements, which often leads to insufficient quality in key areas and waste of background area resources. For example, in power inspections, detailed information of key equipment such as insulators may be lost due to excessive compression, while background areas such as the sky occupy too much bit rate. On the other hand, most of the existing multi-camera collaborative coding methods only consider visible light video streams, and fail to fully utilize the complementary advantages of multi-source data such as infrared and depth. In addition, the traditional layered coding scheme has poor adaptability to dynamic scenes and is prone to quality fluctuations when the eVTOL moves quickly or the scene changes suddenly. Summary of the Invention

[0004] In response to key issues in existing eVTOL video compression technology, such as low multi-source data utilization, insufficient coding efficiency, and poor adaptability to dynamic scenes, the present invention proposes an eVTOL multi-camera collaborative video compression coding method based on multi-source perception and intelligent partitioning. This method innovatively integrates multimodal sensor data and intelligent analysis technology, and achieves efficient compression and transmission of video data through a multi-level collaborative processing architecture. Compared with traditional solutions, the present invention significantly reduces bandwidth requirements and improves the adaptability of the system in complex scenarios while ensuring video quality in key areas. This method overcomes the authenticity and safety problems of the immersive eVTOL driving experience, while solving the core challenges of cost control and large-scale implementation, and promoting the low-altitude economy to accelerate from the pilot stage to commercial application.

[0005] To this end, the present invention proposes an intelligent video compression method for a multi-camera eVTOL system, which combines multi-source sensing and dynamic partitioning technology to achieve efficient video encoding. The method specifically includes the following steps:

[0006] S1, the multi-source perception stage, includes the following steps:

[0007] S1.1: During data synchronization, a dedicated synchronization signal generator generates precise timing trigger pulses, coupled with a high-precision timestamp recording mechanism to ensure consistent timing across all sensor data. An adaptive time deviation compensation algorithm is also employed to effectively eliminate clock drift between sensors, laying a solid foundation for subsequent multi-source data fusion.

[0008] S1.2: The feature extraction process for visible light video utilizes a hybrid architecture that combines deep learning methods with computer vision technology. An improved lightweight convolutional neural network performs multi-level feature analysis on video content, enabling accurate recognition and location of scene objects. Furthermore, combined with advanced optical flow estimation algorithms, a scene motion vector field is constructed to accurately distinguish foreground moving objects from static background areas, providing a key basis for intelligent segmentation.

[0009] S1.3. Multi-source data fusion processing utilizes a three-level progressive fusion strategy. At the data level, spatial alignment of multi-source data is achieved through feature matching and coordinate transformation. At the feature level, the advantageous features of each sensor are extracted and fused. At the decision-making level, probabilistic graphical models are used for comprehensive reasoning to generate a scene understanding report containing rich semantic information. This layered and progressive fusion approach fully preserves the characteristic data of each sensor while achieving effective information complementarity.

[0010] S2, the intelligent partitioning stage, includes the following steps:

[0011] S2.1, dynamic region partitioning uses an adaptive grid partitioning strategy, dynamically adjusting the partitioning granularity based on scene complexity. The partitioning process comprehensively considers multiple dimensions of information, including target detection results, motion activity, temperature distribution characteristics, depth change trends, and visual saliency, to ensure that important areas receive sufficient coding resources. The partitioning results maintain the integrity of key targets while achieving optimal allocation of coding resources.

[0012] S2.2, encoding parameter configuration adopts a differentiated strategy design, formulating corresponding encoding schemes for areas of different importance. Core areas of interest use refined encoding parameter settings to retain rich details; secondary areas use a balanced encoding strategy to balance quality and efficiency; background areas use an efficient compression scheme to maximize bandwidth savings. This hierarchical configuration method significantly improves the utilization efficiency of encoding resources;

[0013] S2.3, partition metadata management, has designed a compact and efficient data structure, using advanced compression coding techniques to represent and store partition information. The metadata content includes key information such as region type identification and encoding parameter indication. By optimizing the storage structure, its proportion in the total bitrate is kept low, ensuring the complete transmission of partition information without significantly affecting the overall bitrate.

[0014] S3, the collaborative coding phase, includes the following steps:

[0015] S3.1, a cross-camera reference frame sharing mechanism, builds an intelligent frame buffer management system. It maintains a reference frame pool through advanced cache replacement strategies, calculates accurate disparity compensation based on depth information, and establishes a global motion model to enable collaborative prediction across multiple views. This sharing mechanism fully utilizes the spatiotemporal correlations between multiple cameras, significantly improving the accuracy and efficiency of inter-frame prediction.

[0016] S3.2, layered coding implementation uses a scalable video coding architecture, dividing video content into a base layer and an enhancement layer. The base layer ensures basic video legibility, while the enhancement layer focuses on improving the quality of key areas. Through a reasonable bitstream organization method, it supports flexible adjustment of decoding quality based on network conditions and terminal capabilities, significantly enhancing the adaptability and practicality of the system.

[0017] S3.3, dynamic bitrate control implements a complete adaptive adjustment solution. By monitoring network status in real time, evaluating channel quality trends, and dynamically adjusting encoding parameters in each region based on optimization theory, this intelligent bitrate allocation mechanism ensures that the system automatically maintains the optimal quality and bandwidth balance when network conditions fluctuate, providing users with a stable viewing experience.

[0018] The innovation of the present invention is mainly reflected in the following aspects:

[0019] (1) This paper proposes a perception framework that deeply integrates multi-source data, breaking through the limitations of traditional single-sensor information. Through the collaborative analysis of visible light, infrared, and depth data, a more comprehensive and accurate understanding of the scene is achieved;

[0020] (2) This paper designs a dynamic intelligent partitioning algorithm, which for the first time introduces temperature characteristics and spatial geometry information into the video partitioning decision process. This method can automatically adjust the partitioning strategy according to the scene content, significantly improving the utilization efficiency of coding resources;

[0021] (3) This invention develops a cross-camera collaborative encoding mechanism that fully utilizes the spatiotemporal correlation between multiple perspectives. Through parallax compensation and reference frame sharing technology, the accuracy of motion prediction is greatly improved;

[0022] (4) This invention builds a layered quality scalable coding architecture that supports flexible adjustment of video quality based on network conditions and terminal requirements. This architecture is particularly suitable for eVTOL video transmission in mobile network environments;

[0023] (5) This invention innovatively designs a dynamic bitrate allocation algorithm based on deep learning. By analyzing scene content complexity and network status changes in real time, it intelligently adjusts the encoding parameters of each video partition. This algorithm breaks through the limitations of traditional fixed bitrate allocation models, achieving an optimal match between encoding resources and network bandwidth, and maintaining video quality in key areas even in the event of sudden network congestion. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is an overall flow chart of the method of the present invention;

[0025] Figure 2 This is a flow chart of the multi-source perception stage of an embodiment of the present invention;

[0026] Figure 3 This is a flow chart of the intelligent partitioning stage of an embodiment of the present invention;

[0027] Figure 4 This is a flowchart of the collaborative encoding stage of an embodiment of the present invention; DETAILED DESCRIPTION

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0029] See also Figure 1 The present invention provides an eVTOL multi-camera collaborative video compression encoding method based on multi-source perception and intelligent partitioning, which includes a multi-source perception stage, an intelligent partitioning stage and a collaborative encoding stage.

[0030] The implementation process of the multi-source perception stage, intelligent partitioning stage and collaborative encoding stage is shown in the following table. Figures 2 to 4 , including the following:

[0031] S1, the specific implementation of the multi-source perception stage:

[0032] S1.1, specific implementation of data synchronization acquisition: The multi-sensor synchronization acquisition system of the present invention is implemented using a hierarchical synchronization architecture. At the bottom layer, an accurate hardware trigger signal is generated by the FPGA chip, and the signal is distributed to each sensor node through a dedicated synchronization bus. Each sensor node is equipped with a high-precision clock counter to record a precise timestamp when a trigger signal is received. In order to eliminate clock drift of different sensor nodes, the system uses a two-way time synchronization protocol (such as the PTP protocol) for periodic clock calibration. During the calibration process, the master node calculates the clock deviation by measuring the round-trip time of the signal, and uses the least squares method to compensate for the deviation, ultimately achieving microsecond-level time synchronization accuracy. For fixed delay differences caused by different transmission paths, the system automatically measures and stores the delay compensation value of each channel through benchmark testing during the initialization phase;

[0033] S1.2, Technical details of visible light feature extraction: Feature extraction of visible light video is implemented using an improved YOLOv5s network architecture. In terms of network design, the input resolution is first adjusted to a moderate size suitable for processing by the eVTOL platform, which not only ensures the accuracy of target detection but also controls the amount of calculation. The backbone of the network uses depthwise separable convolution instead of standard convolution, significantly reducing the number of parameters. A channel attention module is introduced at the neck of the network to obtain channel statistical features through global average pooling, and then generate channel weights through two layers of fully connected layers to enhance the expressive power of important features. The target detection head adopts a multi-scale prediction structure to fuse semantic information at different levels through a feature pyramid. In order to improve processing efficiency, the system implements a dynamic reasoning mechanism to automatically skip the calculation of some convolutional layers according to the complexity of the scene. The optical flow estimation module adopts a sparse optical flow method based on feature points. First, the FAST algorithm is used to detect feature points, and then the LK optical flow algorithm is applied to track the movement of feature points. Finally, the RANSAC algorithm is used to remove abnormal points and construct an accurate motion vector field;

[0034] S1.3, implementation method of multi-source data fusion: Data layer fusion adopts a feature point-based registration method. For visible light and infrared images, SIFT feature points are first extracted and feature descriptors are calculated, then feature matching is performed through the FLANN matcher, and finally the RANSAC algorithm is used to estimate the affine transformation matrix to achieve spatial alignment. The depth data is directly converted to a unified coordinate system through pre-calibrated camera parameters. A dedicated feature extraction network is designed for feature layer fusion. The visible light branch uses CNN to extract texture features, the infrared branch extracts temperature distribution features, and the depth branch calculates surface normal and curvature features. The features of each branch are fused through the splicing layer, and then the important features are screened through the feature selection network. The decision layer fusion constructs a Bayesian network model, which takes the target detection confidence, temperature anomaly probability, depth mutation degree, etc. as input nodes, calculates the comprehensive importance score of each region through the conditional probability table, and finally outputs a three-dimensional scene model with semantic annotations;

[0035] S2. Specific implementation of the intelligent partitioning stage:

[0036] S2.1, Implementation of dynamic area division: The dynamic grid division algorithm uses a quadtree structure to achieve adaptive partitioning. In the initial stage, the image is divided into uniform macroblocks, and then the complexity of each macroblock is evaluated: the maximum confidence of target detection within the block, the mean amplitude of the motion vector, the temperature variance, the depth gradient amplitude and other indicators are calculated, and the pre-trained classifier is used to determine whether further subdivision is needed. The subdivision process adopts a recursive method until the minimum partition granularity is reached or the stopping condition is met. In order to maintain the continuity of the partition boundary, the edge consistency constraint is implemented in the subdivision process to ensure that the difference in partition levels of adjacent blocks does not exceed the preset threshold. The final generated partition result is subjected to morphological post-processing to eliminate isolated small partitions;

[0037] S2.2, technical solution for encoding parameter configuration: differentiated encoding strategies are implemented through a hierarchical encoding controller. For core areas of interest, the encoding controller enables a complete set of encoding tools: including all 35 intra-frame prediction modes, motion estimation with 1 / 4 pixel accuracy, mode decision-making based on rate-distortion optimization, etc. A simplified configuration is used for secondary areas: the number of intra-frame prediction modes is limited, a fast motion estimation algorithm is used, and the rate-distortion optimization threshold is appropriately relaxed. A minimalist encoding scheme is used for the background area: only the DC prediction mode is retained, fixed prediction is used for the motion vector, and all enhancement encoding tools are turned off. The quantization parameters of each area are determined by a lookup table, which is generated based on the rate-distortion characteristic curve obtained through offline training to ensure the optimal quality distribution at the target bit rate;

[0038] S2.3, implementation details of partition metadata management: The metadata adopts a hierarchical coding structure. The top layer is a partition type distribution map, which is compressed using context-based arithmetic coding. The middle layer stores the quantization parameter differences of each partition, using differential pulse code modulation (DPCM) combined with Huffman coding. The bottom layer contains auxiliary information, such as motion prediction restriction flags, reference frame selection preferences, etc. In order to reduce the volume of metadata, the system implements an intelligent update mechanism: metadata updates are triggered only when the partition parameter changes exceed the threshold, and the static area adopts an inter-frame copy strategy. When metadata is embedded in the video stream, a decentralized insertion method is adopted to insert metadata fragments into the header of each video slice to ensure robustness in the event of transmission errors;

[0039] S3, the specific implementation of the collaborative coding stage:

[0040] S3.1, implementation of cross-camera reference frame sharing: the reference frame pool adopts a distributed management architecture. Each camera node maintains a local reference frame cache, and the central node maintains a global reference frame index table. When cross-camera reference is required, the requesting node first queries the index table to obtain the reference frame location information, and then obtains the frame data through the high-speed Internet. The disparity compensation algorithm adopts a hierarchical motion estimation strategy: first, global motion estimation is performed at the low-resolution layer to obtain the initial disparity vector; then local refinement is performed at the high level, using a six-parameter affine motion model. To improve efficiency, the system implements a motion vector prediction mechanism, uses depth map information to derive the disparity search range, and significantly reduces the number of search points. The reference frame quality assessment model comprehensively considers the three dimensions of PSNR, structural similarity, and decoding complexity, and selects the optimal reference frame through weighted scoring;

[0041] S3.2, the specific implementation of layered coding: The base layer coding adopts the traditional inter-frame prediction structure, but limits the use of coding tools to reduce complexity. The enhancement layer coding implements region-based enhancement technology: fine quantization is used for the core area to retain high-frequency details; selective enhancement is used for the secondary area to only strengthen important frequency components. The code stream organization adopts a flexible unit structure. Each coding unit contains three parts: basic data, enhancement data and metadata, which support independent decoding and combination. In order to achieve quality scalability, the system designs a multi-level reconstruction mechanism: the base layer decoder provides basic quality images, the enhancement layer decoder gradually improves the quality of key areas, and finally generates a complete image through the quality fusion module;

[0042] S3.3, Dynamic Rate Control Implementation: The network status monitoring module combines active probing with passive measurement. Active probing measures round-trip latency and packet loss by sending probe packets; passive measurement calculates the transmission latency and throughput of recent data packets. The rate allocation algorithm employs a Lagrangian optimization approach, modeling the rate allocation problem as a constrained optimization problem. The objective function comprehensively considers regional importance weights, quality smoothness, and buffer status. A parameter adjustment mechanism achieves a graded response: short-term fluctuations are accommodated through fine-tuning of quantization parameters; medium-term changes adjust the frame rate; and long-term bandwidth changes trigger resolution switching. To maintain quality stability, the system implements a proactive control mechanism that predicts encoding requirements in the coming seconds based on scene complexity and adjusts parameter settings in advance.

[0043] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. An eVTOL multi-camera collaborative video compression and encoding method based on multi-source perception and intelligent partitioning, applied to the field of eVTOL video processing, includes the following stages: Multi-source perception stage: collect and fuse multi-modal sensor data to obtain complete scene information; Intelligent partitioning stage: Dynamically divide video areas according to scene content and assign differentiated encoding parameters; Collaborative encoding stage: realize cross-camera reference frame sharing and layered bitstream generation. S1, the multi-source perception stage, includes the following steps: S1.1: Synchronously collect multi-source data from visible light cameras, infrared cameras, and depth sensors, establish a time alignment mechanism, and reduce the time deviation of each sensor data; S1.2, extract features from the visible light video stream to identify moving objects and static background areas in the scene; S1.3, integrate infrared data and depth information to build a 3D scene perception model and mark temperature anomaly areas and depth mutation boundaries; S2, the intelligent partitioning stage, includes the following steps: S2.1, based on the multi-source perception results, divide the video image into a core area, a secondary area, and a background area, where the core area contains moving objects and key scene elements; S2.2, assign dynamic quantization parameters (QP) to different areas, with QP22-26 used in the core area, QP28-32 in the secondary area, and QP35-40 in the background area; S2.3, establish a partition metadata table to record the region type and QP parameters of each 64×64 coding tree unit (CTU); S3, the collaborative encoding phase, includes the following steps: S3.1, build a cross-camera reference frame pool to allow cameras with different perspectives to share reference frame data; S3.2 uses a parallax compensation algorithm to process geometric differences between multiple views and achieve cross-camera motion prediction; S3.3, generating a layered code stream, including a base layer, an ROI enhancement layer, and a partition metadata layer, wherein the base layer contains a full-frame low-quality image, and the enhancement layer stores high-frequency components of the core area.

2. The method according to claim 1, wherein The multi-source perception stage ensures the temporal consistency of multi-sensor data acquisition through a hardware synchronization mechanism, adopts clock synchronization technology to eliminate time deviations between devices, and uses a feature matching algorithm to achieve spatial alignment of multi-source data. Finally, a scene perception model containing multi-dimensional information is constructed through data fusion technology, in which visible light sensors provide high-resolution images, infrared sensors supplement temperature features, and depth sensors provide spatial geometric information.

3. The method according to claim 1, wherein The collaborative coding stage realizes the sharing of coding resources among multiple perspectives, optimizes inter-frame prediction by establishing a global motion model and a parallax compensation mechanism, adopts a layered coding architecture to generate a code stream including a base layer and an enhancement layer, and dynamically adjusts the coding parameters of each layer to adapt to different network conditions and terminal requirements.

4. The method according to claim 2, wherein The hardware synchronization mechanism adopts a master-slave clock architecture, which ensures the consistency of the acquisition timing of each sensor through a dedicated synchronization signal, and cooperates with the software timestamp calibration algorithm to further eliminate clock deviation and achieve precise time alignment of multi-source data.

5. The method according to claim 3, wherein The deep learning algorithm adopts a lightweight network architecture, implements multi-scale analysis through a feature pyramid, introduces an attention mechanism to improve the recognition accuracy of key areas, and supports dynamic calculation path adjustment to adapt to scene content of different complexities.

6. The method according to claim 3, wherein The multi-view coding resource sharing mechanism selects the optimal reference frame through a reference frame quality assessment model, establishes a cross-view motion vector prediction method, and adopts an adaptive cache management strategy to maintain the reference frame pool, thereby improving coding efficiency.

Citation Information

Cited By

  • Infrared image low-delay image transmission and intelligent analysis system for unmanned platform

    CN121281196A

  • Method for encoding and decoding spatio-temporal data of three-dimensional space

    CN121327055A

  • A method for encoding and decoding space-time data of a three-dimensional space

    CN121327055B

  • Multi-source detection-oriented high-definition video compression method and system

    CN121691728A

  • A high-definition video compression method and system for multi-source detection

    CN121691728B