Unmanned aerial vehicle multichannel image transmission optimization system based on link quality perception

By building a multi-channel image transmission optimization system for link quality sensing, predicting future link stability and analyzing the changes in image frame content, the problems of lag in link state evaluation and insufficient image content scheduling in drone image transmission are solved, and transmission stability and image quality are improved.

CN120583239APending Publication Date: 2025-09-02SHENZHEN RUIWO MOBILE CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510791937.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

In the parallel transmission of multiple wireless links, the instantaneous evaluation of link status is easily affected by dynamic changes, and the lack of judgment on future link quality trends, resulting in lag or failure of scheduling results, and the lack of in-depth modeling of the changes in image content and timeliness requirements of image data scheduling and compression processes, resulting in problems such as delayed transmission of keyframes and occupation of high-quality channels.

Method used

Build a multi-channel image transmission optimization system for unmanned aerial vehicles based on link quality perception, including prediction modules, modeling modules, decision-making modules and packet modules. Through the timing prediction model, predict future link stability, analyze the changes in image frame content, build the optimal matching relationship between image frames and channels, and realize dynamic scheduling and compression optimization.

Benefits of technology

It improves the stability and image quality assurance capabilities of image transmission, effectively avoids keyframe loss or delay, optimizes transmission quality, reduces encoding redundancy, and realizes the dynamic linkage of the scheduling strategy driven by image content analysis results and link status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120583239A_ABST
    Figure CN120583239A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle multi-channel image transmission optimization system based on link quality perception so as to improve the transmission stability and the image quality guarantee capability of image data in a complex wireless environment. The prediction module constructs a time sequence model based on the historical link quality index of the wireless channel, and outputs a future link stability prediction value; the modeling module analyzes texture change and the like between the image frames, generates an image frame evolution vector, and calculates the aging sensitivity and reconstruction importance score of the image frames according to the image frame evolution vector; the decision-making module fuses the information and generates an optimal matching relation between the image frame and the channel and a transmission priority parameter; the grouping module performs image frame clustering according to the similarity between the priority parameters and the evolution vectors, constructs compression groups and generates corresponding compression configuration files and compression image data; the transmission module schedules the compressed image data to a corresponding wireless channel for transmission; the system can be widely applied to an unmanned aerial vehicle image transmission task with relatively high requirements on real-time performance and image quality in a dynamic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image transmission technology, and in particular to a multi-channel image transmission optimization system for unmanned aerial vehicles based on link quality perception. Background Art

[0002] Existing drone image transmission systems generally utilize multiple wireless links for parallel transmission. This divides image data and sends it to different wireless channels, enabling real-time backhaul and remote presentation. Some solutions evaluate the instantaneous link status of each channel, combining static image compression strategies with fixed scheduling rules to match and distribute image frames to transmission channels, thereby improving transmission efficiency and stability.

[0003] However, existing technologies generally have the following shortcomings: on the one hand, the instantaneous evaluation of link status is easily affected by dynamic changes in the wireless environment, and there is a lack of judgment on future link quality trends, which leads to delayed or invalid scheduling results; on the other hand, the scheduling and compression process of image data lacks in-depth modeling of the changing patterns and timeliness requirements of image content, which is prone to problems such as delayed transmission of key frames and non-key frames occupying high-quality channels, ultimately affecting the overall image quality and timeliness.

[0004] Therefore, it is necessary to propose a new image transmission optimization mechanism to more effectively coordinate the dynamic matching of image content features and link state information in a multi-channel wireless link environment. Summary of the Invention

[0005] This application provides a drone multi-channel image transmission optimization system based on link quality perception to improve the transmission stability and image quality assurance capabilities of image data in complex wireless environments.

[0006] This application provides a UAV multi-channel image transmission optimization system based on link quality perception, including: The prediction module is used to receive the link quality indicators of multiple wireless transmission channels currently connected to the drone in historical periods and build a time series prediction model; based on the time series prediction model, it obtains the link stability prediction value of each wireless transmission channel in the future target transmission period; A modeling module is configured to receive a sequence of image frames captured by a drone, analyze the texture change characteristics, edge motion trajectories, and semantic target consistency between adjacent image frames, and construct an image frame evolution vector that reflects the correlation between the rate of change of image frame content and its temporal sequence; based on the image frame evolution vector, calculate the time sensitivity score and reconstruction importance score of each image frame within the prediction time window; a decision module, configured to receive the link stability prediction value, the time sensitivity score, and the reconstruction importance score, determine the optimal matching relationship between the image frame and each wireless transmission channel by constructing a multi-dimensional weighted mapping function, and obtain a channel allocation instruction and corresponding transmission priority parameters for the image frame based on the optimal matching relationship; a grouping module configured to receive the image frame evolution vector and the transmission priority parameter of the image frame, perform content clustering on image frames having similar transmission priority parameters and image frame evolution vectors, and construct an inter-frame compression grouping structure; set a compression algorithm type, quantization parameter, and inter-frame prediction structure for each compression group based on the change rate information represented in the image frame evolution vector, and obtain a compression profile and corresponding compressed image data; The transmission module is used to receive the compression configuration file and the compressed image data, and distribute each compressed image data to the corresponding wireless transmission channel according to the transmission priority parameter according to the channel allocation instruction.

[0007] The beneficial effects of this application mainly include: (1) By constructing a time series prediction model to obtain the future stability trend of each wireless transmission channel, the channel resource can be proactively scheduled based on the prediction results, effectively avoiding the loss of key frames or transmission delays caused by instantaneous link fluctuations, and improving the anti-interference ability and stability of the overall image transmission system. (2) By modeling the evolution of texture, motion and semantic features between image frames, combined with the time sensitivity of image frames and reconstruction importance scores, high-value image frames can be targeted to high-reliability channels, optimizing the transmission quality of key areas and ensuring image fidelity and task perception effects. (3) Based on the content change rate of image frames and their scheduling priority, a cross-frame compression grouping structure is constructed, and differentiated compression parameters and prediction structures are configured for each group, so that similar frames share the compression strategy, while improving the compression ratio and retaining high-frequency frame details, reducing overall coding redundancy. (4) A data dependency and feedback mechanism is formed between the various modules of the system. The image content analysis results drive the scheduling strategy, the scheduling output controls the compression method, and the compression results determine the transmission path, thereby achieving full-process dynamic linkage and collaborative optimization between image semantics, compression strategy and link status. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 This is a schematic diagram of a drone multi-channel image transmission optimization system based on link quality perception provided in the first embodiment of the present application. DETAILED DESCRIPTION

[0009] The following description sets forth many specific details to facilitate a thorough understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present application. Therefore, the present application is not limited to the specific implementations disclosed below.

[0010] The first embodiment of the present application provides a multi-channel image transmission optimization system for drones based on link quality perception. Figure 1 , which is a schematic diagram of the first embodiment of the present application. Figure 1 The first embodiment of the present application provides a drone multi-channel image transmission optimization system based on link quality perception.

[0011] The UAV multi-channel image transmission optimization system based on link quality perception includes a prediction module 101, a modeling module 102, a decision module 103, a grouping module 104 and a transmission module 105.

[0012] The prediction module 101 is used to receive the link quality indicators of multiple wireless transmission channels currently connected to the drone in a historical period and build a timing prediction model; based on the timing prediction model, it obtains the link stability prediction value of each wireless transmission channel in the future target transmission period.

[0013] The core function of prediction module 101 is to construct a time series prediction model suitable for multi-channel UAV image transmission environments based on historical link status data. Based on this model, it outputs a predicted link stability value for each wireless transmission channel within a future target transmission period, providing a forward-looking reference for image frame scheduling. This module's input is the link quality indicators of the multiple wireless transmission channels currently connected to the UAV over a historical period, and its output is the corresponding link stability prediction value for each wireless transmission channel. The entire module's implementation can be divided into four main stages: data acquisition, feature processing, time series modeling, and prediction output.

[0014] First, during the data collection phase, prediction module 101 reads the link status logs of the multiple wireless transmission channels currently connected to the drone from the physical or MAC layer via a communication interface. These link status logs include, but are not limited to, the following key metrics: packet loss rate per unit time, signal-to-noise ratio, and round-trip delay jitter. These metrics are typically reported periodically by the data link or physical layer of the underlying communication protocol stack and can be centrally collected and cached by software middleware. The system can set a historical time window, such as data from the last 5 to 30 seconds, to serve as the input data set for subsequent modeling.

[0015] During the feature processing phase, prediction module 101 preprocesses the collected raw link quality metrics, including time alignment, gap filling, outlier correction, and normalization. Time alignment synchronizes data from different channels to a common sampling time base. Gap filling can be achieved using linear interpolation or a sliding mean. Outlier correction can be achieved using IQR (interquartile range) elimination or the Z-score method to eliminate abnormal points. Normalization maps each metric to a range of [0, 1] to eliminate dimensional differences. After processing, each channel corresponds to a set of continuous time series samples.

[0016] Subsequently, during the time series modeling phase, prediction module 101 establishes a time series prediction model for each wireless transmission channel. This model can employ traditional statistical methods such as the AR (autoregressive) model and the ARIMA (autoregressive integrated moving average) model, or lightweight machine learning models such as support vector regression (SVR) or long short-term memory (LSTM) networks. LSTM is particularly well-suited for capturing the long-term dependencies and mutational characteristics of link quality indicators, providing superior modeling results compared to standard linear regression. The modeling process consists of two phases: training and validation. The system utilizes a sliding window mechanism to continuously update training samples and retrain or fine-tune model weights at a set interval (e.g., once per second) to adapt to complex and changing flight environments and wireless conditions.

[0017] Finally, in the prediction output stage, the prediction module 101 outputs the link stability prediction value within the future target transmission period based on the constructed timing model and the current state of each channel. The target transmission period can be set to 1 to 2 seconds to cover the decision window for the system image frame scheduling. The prediction value is in the form of a scalar, usually a value between [0,1]. The closer to 1, the more stable and reliable the channel is within the prediction period. This stability prediction value will serve as one of the weighted criteria for the subsequent decision module 103 to match image frames with channels. The prediction module 101 can simultaneously support parallel prediction of multiple channels, and the results are provided to the upper-level modules of the system in the form of vectors for comprehensive scheduling.

[0018] For example, in one embodiment, prediction module 101 uses an autoregressive moving average (ARMA) model as a time series prediction model for short-term prediction of the signal-to-noise ratio (SNR) of wireless transmission channels. The historical time window is set to the most recent 10 seconds, and the system records the SNR values ​​of each wireless channel at a sampling rate of once per second. Therefore, the historical data series input to the model for each channel is 10 sample points long. The system first performs a unit root test on this 10-sample SNR time series. If the series is stationary, it can be directly input into the ARMA model for modeling. The model order (p, q) can be automatically selected using the AIC or BIC criteria. In this example, ARMA(2,1) is selected, indicating that the model contains two autoregressive terms and one moving average term.

[0019] In actual operation, prediction module 101 reads the historical SNR values ​​from seconds t-9 to t to form an input sequence X = [SNR(t-9), ..., SNR(t)]. After inputting this into the ARMA(2,1) model, the output is the predicted value Y for SNR(t+1), which predicts the SNR of the current channel in the next second (t+1). This predicted value serves as one of the link stability indicators for the channel within the 1-second prediction window. Together with other indicators (such as the packet loss rate prediction), it forms a comprehensive link stability prediction for the channel.

[0020] Assume that the signal-to-noise ratio (SNR) values ​​for a particular channel over the past 10 seconds at the current time t are: [17.2, 17.5, 17.1, 16.8, 17.0, 17.3, 17.6, 17.4, 17.2, 17.3] dB. The system builds an ARMA(2,1) model based on this sequence. After fitting the model parameters using the least squares method, the predicted SNR for this channel at t+1 second is calculated to be 17.5 dB. Based on the configured link stability mapping function, the system normalizes this predicted SNR value to a stability score of 0.88 (out of a maximum score of 1.0). This score is then used as the output of prediction module 101 and passed to subsequent decision module 103 for image frame scheduling and channel matching.

[0021] In summary, prediction module 101, through the collection, preprocessing, time-series modeling, and prediction output of historical link quality indicators, forms a practical, quantifiable, and continuously updated link stability assessment mechanism. This provides forward-looking scheduling support for image transmission tasks and improves the adaptability and robustness of the entire system in multi-channel environments. This module does not rely on external servers or high-performance computing platforms and can be implemented locally on the drone as a lightweight model, making it suitable for real-time image transmission scenarios.

[0022] The modeling module 102 is used to receive a sequence of image frames collected by a drone, analyze the texture change characteristics, edge motion trajectory and semantic target consistency between adjacent image frames, and construct an image frame evolution vector that reflects the rate of change of image frame content and the temporal correlation; based on the image frame evolution vector, calculate the time sensitivity score and reconstruction importance score of each frame image within the prediction time window.

[0023] Modeling module 102's function is to establish a vectorized representation of the image's dynamic features based on the content evolution process between image frames. This representation then generates a temporal sensitivity score and a reconstruction importance score for each image frame within a specific prediction time window, providing content-level quantitative support for the subsequent formulation of image frame scheduling and compression strategies. This module's input is a sequence of image frames captured by a drone, and its output is an image frame evolution vector and the temporal sensitivity score and reconstruction importance score calculated from it. The entire module's implementation process primarily includes four sequential processing steps: image frame preprocessing, feature differential extraction, evolution vector construction, and score calculation.

[0024] During the image frame preprocessing stage, modeling module 102 first performs a unified format conversion and size normalization on the raw image frame sequence captured by the drone, ensuring that subsequent feature analysis operations are performed on structurally consistent data. Specifically, each image frame is converted to grayscale or three-channel RGB format and uniformly resized to a fixed resolution, such as 640×480 pixels. Furthermore, if the image sequence has uneven acquisition intervals or fluctuating frame rates, frame time alignment is performed to ensure consistent time intervals between adjacent frames.

[0025] During the feature difference extraction stage, the modeling module 102 calculates three key feature differences for each pair of consecutive image frames, including texture change features, edge motion trajectories, and semantic target consistency. Texture change features can be obtained by calculating the gradient distribution difference, directional entropy, or LBP (local binary pattern) histogram change rate in local image regions, reflecting changes in local complexity at the image detail level. Edge motion trajectories are obtained by performing Canny edge detection on the image and calculating the displacement vector of edge points between frames, further extracting statistical information such as the global average motion amplitude and local displacement direction distribution. Semantic target consistency can be achieved by identifying salient object regions in each frame using an object detection algorithm (such as YOLO or MobileNet-SSD) and calculating their intersection over union (IoU), center offset, and scale change rate between adjacent frames to measure the continuity of semantic regions in the temporal dimension.

[0026] After extracting the above three types of features between each frame, modeling module 102 traverses the entire image frame sequence using a sliding window. For each frame, a composite vector is constructed that contains the differences between it and the preceding and following frames, forming an image frame evolution vector. This evolution vector is a multidimensional feature set, with dimensions including but not limited to statistics such as texture change intensity, edge motion amplitude, semantic consistency score, and standard deviation of inter-frame displacement direction. Each image frame corresponds to a set of vectors that characterize the content change trend and degree of semantic change.

[0027] Based on the image frame evolution vector, modeling module 102 further introduces a scoring function to calculate the temporal sensitivity score and reconstruction importance score of the image frame within the prediction time window. The temporal sensitivity score is used to determine whether the image frame should be transmitted first. Its calculation criteria include, but are not limited to, the intensity of inter-frame motion (e.g., the score increases when the motion amplitude exceeds a set threshold), the significance of semantic changes (e.g., the displacement or disappearance of the detected object), and the time interval between the key frame and the frame. This score is typically a real number between 0 and 1. Higher values ​​indicate a frame with greater sensitivity to temporal response and should be prioritized for transmission. The reconstruction importance score measures the contribution of the image frame to restoring visual coherence and semantic integrity in the image sequence. Its evaluation criteria include the degree of semantic structure difference between the current frame and its neighboring frames, whether it contains independent object regions, and whether it is at a transition point in image continuity. The score is also a number between 0 and 1. Higher values ​​indicate a frame with greater irreplaceability in the reconstruction process at the receiving end, and higher compression quality or greater error tolerance should be configured.

[0028] In one specific embodiment, modeling module 102 receives a sequence of video frames captured by a drone at 30 frames per second. The system processes the image frame sequence using a sliding window approach, where a three-frame window is formed, consisting of the current frame and the frames before and after it. The system then extracts the change features between the current frame and adjacent frames to construct an image frame evolution vector.

[0029] Assume that in a sequence of images captured by a drone, frame 75 is the current frame. The modeling module extracts the changes between frames 74, 75, and 76, primarily including the degree of texture change, edge motion, semantic object changes, and coverage changes.

[0030] After image preprocessing, the module first analyzes the degree of texture change. In frame 75, the background changes from the clearing in the previous frame to a complex forested area. Using the gray-level co-occurrence matrix, the system calculates an image energy difference of 0.24 and a directional entropy difference of 0.16, indicating significant texture change. Therefore, the system assigns a texture change score of 0.7 (out of a maximum score of 1.0) for this category.

[0031] Next, the module detected edge changes, using the Canny algorithm to extract edges and calculate edge point displacement. The average edge displacement reached 11.8 pixels, and the variance of the edge displacement direction was 0.72, indicating significant global motion. Based on the configured mapping, the system assigned an edge motion score of 0.8 for this frame.

[0032] The system then used YOLOv5 to detect semantic objects in the image. Two people were detected in frame 74, while a running object was added in frame 75. The object remained in frame 76. The change in Intersection over Union (IoU) between the three frames was 0.35, and the area occupied by the person increased from 12% in the previous frame to 18% in the current frame. Based on this, the system assigned a semantic consistency score of 0.6, an object number change score of 0.9, and an object coverage change score of 0.8.

[0033] The module combines these scores into the image frame evolution vector of the current frame, namely: [0.7 (texture change), 0.8 (edge ​​motion), 0.6 (semantic consistency), 0.9 (number of object changes), 0.8 (object coverage change)].

[0034] Based on the evolution vector, the modeling module begins calculating the temporal sensitivity score. The system configures a linear weighting function, assigning higher weights to edge motion, target number change, and target coverage change, at 0.35, 0.3, and 0.2, respectively. The other dimensions have a combined weight of 0.15. Substituting this vector into the scoring function yields the temporal sensitivity score for the current frame: 0.7×0.05 + 0.8×0.35 + 0.6×0.1 + 0.9×0.3 + 0.8×0.2 = 0.765; This indicates that the frame changes quickly and should be transmitted first.

[0035] The module then calculates a reconstruction importance score. This score focuses on the amount of independent information in a frame and its support for subsequent images. The system assigns higher weights to texture change, semantic consistency, and object coverage change, at 0.3, 0.3, and 0.25, respectively, while the remaining dimensions receive 0.15. Substituting the same evolution vector, the reconstruction importance score is calculated as: 0.7×0.3 + 0.8×0.05 + 0.6×0.3 + 0.9×0.1 + 0.8×0.25 = 0.74; This value indicates that the current frame is critical to the content continuity and semantic understanding of subsequent images, so high quality needs to be retained during compression.

[0036] Ultimately, the modeling module outputs the image frame's evolution vector, a temporal sensitivity score of 0.765, and a reconstruction importance score of 0.74, which are then fed to the decision-making and compression modules. This processing approach ensures that both scheduling and compression decisions are quantitatively informed by image changes and is suitable for the computationally constrained environment of real-time drone image transmission.

[0037] Through the above-described process, modeling module 102 can efficiently and stably generate a set of image frame evolution vectors for each frame in the image sequence, representing its content variation trends. Based on these vectors, it can calculate scoring metrics to support transmission scheduling and compression parameter configuration. The data output by this module directly serves as input to decision module 103 and grouping module 104, ensuring that the system fully considers the coupling between content variation characteristics and transmission timeliness requirements when processing image frames, thereby effectively improving the performance, adaptability, and image quality assurance of the entire image transmission system.

[0038] Furthermore, the modeling module is specifically used to: After receiving a sequence of image frames captured by a drone, for each image frame, the texture gradient variability, edge displacement direction variance, target number change rate, and semantic area IoU change rate between the image frame and its preceding and following adjacent frames are extracted. The extracted values ​​are combined to form an image frame evolution vector, which is used to characterize the rate of change of image frame content and the temporal correlation. Based on the image frame evolution vector, constructing a time-domain weighted vector including a time-attenuation weight, wherein the time-attenuation weight is set according to the frame spacing of the image frame relative to the current frame, and is used to perform temporal weight modulation on each dimensional feature in the image frame evolution vector to obtain a time-weighted image frame evolution vector; Inputting the time-weighted image frame evolution vector into a temporal sensitivity scoring function, the scoring function calculates the temporal sensitivity score of the image frame within the prediction time window based on the edge displacement direction variance and the semantic region IoU change rate, which is used to quantify the image frame's tolerance to transmission delay; The time-weighted image frame evolution vector is input into a reconstruction importance scoring function. The reconstruction importance scoring function uses a nonlinear function to calculate the reconstruction importance score of the image frame based on the trend changes of the target quantity change rate and the semantic area IoU change rate in multiple consecutive image frames. The score is used to evaluate the contribution of the image frame to the restoration of the image sequence structure. The time sensitivity score and the reconstruction importance score are provided to the decision module as output results of the modeling module.

[0039] During implementation, after receiving a sequence of image frames captured by a drone, the modeling module first uses the current image frame as the reference frame and selects several adjacent frames before and after it as analysis targets. Typically, a sliding window approach is used, with two frames before and after each, to construct a five-frame analysis window. Within this analysis window, for each target image frame, the system performs the following feature extraction operations: First, texture gradient variability is extracted. This feature measures the temporal instability of the image's texture structure. Specifically, a gradient magnitude image is generated from the gradient value of each pixel in the image frame (for example, using the Sobel operator). The gradient image difference of the current frame is then calculated using the frame mean within the window as a reference. This difference is defined as the standard deviation of the pixel gradient magnitude difference between the current frame and the adjacent frames, reflecting the degree of abrupt change in the texture structure.

[0040] The texture gradient variability is calculated as follows: First, the current image frame is converted to grayscale. The Sobel operator is then used to extract its horizontal and vertical gradient images. For each pixel, a gradient magnitude map is calculated by taking the square root of the sum of the squares of the horizontal and vertical gradients at each pixel. This process is repeated for each of the five frames preceding and following the current frame, generating the corresponding gradient magnitude map. Next, the pixel-by-pixel difference in gradient magnitude between the current frame and the other four frames at the same location is calculated. The standard deviation of these differences is then used as the texture gradient variability value for that frame, reflecting the degree of texture variation relative to adjacent frames.

[0041] Next, we extract the variance of edge displacement directions. We perform Canny edge detection on the image frames to obtain a set of edge pixels. We then use optical flow to estimate the displacement directions of edge points in the temporal dimension. Finally, we calculate the variance of these displacement directions in angular space. A larger variance indicates complex, multi-directional motion in the image.

[0042] The calculation method of the edge displacement direction variance is as follows: First, the Canny edge detection algorithm is used on the current image frame to extract all edge pixels. Then, the current frame and the previous frame are paired together. Sparse motion estimation is performed on the edge pixels in the current frame using an optical flow method (such as the Lucas-Kanade algorithm). The displacement direction of each edge point is obtained, that is, the angle of motion of the point between the two frames. After collecting the motion directions of all edge points, the dispersion of these angles is calculated. More dispersed directions indicate a greater presence of moving objects or camera disturbances in different directions in the scene, while more concentrated directions indicate more consistent motion. The dispersion of these angles is used as the edge displacement direction variance, which represents the complexity of the motion direction of the edge region in the current frame.

[0043] At the same time, the target number change rate is extracted, and the target detection network (such as YOLO) is used to count the changes in the number of effective targets detected in consecutive image frames, and the change rate of the current frame is defined as: ; This formula calculates the rate of change of the number of objects detected in an image frame over time, reflecting the relative change in the number of objects detected in the scene in the current frame relative to the previous frame. This value can be used to determine whether significant events such as an increase (e.g., the entry of a new object) or a decrease (e.g., the departure of an object) have occurred in the scene, thus providing a basis for determining the importance of reconstruction in subsequent image scheduling strategies.

[0044] in, Indicates the number of targets detected in the current frame, Indicates the number of targets in the previous frame. is the target quantity change rate, indicating the current image frame Relative to the previous frame The relative percentage change in the number of targets detected.

[0045] Further extract the semantic region loU change rate, calculate the intersection over union (IoU) between each detected target in the current frame and the bounding box of the same target in the previous frame, and average the IoU of all targets. Let the current frame be , the previous frame is , then the rate of change is: ; This formula calculates the relative rate of change of the average Intersection over Union (IoU) of semantic regions within a time series of image frames. It measures the magnitude of the change in the position of the object detection box between the current and previous frames, reflecting whether the object appears, moves, disappears, or undergoes shape change. It is a key metric for assessing structural changes in image frames. is the semantic region IoU change rate, reflecting the image frame Relative to the previous frame The relative change rate of the average IoU of the detected objects in . is the average IoU between each target in the current frame and the target in the previous frame. is the average matching IoU of the target in the previous frame.

[0046] The above four features are combined into the image frame evolution vector: ; in Represents texture gradient variability; represents the variance of edge displacement direction; Indicates the target quantity change rate; Indicates the semantic region loU change rate.

[0047] In order to introduce time series weights, the system introduces time domain weighted processing to the evolution vector, and sets the time window center as , including the two frames before and after, a total of five frames, the weight function is defined as: ; This formula is a time-domain weighting function that assigns different weights to the eigenvalues ​​of image frames from different time points (i.e., different frame intervals) in the image frame evolution vector. This emphasizes the importance of neighboring frames and reduces the influence of distant frames on the modeling results. Its core purpose is to simulate the "temporal memory decay" mechanism: information from adjacent frames has the greatest impact, and the greater the distance, the lower the weight. This weighting strategy is suitable for modeling feature evolution over short periods of time in image sequences, and is particularly well-suited for target continuity analysis, predictive stability modeling, and input enhancement of scoring functions in drone vision.

[0048] The absolute value of the frame distance, indicating the distance between the target frame currently being processed and the reference frame (usually the intermediate frame ). Let the reference frame be , the target frame is ,but ; For example, if the current frame is 100, the distance between the 98th frame and it is .

[0049] is the attenuation coefficient, usually set between 0.5 and 1.0.

[0050] is the time decay weight value.

[0051] The weighted image frame evolution vector The dimensions are defined as: ; Obtained This is the image frame evolution vector corrected by time decay, which is used for subsequent scoring. This formula is used to calculate the first Dimensions in the current frame The purpose is to perform "time series smoothing" on the feature values ​​of the current frame. By fusing the information of the previous and next frames, the feature expression is made more continuous and stable, and the ability to identify abnormal changes is enhanced.

[0052] result Represents: image frame In the The weighted value of the five-frame time window in each dimension, combined with the time domain attenuation weight, will be used as the input of subsequent scoring functions (such as the time sensitivity scoring function and the reconstruction importance scoring function).

[0053] is the time decay weight, which is calculated according to the above weight function calculation formula. Represents a frame Middle The original value of the feature dimension; Next, the weighted image frame evolution vector is input into the temporal sensitivity scoring function. The temporal sensitivity scoring function is calculated based on the edge displacement direction variance and the loU change rate, and is defined as: ; This formula is used to calculate the image frame Time sensitivity score This evaluates the image frame's sensitivity to transmission delays. A higher score indicates that delayed transmission of the frame could significantly impact system target recognition, scene reconstruction, or task response. Therefore, the frame should be assigned to a wireless channel with higher link quality and given priority for transmission.

[0054] The input of this function comes from the two dimensions of the image frame evolution vector and has been time-weighted, that is, and , respectively reflecting the inter-frame motion complexity and semantic region instability. They are calculated according to the following formulas: ; ; is the edge displacement direction variance, representing the image frame The variance of the edge displacement direction in . is the average IoU of the semantic target of the current frame, indicating the frame The average IoU between all detected objects in the frame and the corresponding objects in the previous frame.

[0055] in, and is the weighting coefficient, the recommended value is .

[0056] Then, the same weighted vector is input into the reconstruction importance scoring function, which uses a nonlinear form to express the importance of the target change trend: ; This formula is used to calculate the image frame Reconstruction importance score of , which measures whether a frame is located at a point with sudden semantic structure changes, densely changing objects, or a key frame within the entire image sequence. This allows for prioritization of this frame when transmission bandwidth is limited, ensuring semantic continuity and structural restoration capabilities at the receiver. Unlike the temporal sensitivity scoring function, the reconstruction importance score focuses more on the structural contribution of the image content itself to subsequent tasks such as image sequence decoding, splicing, and semantic analysis.

[0057] is the weighted target quantity change rate, indicating the frame The weighted average of the target quantity change within the time window; the calculation formula is as follows: ; is the target quantity change rate, expressed in image frames In the , the number of targets is relative to the previous frame rate of change; is the first-order difference of the weighted IoU change rate, which indicates the degree of mutation of the time-weighted semantic region IoU change rate sequence in the current frame. The calculation formula is as follows: ; is the average IoU of the semantic target of the current frame, indicating the frame The average IoU between all detected objects in the frame and the corresponding objects in the previous frame.

[0058] is the amplification factor, the recommended value is 2; 、 are weighting coefficients, and the recommended values ​​are 0.6 and 0.4 respectively.

[0059] , for example, setting ; The above scoring function outputs two quantities: : Time sensitivity score, used to indicate the sensitivity of the image frame to transmission delay; : Reconstruction importance score, which is used to indicate the contribution of the image frame to the continuity of the semantic structure in the entire image sequence.

[0060] Finally, the modeling module outputs the above two scoring results to the decision module to construct the optimal matching relationship between the image frame and each wireless transmission channel.

[0061] The decision module 103 is used to receive the link stability prediction value, the time sensitivity score and the reconstruction importance score, and determine the optimal matching relationship between the image frame and each wireless transmission channel by constructing a multi-dimensional weighted mapping function; based on the optimal matching relationship, obtain the channel allocation instruction of the image frame and the corresponding transmission priority parameter.

[0062] Decision module 103 comprehensively processes the data output by prediction module 101 and modeling module 102. Based on the link stability prediction value, time sensitivity score, and reconstruction importance score, it constructs a reasonable and effective image frame scheduling strategy, and generates channel allocation instructions and transmission priority parameters for the image frames. This module is a key link in the fusion of image content perception and network status perception. Its core task is to establish a quantitative mapping mechanism from multidimensional inputs to scheduling decisions, ensuring that each image frame is assigned to the most appropriate wireless transmission channel and has a priority identifier that can be used for scheduling.

[0063] In a typical implementation, decision module 103 first synchronously receives the link stability prediction value from prediction module 101. This stability prediction value estimates the availability and reliability of each wireless transmission channel within the next transmission cycle, typically expressed as a numerical score, for example, between 0 and 1, with higher values ​​indicating more stable channel transmission quality. Simultaneously, modeling module 102 provides two scores for each image frame: a time sensitivity score indicating the image frame's sensitivity to transmission delay, and a reconstruction importance score reflecting the image frame's structural value in restoring the image sequence.

[0064] The decision module 103 jointly models the above three types of scoring information and constructs a multi-dimensional weighted mapping function. The input dimensions of this function include the time sensitivity score of the image frame, the reconstruction importance score, and the link stability prediction value of each wireless channel. For specific implementation, the module establishes a scoring matrix, in which the rows represent image frames and the columns represent optional channels. For each element in the matrix, the system calculates a comprehensive matching score based on the score of the image frame and the link prediction value of the channel. The matching score can be calculated by a linear weighted method. For example, the time sensitivity weight is set to 0.4, the reconstruction importance weight is set to 0.3, and the link stability weight is set to 0.3. The corresponding multiplication and summation are performed to obtain the degree of adaptation of the image frame on the channel.

[0065] The decision module traverses all combinations of image frames and all channels to form a complete frame-channel matching score table. Then, under the premise of satisfying the maximum matching principle, the module uses a greedy method or a local optimal strategy to assign a unique target transmission channel to each image frame. If there are multiple channels with similar scores, the one with higher link stability is prioritized to ensure the success rate of transmission. At the same time, the system assigns a transmission priority parameter to each image frame based on its comprehensive score and its relative position in the entire image sequence. The priority parameter can be encoded in an ascending integer form, with smaller values ​​indicating higher priority. For example, image frames with high importance and timeliness will be assigned priority 1 or 2, while image frames with less content changes and tolerance for delays may be assigned priority 4 or 5.

[0066] Ultimately, decision module 103 outputs two results: a channel allocation instruction for the image frame, which explicitly identifies the wireless channel over which each image frame should be transmitted; and a transmission priority parameter for the image frame, which is used to schedule and allocate bandwidth based on priority in scenarios where multiple channels are transmitted simultaneously. These two outputs are directly passed to grouping module 104 and transmission module 105, ensuring that the system maintains efficient and reliable image transmission performance even in environments with dynamically changing wireless resources and varying image content.

[0067] To further illustrate the specific implementation of the decision module 103, the following uses a clear example to demonstrate its complete workflow, including how to construct a multidimensional weighted mapping function, how to determine the optimal matching relationship between the image frame and each wireless transmission channel based on the score, and ultimately generate channel allocation instructions and transmission priority parameters.

[0068] Assume that a drone is currently performing a city road monitoring mission. The system is processing three frames of images, numbered Frame A, Frame B, and Frame C. There are three available wireless transmission channels, Channel 1, Channel 2, and Channel 3. Prediction module 101 has output link stability prediction values ​​for the three channels within the next 1-second prediction window: 0.92, 0.75, and 0.60, respectively. Higher values ​​indicate more reliable transmission. Simultaneously, modeling module 102 calculates the following scores for each of these three frames: The time sensitivity score of frame A is 0.85, and the reconstruction importance score is 0.70; the time sensitivity score of frame B is 0.50, and the reconstruction importance score is 0.80; the time sensitivity score of frame C is 0.40, and the reconstruction importance score is 0.45.

[0069] The decision module 103 uses the following linear weighted mapping function to calculate the matching score of the image frame on each channel: matching score = 0.4 × temporal sensitivity score + 0.3 × reconstruction importance score + 0.3 × channel stability prediction value.

[0070] According to this function, the system calculates the matching score between each frame image and each channel as follows: Frame A: Channel 1: 0.4×0.85 + 0.3×0.70 + 0.3×0.92 = 0.34 + 0.21 + 0.276 = 0.826; Channel 2: 0.4×0.85 + 0.3×0.70 + 0.3×0.75 = 0.34 + 0.21 + 0.225 = 0.775; Channel 3: 0.4×0.85 + 0.3×0.70 + 0.3×0.60 = 0.34 + 0.21 + 0.18 = 0.73; Frame B: Channel 1: 0.4×0.50 + 0.3×0.80 + 0.3×0.92 = 0.20 + 0.24 + 0.276 = 0.716; Channel 2: 0.4×0.50 + 0.3×0.80 + 0.3×0.75 = 0.20 + 0.24 + 0.225 = 0.665; Channel 3: 0.4×0.50 + 0.3×0.80 + 0.3×0.60 = 0.20 + 0.24 + 0.18 = 0.62; Frame C: Channel 1: 0.4×0.40 + 0.3×0.45 + 0.3×0.92 = 0.16 + 0.135 + 0.276 = 0.571; Channel 2: 0.4×0.40 + 0.3×0.45 + 0.3×0.75 = 0.16 + 0.135 + 0.225 = 0.52; Channel 3: 0.4×0.40 + 0.3×0.45 + 0.3×0.60 = 0.16 + 0.135 + 0.18 = 0.475; Based on the aforementioned scoring matrix, the system selects the optimal match, assigning each frame to the channel with the highest matching score among all channels, ensuring that the same channel is not occupied by multiple frames simultaneously. In this example, frame A has the highest matching score with channel 1, 0.826, so frame A is assigned to channel 1. Frame B's matching score for channel 2 is slightly lower than channel 1, but since channel 1 is already occupied by frame A, frame B chooses channel 2 (score of 0.665). Frame C has channel 3 available, with a corresponding score of 0.475.

[0071] Next, the system prioritizes the image frames based on their combined scores. Priority is determined by the matching score, with higher values ​​indicating higher priority. The matching score ranking is: Frame A (0.826) > Frame B (0.665) > Frame C (0.475), resulting in corresponding priority parameters of 1, 2, and 3, respectively.

[0072] Ultimately, the decision module 103 outputs the following results: the channel assignment instruction for frame A is channel 1, with a priority parameter of 1; the channel assignment instruction for frame B is channel 2, with a priority parameter of 2; and the channel assignment instruction for frame C is channel 3, with a priority parameter of 3. This result is transmitted to the grouping module 104 and the transmission module 105 for image compression configuration and parallel scheduling execution.

[0073] From the above specific examples, it can be seen that the decision module 103 completes the intelligent scheduling between image frames and multi-channel resources with a clear input structure, a computable weight function and a determined priority allocation mechanism, ensuring the dynamic coupling of image content value and link status, and enabling the system to have stable and efficient image transmission capabilities in complex wireless environments.

[0074] Furthermore, the decision module is specifically used to: receiving the time sensitivity score and the reconstruction importance score output by the modeling module, and the link stability prediction value of each wireless transmission channel in the future target transmission period output by the prediction module; Based on the timeliness sensitivity score, reconstruction importance score, and link stability prediction value, a multidimensional weighted mapping function is constructed. The multidimensional weighted mapping function adopts a dynamic weight adjustment mechanism that is adaptive to the task type. The weight of the timeliness sensitivity score is adjusted according to the current task type. When the task type is an object recognition task, the weight of the timeliness sensitivity score related to the semantic region IoU change rate is increased to enhance the score result's responsiveness to sudden changes in the image frame structure. The output of the multidimensional weighted mapping function is used as input to generate a transmission priority parameter for each image frame. The transmission priority parameter is a continuously adjustable nonlinear score. A key frame protection factor is introduced during the generation process. The key frame protection factor is determined by the reconstruction importance score and the semantic area IoU change trend of the image frame within the time window. It is used to enhance the priority of the reconstructed key frame in the transmission sorting. The transmission priority parameter and the link stability prediction value are input into a soft maximum matching strategy module, and a channel allocation calculation is performed between the image frame and multiple wireless transmission channels, allowing image frames with similar transmission priority parameters to be simultaneously mapped to multiple wireless transmission channels with higher link stability prediction values, thereby achieving redundant fault-tolerant transmission across channels; The channel allocation instructions and corresponding transmission priority parameters of the output image frames are used as input parameters of the subsequent grouping module and the transmission module to achieve multi-channel image transmission optimization driven by the joint efforts of image frame content scoring and link status prediction.

[0075] In the link quality-aware multi-channel image transmission optimization system for drones described in this paper, the decision-making module plays a bridging role in the system's image scheduling chain. This module receives intermediate processing results from the modeling and prediction modules. Based on these results and the current mission scenario, it constructs a multidimensional weighted mapping function and a nonlinear priority generation mechanism to make quantitative decisions about the data scheduling priority and transmission path for each image frame. Ultimately, this decision is passed to the subsequent grouping and transmission modules, achieving end-to-end image transmission control.

[0076] First, the decision module receives two content scoring metrics from the modeling module: a time sensitivity score and a reconstruction importance score. The former measures the image frame's tolerance for transmission delays. A higher value indicates a more urgent transmission of the image frame, and delays may lead to semantic distortion or delayed task feedback. The latter measures the importance of the image frame to structural semantic restoration within the entire image sequence. A higher value indicates a frame's greater contribution to maintaining semantic continuity and scene coherence, such as frames containing structural changes such as the appearance, disappearance, splitting, and aggregation of objects. Simultaneously, the decision module receives a link stability prediction from the prediction module. This prediction is an estimate of the stability of each wireless transmission channel within the future target transmission period. It reflects the channel's comprehensive performance, such as packet loss rate, signal-to-noise ratio, and delay fluctuation, and is used to determine whether the channel is suitable for transmitting the current frame.

[0077] The core processing logic of the decision module lies in constructing a multi-dimensional weighted mapping function. This function takes as input the image frame's time sensitivity score, reconstruction importance score, and the predicted link stability of the wireless transmission channel. It outputs a continuous score, which serves as the basis for the comprehensive scheduling priority of the image frame in a specific task environment. Unlike traditional scoring methods, this system adopts a task-type adaptive dynamic weight adjustment mechanism. This mechanism first establishes a set of basic weight templates and then adjusts the importance of different scoring dimensions based on the specific task being performed by the system (such as real-time target tracking, map construction, and environmental monitoring). For example, in real-time target recognition tasks, the system will significantly increase the weight of the time sensitivity score, especially the portion directly related to the rate of change of the intersection over union (IoU) of semantic regions, thereby enhancing the system's scoring responsiveness to frames with sudden changes in semantics. If the task shifts to stable monitoring or non-real-time mapping, the weight of the reconstruction importance score can be appropriately increased, allowing the system to prioritize the preservation of the overall scene structure rather than transmission latency.

[0078] This multidimensional weighted mapping function is essentially a linear-nonlinear combination mapping of task-driven constraints and state feedback scores. In the initial calculation, each input score item undergoes normalization preprocessing to eliminate scale differences caused by different sources. Each score value is then weighted according to the dynamic weight configuration of the current task to form a combined score vector. In certain situations, such as when the link status is extremely unstable or the difference in image frame scores is not obvious, the system will automatically trigger an anomaly correction mechanism to perform weighted corrections on the score output to avoid abnormal resource scheduling deviations.

[0079] After the mapping function is calculated, the system uses its output as input to generate the image frame transmission priority parameter. This transmission priority parameter is not a simple hierarchical classification identifier, but rather a continuously adjustable nonlinear floating-point value. To enhance its ability to distinguish key image frames, the system introduces a dynamic adjustment factor, called the key frame protection factor, into its generation process. This factor is calculated by jointly analyzing the reconstruction importance score of an image frame and the trend of changes in the semantic area Intersection over Union (IoU) of that frame within a time window. Specifically, the system analyzes the IoU change rate of image frames in the time series—the difference between the IoU change rate of the current frame and the previous frame—to determine whether the current frame is at a semantic structure jump point. If this difference increases significantly, indicating a sudden change in the semantic area (such as rapid object movement, occlusion, or displacement), the system increases the key frame protection factor for that image frame accordingly, assigning it a higher weight in the generated nonlinear score, ensuring that it receives higher transmission priority in subsequent scheduling.

[0080] When the transmission priority parameter is generated, the system inputs it together with the aforementioned link stability prediction value into the soft maximum matching strategy module. The design of this module is based on the optimization of the instability of the traditional maximum score matching mechanism. In the traditional method, the system will assign each image frame only to the channel with the highest score. Although this method is simple, it is easy to cause score conflicts or load offsets when facing image frames with similar scores or wireless channels with unstable link status. The soft maximum matching strategy introduced in the present invention allows multiple image frames to share link resources and allows channel allocation to slide and adjust between multiple image frames with similar scores, thereby achieving higher scheduling flexibility and transmission fault tolerance. Specifically, during the channel allocation process, the system will perform similarity clustering on the transmission priority parameters of all image frames, classify frames with similar scores into the same group, and assign them to multiple channels with higher link stability prediction values, forming a many-to-many buffer mapping structure between channels and frames.

[0081] This mechanism not only enhances the redundancy of image transmission at the channel level, but also improves stability and responsiveness in highly dynamic scenarios. For example, within a transmission cycle, if the difference in transmission priority between two image frames is less than a preset threshold, the system allows them to be simultaneously assigned to two different channels and, if sufficient resources are available, they are sent repeatedly to ensure that at least one version successfully reaches the receiving end. This ensures that even momentary jitter or interruption on a channel will not affect the reception of key image frames, greatly improving the robustness of the system.

[0082] Ultimately, the decision module outputs the channel allocation instructions and transmission priority parameters for each image frame to the subsequent grouping and transmission modules. The channel allocation instructions clearly indicate which wireless transmission channels the image frame should be sent through, while the transmission priority parameters serve as a reference for sorting during compression grouping. This controls the granularity of the compression encoding, the inter-frame reference relationship, and the delay strategy. This ensures that when frame resources are limited, key frames are prioritized, while non-key frames are discarded or compressed, achieving end-to-end multi-channel image scheduling optimization.

[0083] In summary, the decision module described in the present invention constructs a task-adaptive multidimensional scoring model, introduces a continuous priority scoring and key frame protection mechanism, and combines the soft maximum matching strategy with the link prediction coupling mechanism. While ensuring the integrity of image content expression, it fully considers the differences in wireless channel states and the real-time requirements of the task, and realizes efficient joint scheduling control of image content and link quality.

[0084] The grouping module 104 is used to receive the image frame evolution vector and the transmission priority parameter of the image frame, perform content clustering on the image frames with similar transmission priority parameters and similar image frame evolution vectors, and construct a cross-frame compression grouping structure; based on the change rate information represented in the image frame evolution vector, set the compression algorithm type, quantization parameter and inter-frame prediction structure for each compression group, and obtain the compression profile and the corresponding compressed image data.

[0085] The grouping module 104 is the core task of optimizing image frame compression within the entire system. Its operation involves not only clustering image frames at the content level but also generating compression profiles for cross-frame structures, thereby providing efficient and targeted encoding output for final image data transmission. Its inputs come from the modeling module 102 and the decision module 103, respectively, including the image frame evolution vectors and transmission priority parameters. Its outputs include a customized compression profile and compressed image data for each group of image frames, which are used by the transmission module 105.

[0086] In actual operation, the grouping module 104 first receives the image frame evolution vector and the transmission priority parameters corresponding to each frame of the image. The system sets a clustering window, for example, containing N consecutive frames of images, usually 5 to 10 frames. The module scans the image frames in the window in turn and calculates the similarity between the image frame evolution vectors of each frame. The similarity calculation method can be based on indicators such as Euclidean distance, cosine similarity or Mahalanobis distance. If the evolution vectors of two image frames have similar numerical change trends, it means that they have similarities in texture, motion and semantic structure. In addition, the system will also refer to the transmission priority parameters of these image frames. If the priority difference is within the set threshold range, for example, no more than two priority levels, it is considered that these image frames meet the prerequisite for compression grouping of the same content.

[0087] Image frames that meet the above conditions are grouped together in a single compression group, thus constructing a cross-frame compression group structure. Each compression group typically contains several frames, which are relatively consistent in terms of content change trends and scheduling priorities, making them suitable for a unified compression strategy. The system analyzes each compression group as a whole, extracting the average and variation of the image frame evolution vectors within the group as a basis for determining the compression strategy.

[0088] Next, grouping module 104 determines the specific compression algorithm to use based on the overall rate of change of the image frames within the compression group. For example, when the image frame change rate is low, the system prioritizes encoding using an inter-frame prediction structure (such as a B-frame structure). When the change rate is high, the system switches to a compression structure primarily based on intra-frame coding (such as a hybrid structure of all I frames or P frames) to reduce prediction error accumulation. After determining the compression algorithm type, the system further sets the quantization parameter and inter-frame prediction depth. The quantization parameter is set based on the importance score of the image frame reconstruction. A higher score indicates a finer quantization level, a lower encoding compression ratio, and greater image quality preservation. The inter-frame prediction depth is determined by the maximum rate of image frame change. The lower the rate, the higher the prediction depth can be to improve the compression ratio. If the content changes significantly, the number of prediction levels is limited to prevent error diffusion.

[0089] Finally, the grouping module 104 generates a compression profile for each compressed group, detailing the compression algorithm type, quantization parameter settings, inter-frame prediction structure, coded frame order, and error-tolerance parameters used for the group. The compressed image data according to these profiles is then output as compressed image data. The compression profile and compressed image data are then sent to the transmission module 105 for subsequent bandwidth scheduling and transmission based on channel allocation instructions and priority parameters.

[0090] Through the above mechanism, the grouping module 104 can not only realize the effective clustering of image frames with similar content, but also flexibly configure compression strategies for different groups according to evolution trends and scheduling requirements, thereby retaining the key content of the image to the greatest extent while meeting real-time transmission requirements, thereby improving the system's transmission efficiency and visual reconstruction quality in complex dynamic scenes.

[0091] To further illustrate the working mechanism of the grouping module 104, a specific example is used below to fully demonstrate its operation process in the actual image transmission process.

[0092] Assume that a drone is performing a high-altitude patrol mission. The current system needs to process a sequence of six consecutive image frames, numbered Frames A through F. Modeling module 102 generates an image frame evolution vector for each frame, representing the changing trends in texture, edges, and semantics between frames. Simultaneously, decision module 103 assigns a transmission priority parameter to each frame: Frames A and B are assigned priority 1, Frames C and D are assigned priority 2, and Frames E and F are assigned priority 3.

[0093] Grouping module 104 first analyzes the evolution vectors of these image frames. Assume that the numerical differences between the evolution vectors of frames A and B in six dimensions are very small, with the Euclidean distance below the set threshold of 0.15. Furthermore, their transmission priority parameters are identical, both 1. Based on this, the module determines that frames A and B have similar image content evolution and the same transmission urgency, meeting the conditions for forming a compressed group. Therefore, they are classified as compressed group 1.

[0094] Similarly, the evolution vectors of frames C and D show similar semantic change rates and texture changes. Although the movement directions are slightly different, the overall change amplitude is consistent. The distance is 0.18, which is lower than the upper limit of 0.2 allowed for clustering. In addition, the transmission priority of both is 2, which also meets the clustering conditions, so they are classified into compression group 2.

[0095] The changes in content structure between frames E and F are more complex. Frame E contains distinct building edges and moving objects, while frame F captures a large, featureless area. The two exhibit a high difference in evolution vectors (a distance of 0.32), and frame E's reconstruction importance score is significantly higher than that of frame F. Although the two frames share the same transmission priority parameters, the module refuses to merge them based on their content dissimilarity and therefore treats them as separate groups: frame E forms compressed group 3, and frame F forms compressed group 4.

[0096] During the compression configuration phase, the module analyzes the average evolution vectors for each compression group. For compression group 1, because frames A and B change slowly and the target position is stable, the system selects a lightweight inter-frame prediction structure based on P frames and sets the quantization parameter (QP) to 28, maintaining a moderate compression ratio. For compression group 2, frames C and D exhibit slight target displacement, but the image structure is stable. Therefore, the P frame structure is still used, but the quantization strength is appropriately reduced, setting the QP to 26. For compression group 3, frame E has complex content and significant target variation, so the system uses I-frame encoding and sets a lower quantization parameter (QP) of 22 to preserve image detail. For compression group 4, frame F has stable content, so the module selects a strong compression scheme, using a P frame structure and setting the QP to 32.

[0097] For each of the four compression groups, the module generates a corresponding compression profile, which explicitly records information such as the selected coding structure (I-frame or P-frame), quantization parameter (QP), GOP length (if any), prediction depth, and reference frame index relationship. The module then performs image compression operations according to the profile and outputs compressed image data. The compressed image data for each group is mapped to its compression profile and sent to the transmission module 105 for subsequent scheduling according to channel allocation instructions and priority parameters.

[0098] Through this specific process, the grouping module 104 makes full use of the image frame evolution vector and transmission scheduling information to complete the complete chain from content clustering to parameter setting to data generation. It not only improves the targetedness of image compression, but also ensures that in scenarios with limited link resources, key images can be prioritized with higher quality, thereby significantly enhancing the stability and task adaptability of the system's overall image transmission.

[0099] Furthermore, the grouping module specifically includes the following processing steps: receiving the image frame evolution vectors and the transmission priority parameters of the image frames, constructing a combined similarity measurement function based on the difference between the cosine similarity between the image frame evolution vectors and the transmission priority parameters, performing cross-frame content clustering between the image frames under a set clustering similarity combination threshold constraint, and generating a plurality of cross-frame compression grouping structures, wherein each cross-frame compression grouping structure includes image frames having similar image frame evolution vectors and similar transmission priority parameters; Performing intra-group statistical analysis on the image frame evolution vectors of the image frames included in each of the cross-frame compression grouping structures, calculating the intra-group variance of the image frame evolution vectors of the compression grouping structure, and using the intra-group variance as an image content change intensity index. The compression algorithm type of the compression grouping structure is set according to the image content change intensity index. When the intra-group variance is higher than a preset change intensity threshold, the compression grouping structure is set to adopt a full intra-frame coding structure with an inter-frame reference depth of zero; when the intra-group variance is lower than a preset change smoothness threshold, the compression grouping structure is set to adopt a prediction chain compression structure, and the inter-frame prediction depth of the prediction chain is further set according to the content complexity reflected in the image frame evolution vector. Counting the average reconstruction importance scores of all image frames in each of the compressed group structures, and determining whether the average is higher than a preset reconstruction key frame threshold; when the average is higher than the threshold, identifying the compressed group structure as a high reconstruction importance frame group, and configuring a forward error correction redundant segment structure for it, so as to introduce redundant protection information into the compressed image data to improve image reconstruction capability in channel instability scenarios; Writing the compression algorithm type, inter-frame prediction depth, and error correction redundancy segment structure configuration results of each of the compression group structures into a compression profile, and generating compressed image data corresponding to the compression profile, wherein the compression profile is used to indicate an image frame transmission strategy in a subsequent channel scheduling phase; The compression profile and the compressed image data are output, and the compression profile is used as a control basis for the transmission module to distribute image frames by channel, thereby realizing the joint optimization of the image frame evolution vector, transmission priority parameter and image content change trend in compression strategy and channel scheduling.

[0100] In the present invention's link-quality-aware multi-channel image transmission optimization system for drones, the grouping module assumes intermediate decision-making responsibility for image frame compression processing. Its design directly impacts the transmission efficiency and image restoration quality of subsequent channel scheduling. This module not only rationally groups image frames with strong semantic continuity and similar scheduling priorities, but also selects an adaptive compression structure and coding depth based on the changing trends of image content. It also configures a fault-tolerant forward error correction mechanism on key image frame sets, thereby forming a content-driven image frame grouping strategy that balances image compression efficiency with link transmission robustness.

[0101] First, the module receives the image frames to be processed before compression and obtains two types of input data: the first is the image frame evolution vector, which is generated by the modeling module and contains the content evolution characteristics between the image frame and adjacent frames in terms of texture gradient variability, edge displacement direction variance, target number change rate, and semantic area IoU change rate; the second is the transmission priority parameter, which is generated by the decision module and is used to quantify the importance of each frame in the current network scheduling scenario. The higher the value, the more priority the frame should be scheduled and transmitted as soon as possible. These two types of data reflect the dynamic state of the image frame at the image content level and its ranking position in the transmission strategy. They are the core basis for constructing image frame clustering groups.

[0102] After receiving the above data, the system first constructs a combined similarity measurement function in the grouping module. This function is designed to comprehensively consider the content similarity and scheduling requirement consistency between image frames. Content similarity is measured by the cosine similarity between the image frame evolution vectors, that is, the directional angle between the two image frames in the vector space reflects their similarity in texture, motion, semantic changes, etc.; scheduling consistency is characterized by the absolute difference between the transmission priority parameters corresponding to the image frames. The closer the priorities of the two frames, the higher the value of their joint processing. By setting a joint threshold for combined similarity judgment, the system screens out image frame pairs that have both content similarity and scheduling consistency and classifies them into the same compression group. In this way, through joint analysis based on image frame evolution vectors and transmission priority parameters, the system can find the optimal balance between spatial semantic consistency and transmission scheduling tightness.

[0103] Each generated compression group structure (also known as a cross-frame compression group structure) will contain several image frames with similar content evolution trends and relatively consistent scheduling priorities. To further optimize subsequent compression structure selection, the system performs intra-group statistical analysis on the image frame evolution vectors in each compression group, calculating its intra-group variance across all dimensions. This intra-group variance is used to characterize the intensity of content change within the group of image frames. A larger variance indicates more significant differences in texture, motion direction, semantic structure, and other aspects within the group, and a more dramatic degree of change. A smaller variance indicates a stable inter-frame structure, good image sequence continuity, and greater inter-frame compression potential. The system uses this intra-group variance as an indicator of the intensity of image content change and formulates corresponding compression strategies accordingly.

[0104] Specifically, when the intra-group variance exceeds the system-set threshold for change intensity, indicating that the content of the group of image frames fluctuates dramatically and that inter-frame redundancy is low, the system sets the compression group structure to adopt a full intra-frame coding structure with an inter-frame reference depth of zero. That is, all image frames in the group are encoded as I frames to avoid image distortion caused by the diffusion of prediction errors. When the intra-group variance falls below the system-set threshold for stable content change, indicating that the group of image frames has a high degree of content continuity and meets the conditions for establishing a prediction chain, the system sets the compression group to adopt an inter-frame prediction structure and automatically determines the appropriate inter-frame prediction depth based on the dimensional information representing the content complexity in the evolution vector of the group of image frames. The lower the content complexity, the larger the prediction depth can be set to achieve a higher compression ratio; if the content complexity is relatively high, the system will limit the prediction depth to avoid the negative impact of error accumulation in long-chain predictions on image quality.

[0105] Next, the system further evaluates whether the compressed group structure belongs to a high-reconstruction-importance frame group. To this end, the system calculates the average reconstruction importance score of the image frames contained in each compressed group structure. This score is derived from the modeling module and reflects the importance of the image frame to the restoration of the scene structure based on the trend of changes in the number of objects and the fluctuation trend of the semantic area IoU. If the average reconstruction importance score of a compressed group exceeds the set critical value, it means that the group of image frames belongs to a key semantic component node in the image sequence, such as a time period with frequent object convergence, movement, or occlusion events. At this time, the system identifies the compressed group as a high-reconstruction-importance frame group and configures it with a forward error correction redundant segment structure, that is, introducing additional data redundancy in the compressed image data to improve the image frame recovery capability after data packet loss during subsequent transmission. This redundant structure can be implemented using simple parity bits, low-density error correction codes, or prediction residual backup, and is ultimately written into the compression profile as a structured tag.

[0106] After the compression strategy is fully configured, the system writes the compression algorithm type (i.e., inter-frame prediction or intra-frame compression), inter-frame prediction depth (e.g., set to 2 or 4 frames per group), and error correction redundancy segment structure configuration results into the compression profile. This profile not only guides the compression module's specific encoding operations on image frame data but also serves as important auxiliary information for the subsequent channel scheduling module when formulating image frame transmission plans. Different compression structures correspond to different data volumes, delay tolerance, and packet loss recovery capabilities. Therefore, the scheduling module needs to refer to the structural parameters in the compression profile to determine how to properly arrange the transmission order, fragmentation strategy, and retransmission logic in a multi-channel environment.

[0107] After completing the above configuration and compression operations, the system outputs the compression profile and compressed image data corresponding to each compression group structure. Among them, the compression profile, as a set of structured metadata, describes in detail the compression logic and fault tolerance mechanism adopted by each group of image frames, and also marks the priority, redundancy and prediction logic of the group of image frames. This information will provide constraints and guidance at the transmission strategy level when the image data enters the transmission module. For example, for frame groups using inter-frame prediction structures, the scheduling module must ensure that their reference frames are transmitted first; for key frame groups configured with redundant segments, the scheduling module can improve the transmission success rate through priority multi-channel forwarding, redundant distribution, etc., thereby realizing the coordinated integration of image compression optimization and wireless link scheduling.

[0108] In summary, the grouping module described in this invention forms a cross-frame content clustering mechanism by introducing a joint similarity metric between image frame evolution vectors and transmission priority parameters. This mechanism, combined with the intensity of content change within a group and the importance of semantic reconstruction, enables full-process optimization control of compression structure selection, prediction depth control, keyframe redundancy configuration, and compression configuration output. This module not only enhances compression efficiency at the content clustering level and ensures transmission security at the structural configuration level, but also provides a quantifiable and callable control basis for the subsequent channel scheduling module.

[0109] The transmission module 105 is configured to receive the compression configuration file and the compressed image data, and distribute each compressed image data to a corresponding wireless transmission channel according to a transmission priority parameter based on the channel allocation instruction.

[0110] Transmission module 105 is responsible for dispatching and allocating the compressed profile and compressed image data output from grouping module 104 to the corresponding wireless transmission channel based on the channel allocation instructions and transmission priority parameters provided by decision module 103, completing the final data downlink transmission. This module is at the very end of the data flow in the entire system, but its functions directly affect the reliable delivery of image data, transmission efficiency, and the implementation of the image frame and channel matching strategy.

[0111] During system operation, the transmission module 105 first receives the output data from the grouping module 104, including the compression profile and compressed image data corresponding to each compressed group. The compression profile clearly records the encoding structure, quantization parameters, and inter-frame reference relationships used for the image group, while the compressed image data is the data stream actually compressed based on this configuration. Simultaneously, the transmission module also receives the channel allocation instruction generated by the decision module 103. This instruction specifies a target wireless transmission channel for each image frame or compressed group, and provides corresponding transmission priority parameters.

[0112] The transmission module matches image groups with channel instructions and places the compressed image data into the corresponding channel buffer queue according to the channel number identified in the instruction (for example, channel 1, channel 2, or channel 3). Each channel independently maintains a queue of data to be sent. The order of image data in the queue strictly follows the transmission priority parameter. The lower the priority value, the higher the transmission order. The system supports dynamic reordering. When multiple compressed image data are assigned to the same channel, the transmission module automatically adjusts their queue position based on their priority parameters to ensure that high-priority image frames are transmitted first, thereby achieving low-latency transmission of delay-sensitive content.

[0113] During the data scheduling phase, the transmission module dynamically allocates bandwidth resources based on air interface status feedback from the wireless modulation subsystem, combined with each channel's actual bandwidth utilization and real-time status (e.g., whether it's experiencing interference, retransmission, or availability). For example, if the instantaneous available rate on a channel drops, the transmission module will temporarily suspend the transmission of low-priority frames on that channel and prioritize high-priority data on another channel to ensure timely delivery of critical image frames within limited bandwidth.

[0114] In addition, to improve data transmission reliability, the transmission module can also use redundant coding or forward error correction (FEC) for some key frame data based on the redundancy strategy in the compression profile, and repeatedly send it through a secondary or backup channel when the channel is congested, ensuring recoverability during the image reconstruction phase. For compressed image data that supports fragmented transmission, the transmission module can also decompose the image frame into multiple transmission units (such as NALU units or fixed-length data blocks) and forward them in fragments between multiple channels according to the scheduling strategy, thereby improving the utilization of channel resources.

[0115] When the image frame is sent, the transmission module records the sending status, actual time, number of retransmissions and other data, and sends a completion notification to the system to provide feedback data support for subsequent scheduling modules and compression modules, realizing adaptive updates of the system scheduling strategy.

[0116] During execution, the entire transmission module 105 maintains a high degree of data consistency and strategic coordination with other modules in the system, ensuring a close integration of priority control of image compression content with dynamic allocation of channel resources. This not only ensures the timely transmission of critical image data in environments with limited bandwidth resources and fluctuating channel quality, but also improves the transmission stability, robustness, and scalability of the entire drone image transmission system. Through task queues, channel mapping, scheduling control, and feedback collection, this module constructs a complete and implementable wireless image data scheduling system. This system possesses excellent engineering practicality and real-time operation capabilities, making it suitable for drone platforms with high requirements for precise bandwidth management and latency control in multi-channel wireless image transmission scenarios.

[0117] To illustrate the working process of the transmission module 105 in more detail, the following fully describes the entire data receiving, scheduling and sending process with a practical example.

[0118] Consider a drone performing a high-altitude traffic monitoring mission. The system is currently processing four compressed image packets, output by packetization module 104. Each packet contains a compression profile and compressed image data. Based on the scheduling results generated by decision module 103, the system also obtains the channel assignment instructions and transmission priority parameters for each image packet. Specifically, image packets A and B are assigned to wireless channel 1, with priorities of 1 and 3, respectively; image packet C is assigned to wireless channel 2, with priority 2; and image packet D is assigned to wireless channel 3, with priority 1.

[0119] Transmission module 105 first places the four image packets into the buffer queues corresponding to their respective channels. For channel 1, the module compares the priority parameters of image packets A and B before writing them into the queue. It determines that image packet A has a higher priority and therefore places it before image packet B. Since channels 2 and 3 only receive one image packet, no reordering is required in the queues.

[0120] The compression profile for image group A specifies the use of a P-frame structure, a quantization parameter of 26, and requires transmission in complete frames. The compression profile for image group C enables fragmentation, splitting the compressed image data into several equal-length segments, each no longer than 1024 bytes, and requiring them to be sent sequentially. Based on these configurations, the transport module numbers the data slices for image group C into a queue and includes reassembly information to help the receiving end recover the frame structure.

[0121] During the actual transmission process, the transmission module dynamically allocates bandwidth resources based on the channel availability status reported by the wireless module. At this point, channel 1 is stable and available, channel 2 is experiencing mild interference, and channel 3 is briefly congested. The system prioritizes dispatching image packet A from channel 1's queue and transmits it in full, frame-by-frame order. Due to significant signal fluctuations on channel 2, the module transmits the data slices of image packet C sequentially, numbered, and checks for acknowledgment after each slice. Upon detecting a timeout for the acknowledgment of the fourth slice, the system immediately triggers the automatic retransmission mechanism, retransmitting only the fourth slice. Due to insufficient instantaneous bandwidth on channel 3, the module temporarily suspends transmission of image packet D and attempts to use the secondary channel to transmit a portion of the key data for image packet D during an idle period on channel 1, while simultaneously recording the asynchronous transmission status.

[0122] Throughout the entire process, the module continuously tracks parameters such as average throughput, number of retransmissions, and transmission latency for each channel. It also generates a transmission report after each image packet is sent, marking the status as "completed," "partially completed," or "retransmitting." This transmission report, along with channel utilization and frame loss information, is fed back to the upper-level management module for optimization of subsequent scheduling strategies.

[0123] Through this specific workflow, the transmission module 105 completely realizes the entire process from receiving compressed image data, to queue scheduling according to channel allocation instructions, to sending according to transmission priority, and executing reorganization and retransmission control based on configuration, ensuring that the system can achieve high-efficiency, low-latency, and high-reliability image transmission in a multi-channel environment, and effectively guaranteeing the transmission quality and task real-time performance of key image content under complex communication conditions.

[0124] Furthermore, the transmission module is specifically configured to: receiving the compression profile and the compressed image data, and performing initial scheduling of the compressed image data of each image frame according to the channel allocation instruction; during the scheduling process, if the current status of all wireless transmission channels does not meet the transmission reliability requirement corresponding to the link stability prediction value output by the prediction module, constructing an image frame importance queue, and performing weighted sorting on the image frame importance queue according to the transmission priority parameter and the reconstruction importance score to form a scheduling priority list for the image frames; Inputting high-priority image frames in the image frame importance queue into an adaptive reordering strategy module, dynamically reconstructing channel allocation instructions for corresponding image frames based on currently available backup wireless transmission channel resources, comparing the reconstructed channel allocation instructions with the original channel allocation instructions, and writing a flag field in the compression profile to indicate that the channel allocation of the image frame has been updated; Reading the compression group structure type and the average reconstruction importance score recorded in the compression profile, executing a multi-channel redundant copy distribution strategy for image frames marked as high reconstruction importance frame groups, simultaneously transmitting compressed image data of the image frames on two or more wireless transmission channels with higher link stability prediction values, and synchronously writing the compression profile into a distribution mapping table of the image frames on different channels for receiving end recovery and subsequent retransmission control; During the transmission of image frames through each wireless transmission channel, the actual packet loss rate of each channel is recorded in real time, and the deviation between the packet loss rate and the corresponding link stability prediction value is calculated. The deviation value is used as a dynamic adjustment factor for channel credibility and fed back to the decision module for correcting the link reliability weight during the generation of channel allocation instructions for subsequent image frames; The actual transmission records of the channels used by the image frames in the final scheduling process, the adaptive reordering history, the redundant copy channel mapping, and the link deviation statistics are written into the scheduling log as a historical basis for dynamically updating the compression configuration template and the image frame scheduling strategy during system operation, thereby realizing a robust scheduling process for the image frames driven jointly by the compression configuration file and link status feedback.

[0125] In the link quality-aware multi-channel image transmission optimization system for unmanned aerial vehicles (UAVs) described in this paper, the transmission module, as the terminal execution link of the image scheduling process, bears the key responsibility of distributing compressed image data to specific wireless transmission channels based on scheduling decision results. This module not only needs to strictly execute the channel allocation instructions output by the decision module but also possesses robustness capabilities under complex dynamic link conditions. In particular, when wireless channels are simultaneously unavailable or the link status is severely degraded, it can implement image frame transmission priority reordering, adaptive channel replacement, multi-channel redundant transmission, and feedback recording of scheduling results, providing adaptive reconstruction and optimization closed-loop operation support for the entire system.

[0126] First, during the initial operation phase, the transmission module receives the compressed image data output by the compression module and its corresponding compression profile. This profile contains key scheduling information, including channel allocation instructions, compression structure parameters, prediction chain structure, and redundancy strategy settings for each image frame. Under normal circumstances, the transmission module directly distributes the compressed image data to the corresponding wireless transmission channel according to these channel allocation instructions. However, in actual operation, due to the dynamic fluctuations of the wireless channel, it is possible that all currently allocated transmission channels will be unavailable within a specific scheduling period. For example, if a drone enters a communication blind spot during flight, or if multiple channels experience severe interference, congestion, or attenuation simultaneously, the system will detect that the actual channel transmission performance fails to meet the minimum reliability threshold corresponding to the original link stability prediction value. In this case, continuing with the original scheduling strategy will result in image frame transmission failure. Therefore, the transmission module immediately initiates an image frame reordering mechanism to ensure that limited transmission resources prioritize the most critical data content.

[0127] The implementation of this mechanism first relies on the construction of an image frame importance queue. Specifically, the system extracts the transmission priority parameters and reconstruction importance scores corresponding to each frame from the set of image frames currently waiting to be transmitted. The transmission priority parameters are usually calculated by the system during the decision-making phase based on time sensitivity and link conditions, representing the urgency of the frame in scheduling; while the reconstruction importance score reflects the importance of the image frame to the restoration of the semantic structure, especially in intervals with drastic scene changes or frequent target activities. This value will increase significantly. The transmission module weightedly fuses the above two indicators to form an image frame importance metric, and sorts all image frames to be transmitted accordingly to form a scheduling priority list.

[0128] The image frames at the top of the priority list will be selected into the adaptive rescheduling strategy process. This strategy aims to reassign high-priority image frames to the backup wireless transmission channels that are still available in the current system. To this end, the system will scan all connected wireless channels in the current communication module in real time, eliminate the set of channels that are unavailable or below the reliability threshold, and select the channel resources that best match the image frame link requirements from the remaining candidate channels for scheduling replacement. Once the new channel allocation is completed, the system will compare the updated result with the original channel allocation instruction of the image frame and set up a reconfiguration mark field in the compression configuration file to indicate that the channel scheduling logic of the frame has been dynamically adjusted. This mark field is not only used to update the status of the local control logic, but also provides a basis for the receiving end to identify redundant links and judge the channel switching history during the image reconstruction process.

[0129] After the channel scheduling reconstruction is completed, the system will start a multi-channel redundant copy distribution strategy for image frames marked as high reconstruction importance frame groups. This process depends on the compression group structure type of the image frame recorded in the compression profile and the average reconstruction importance score of the group. If the system determines that the value is higher than the set reconstruction key frame threshold, it means that the image frame plays a critical role in semantic continuity or visual stability. Therefore, the system will simultaneously send compressed image data copies of the frame on two or more wireless channels with higher link stability prediction values. The content of the replica data is completely consistent, and its purpose is to improve the anti-packet loss performance. Even if a transmission channel has problems such as packet loss or delay jitter during the transmission process, it can be successfully transmitted on another channel, ultimately improving the end-to-end image frame arrival rate.

[0130] To ensure the receiving end can correctly identify and interpret the source channel and replica status corresponding to the received image data, the system synchronously writes an image frame redundant replica channel distribution mapping table into the compression profile. This mapping table clearly identifies the channel to which each image frame is sent, whether it belongs to the primary channel or a replica channel, and whether there is an alternative scheduling. This information is used by the receiving end to guide the image frame processing process, including priority arrival, de-duplication sorting, and error recovery, ensuring the accuracy and sequence integrity of the final image frame reconstruction.

[0131] Throughout the actual image data transmission process, the system maintains real-time monitoring of the status of each wireless transmission channel, paying particular attention to the packet loss rate of each channel. The transmission module records the success or failure of each image frame transmission and counts the total number of packet losses and valid transmissions per unit time for each channel, dynamically calculating the actual packet loss rate for each channel. Simultaneously, the system compares this real-time packet loss rate with the link stability prediction output by the prediction module and calculates the deviation between the two. This deviation reflects the accuracy of the prediction model's assessment of the current channel status. If the deviation continues to increase, it indicates a sudden change in the link status or a delayed response from the prediction model. To enable system self-adjustment, the transmission module feeds each channel's predicted deviation value back to the decision module as a dynamic channel credibility adjustment factor, allowing the channel's weight to be lowered, replaced, or corrected through reinforcement learning during the generation of channel allocation instructions for subsequent image frames.

[0132] After completing the aforementioned scheduling and transmission operations, the system also logs the entire process of the image frame during the final scheduling execution phase, forming a set of structured scheduling log data. This log includes, but is not limited to: the final valid channel allocation result for each image frame, whether adaptive reordering occurred, a list of channels participating in replica redundancy, channel packet loss rate statistics, and link prediction deviation curves. This log data is not only used to monitor the system's operating status and trace issues, but also serves as an important basis for subsequent adaptive updates to the system's compression configuration template and optimization of the channel scheduling model structure. Based on the scheduling log, the system can implement operations such as online learning, dynamic parameter adjustment, and channel credibility database reconstruction, forming a closed-loop control architecture of "perception-scheduling-feedback-reoptimization," significantly improving the system's real-time responsiveness and environmental adaptability.

[0133] In summary, the transmission module not only distributes compressed image frame data to various wireless channels but also provides a multi-layered response mechanism for uncertain scenarios such as link failure, scheduling anomalies, and data congestion. This includes establishing image frame importance queues, adaptive scheduling reordering, sending redundant copies across multiple channels, feedback on link prediction deviations, and scheduling log archiving. This series of functional designs and technical processes significantly differs from the static channel binding and fixed scheduling structures of existing image transmission systems, and offers significant technical benefits in improving image communication quality, enhancing key frame protection, and optimizing system robustness.

[0134] Although the present application is disclosed as above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.

Claims

1. A multi-channel image transmission optimization system for UAVs based on link quality perception, characterized in that: include: The prediction module is used to receive the link quality indicators of multiple wireless transmission channels that the drone is currently connected to in the historical period and build a time series prediction model; Obtaining a link stability prediction value of each wireless transmission channel within a future target transmission period according to the timing prediction model; A modeling module is configured to receive a sequence of image frames captured by a drone, analyze the texture change characteristics, edge motion trajectories, and semantic target consistency between adjacent image frames, and construct an image frame evolution vector that reflects the correlation between the rate of change of image frame content and its temporal sequence; based on the image frame evolution vector, calculate the time sensitivity score and reconstruction importance score of each image frame within the prediction time window; a decision module, configured to receive the link stability prediction value, the time sensitivity score, and the reconstruction importance score, determine the optimal matching relationship between the image frame and each wireless transmission channel by constructing a multi-dimensional weighted mapping function, and obtain a channel allocation instruction and corresponding transmission priority parameters for the image frame based on the optimal matching relationship; a grouping module, configured to receive the image frame evolution vector and the transmission priority parameter of the image frame, perform content clustering on image frames with similar transmission priority parameters and similar image frame evolution vectors, and construct a cross-frame compression grouping structure; Based on the change rate information represented in the image frame evolution vector, the compression algorithm type, quantization parameter and inter-frame prediction structure are set for each compression group, and a compression profile and corresponding compressed image data are obtained; The transmission module is used to receive the compression configuration file and the compressed image data, and distribute each compressed image data to the corresponding wireless transmission channel according to the transmission priority parameter according to the channel allocation instruction.

2. The UAV multi-channel image transmission optimization system based on link quality perception according to claim 1 is characterized in that: The modeling module is specifically used for: After receiving a sequence of image frames captured by a drone, for each image frame, the texture gradient variability, edge displacement direction variance, target number change rate, and semantic area IoU change rate between the image frame and its preceding and following adjacent frames are extracted. The extracted values ​​are combined to form an image frame evolution vector, which is used to characterize the rate of change of image frame content and the temporal correlation. Based on the image frame evolution vector, constructing a time-domain weighted vector including a time-attenuation weight, wherein the time-attenuation weight is set according to the frame spacing of the image frame relative to the current frame, and is used to perform temporal weight modulation on each dimensional feature in the image frame evolution vector to obtain a time-weighted image frame evolution vector; Inputting the time-weighted image frame evolution vector into a temporal sensitivity scoring function, the scoring function calculates the temporal sensitivity score of the image frame within the prediction time window based on the edge displacement direction variance and the semantic region IoU change rate, which is used to quantify the image frame's tolerance to transmission delay; The time-weighted image frame evolution vector is input into a reconstruction importance scoring function. The reconstruction importance scoring function uses a nonlinear function to calculate the reconstruction importance score of the image frame based on the trend changes of the target quantity change rate and the semantic area IoU change rate in multiple consecutive image frames. The score is used to evaluate the contribution of the image frame to the restoration of the image sequence structure. The time sensitivity score and the reconstruction importance score are provided to the decision module as output results of the modeling module.

3. The UAV multi-channel image transmission optimization system based on link quality perception according to claim 1 is characterized in that: The decision module is specifically used for: receiving the time sensitivity score and the reconstruction importance score output by the modeling module, and the link stability prediction value of each wireless transmission channel in the future target transmission period output by the prediction module; Based on the timeliness sensitivity score, reconstruction importance score, and link stability prediction value, a multidimensional weighted mapping function is constructed. The multidimensional weighted mapping function adopts a dynamic weight adjustment mechanism that is adaptive to the task type. The weight of the timeliness sensitivity score is adjusted according to the current task type. When the task type is an object recognition task, the weight of the timeliness sensitivity score related to the semantic region IoU change rate is increased to enhance the score result's responsiveness to sudden changes in the image frame structure. The output of the multidimensional weighted mapping function is used as input to generate a transmission priority parameter for each image frame. The transmission priority parameter is a continuously adjustable nonlinear score. A key frame protection factor is introduced during the generation process. The key frame protection factor is determined by the reconstruction importance score and the semantic area IoU change trend of the image frame within the time window. It is used to enhance the priority of the reconstructed key frame in the transmission sorting. The transmission priority parameter and the link stability prediction value are input into a soft maximum matching strategy module, and a channel allocation calculation is performed between the image frame and multiple wireless transmission channels, allowing image frames with similar transmission priority parameters to be simultaneously mapped to multiple wireless transmission channels with higher link stability prediction values, thereby achieving redundant fault-tolerant transmission across channels; The channel allocation instructions and corresponding transmission priority parameters of the output image frames are used as input parameters of the subsequent grouping module and the transmission module to achieve multi-channel image transmission optimization driven by the joint efforts of image frame content scoring and link status prediction.

4. The UAV multi-channel image transmission optimization system based on link quality perception according to claim 1 is characterized in that: The grouping module is specifically used for: receiving the image frame evolution vectors and the transmission priority parameters of the image frames, constructing a combined similarity measurement function based on the difference between the cosine similarity between the image frame evolution vectors and the transmission priority parameters, performing cross-frame content clustering between the image frames under a set clustering similarity combination threshold constraint, and generating a plurality of cross-frame compression grouping structures, wherein each cross-frame compression grouping structure includes image frames having similar image frame evolution vectors and similar transmission priority parameters; performing intra-group statistical analysis on image frame evolution vectors of image frames included in each of the cross-frame compression grouping structures, calculating an intra-group variance of the image frame evolution vectors of the compression grouping structure, using the intra-group variance as an image content change intensity indicator, setting a compression algorithm type for the compression grouping structure based on the image content change intensity indicator, and setting the compression grouping structure to adopt a full intra-frame coding structure with an inter-frame reference depth of zero when the intra-group variance is higher than a preset change intensity threshold; When the intra-group variance is lower than a preset stable change threshold, setting the compression grouping structure to adopt a prediction chain compression structure, and further setting the inter-frame prediction depth of the prediction chain according to the content complexity reflected in the image frame evolution vector; Counting the average reconstruction importance scores of all image frames in each of the compressed group structures, and determining whether the average is higher than a preset reconstruction key frame threshold; when the average is higher than the threshold, identifying the compressed group structure as a high reconstruction importance frame group, and configuring a forward error correction redundant segment structure for it, so as to introduce redundant protection information into the compressed image data to improve image reconstruction capability in channel instability scenarios; Writing the compression algorithm type, inter-frame prediction depth, and error correction redundancy segment structure configuration results of each of the compression group structures into a compression profile, and generating compressed image data corresponding to the compression profile, wherein the compression profile is used to indicate an image frame transmission strategy in a subsequent channel scheduling phase; The compression profile and the compressed image data are output, and the compression profile is used as a control basis for the transmission module to distribute image frames by channel, thereby realizing the joint optimization of the image frame evolution vector, transmission priority parameter and image content change trend in compression strategy and channel scheduling.

5. The UAV multi-channel image transmission optimization system based on link quality perception according to claim 1 is characterized in that: The transmission module is specifically used for: receiving the compression profile and the compressed image data, and performing initial scheduling of the compressed image data of each image frame according to the channel allocation instruction; during the scheduling process, if the current status of all wireless transmission channels does not meet the transmission reliability requirement corresponding to the link stability prediction value output by the prediction module, constructing an image frame importance queue, and performing weighted sorting on the image frame importance queue according to the transmission priority parameter and the reconstruction importance score to form a scheduling priority list for the image frames; Inputting high-priority image frames in the image frame importance queue into an adaptive reordering strategy module, dynamically reconstructing channel allocation instructions for corresponding image frames based on currently available backup wireless transmission channel resources, comparing the reconstructed channel allocation instructions with the original channel allocation instructions, and writing a flag field in the compression profile to indicate that the channel allocation of the image frame has been updated; Reading the compression group structure type and the average reconstruction importance score recorded in the compression profile, executing a multi-channel redundant copy distribution strategy for image frames marked as high reconstruction importance frame groups, simultaneously transmitting compressed image data of the image frames on two or more wireless transmission channels with higher link stability prediction values, and synchronously writing the compression profile into a distribution mapping table of the image frames on different channels for receiving end recovery and subsequent retransmission control; During the transmission of image frames through each wireless transmission channel, the actual packet loss rate of each channel is recorded in real time, and the deviation between the packet loss rate and the corresponding link stability prediction value is calculated. The deviation value is used as a dynamic adjustment factor for channel credibility and fed back to the decision module for correcting the link reliability weight during the generation of channel allocation instructions for subsequent image frames; The actual transmission records of the channels used by the image frames in the final scheduling process, the adaptive reordering history, the redundant copy channel mapping, and the link deviation statistics are written into the scheduling log as a historical basis for dynamically updating the compression configuration template and the image frame scheduling strategy during system operation, thereby realizing a robust scheduling process for the image frames driven jointly by the compression configuration file and link status feedback.

Citation Information

Cited By

  • Camera calibration method

    CN121639819A

  • A camera calibration method

    CN121639819B

  • Multi-modal data transmission method and system for unmanned area operation

    CN121792629A

  • A multi-modal data transmission method and system for operating in an unmanned area

    CN121792629B

  • Network and task dual-driven unmanned aerial vehicle image instant return method in complex environment

    CN121967808A