A method for real-time transmission of river and lake inspection images based on UAV swarms

By using neural networks to predict channel gain and a distributed multi-agent system to optimize the transmission path, combined with adaptive modulation and coding and cooperative retransmission, the problem of unstable image transmission in UAV swarm river and lake inspection was solved, achieving efficient and reliable real-time image transmission.

CN121864953BActive Publication Date: 2026-05-26山东黄河顺成水利水电工程有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
山东黄河顺成水利水电工程有限公司
Filing Date
2026-03-18
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

The stability of image transmission links in river and lake inspections by drone swarms faces challenges. Due to complex terrain, electromagnetic environment and external interference, transmission is unstable and cannot provide continuous and reliable visual information support.

Method used

By constructing a neural network-based predictive channel gain matrix, combining a distributed multi-agent system and an adaptive modulation and coding mechanism, a multi-hop optimal transmission path is generated, and collaborative retransmission is implemented when data transmission fails. The transmission strategy is optimized using linear network coding and online learning.

Benefits of technology

It significantly improves the reliability of image transmission by UAV swarms in complex river and lake environments, avoids transmission interruptions, ensures the continuity and clarity of video footage, and achieves efficient and robust real-time transmission with low latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121864953B_ABST
    Figure CN121864953B_ABST
Patent Text Reader

Abstract

This invention relates to a real-time transmission method for river and lake inspection images based on UAV swarms, specifically in the field of river and lake inspection image transmission. By combining forward-looking channel prediction and intelligent path planning, the reliability of image transmission by UAV swarms in complex river and lake environments is significantly improved. Its adaptive modulation and coding mechanism, which integrates real-time measurement and prediction information, effectively avoids transmission interruptions caused by rapid channel fading, ensuring the continuity and clarity of video images. At the same time, the collaborative network coding retransmission and online learning closed loop based on actual transmission feedback enable the entire system to continuously self-optimize and enhance its adaptability to dynamic environments, thereby achieving efficient and robust real-time transmission of inspection image data while ensuring low latency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of river and lake inspection image transmission, and more specifically, to a method for real-time transmission of river and lake inspection images based on a swarm of unmanned aerial vehicles (UAVs). Background Technology

[0002] In recent years, the use of drone swarms for routine inspections of rivers, lakes, reservoirs, and other water bodies has become an important means of ecological environment monitoring, hydrological data collection, water safety supervision, and emergency response. Compared with fixed camera monitoring, drone swarms have outstanding advantages such as mobility, wide field of view, and large coverage. They can efficiently acquire high-definition real-time image data of the water surface, shoreline, and surrounding environment, thereby enabling dynamic monitoring of targets such as pollution discharge, illegal fishing, eutrophication of water bodies, and the status of flood control facilities. This application essentially constitutes a large-scale, mobile aerial closed-circuit television system, which places extremely high demands on the continuity, real-time performance, and reliability of video image data.

[0003] However, when applying drone swarms to actual river and lake inspection tasks, the stability of their image transmission links faces severe challenges. The terrain and electromagnetic environment of river and lake areas are often very complex. Inspection areas frequently span vast water surfaces and are surrounded by obstacles such as mountains, high dams, bridges, wind turbines, and dense vegetation. The water surface causes strong specular reflection of the electromagnetic wave signals transmitted by the drones downlink, while the aforementioned obstacles result in severe shielding and diffuse reflection. This environment causes significant multipath effects in the wireless channel, with reflected and direct signals coherently superimposed at the receiver, easily leading to deep fading and inter-symbol interference. Simultaneously, co-channel or adjacent-channel signals from civilian wireless equipment and other nearby communication systems also constitute external interference. Existing systems based on conventional Wi-Fi or public... The image transmission scheme for drones in mobile communication networks is not optimized for such harsh non-line-of-sight, highly reflective, and dynamically changing channels in terms of protocol design and physical layer technology. Its resistance to multipath fading and interference is limited. In addition, due to limitations in payload and power consumption, the communication modules carried by drones typically have low transmit power and antenna gain, which further exacerbates the instability of the link under long-distance transmission. The direct consequence is that in complex scenarios, the returned video stream will experience frequent bit rate fluctuations, high bit error rates, or even transmission interruptions. This leads to the monitoring center receiving images that are choppy, have increased pixelation, or lose real-time performance. It cannot provide continuous and reliable visual information support for critical tasks such as pollution incident early warning, illegal behavior evidence collection, and emergency command, which seriously restricts the practical effectiveness and reliability of the entire inspection system. Summary of the Invention

[0004] This invention addresses the technical problems existing in the prior art by providing a method for real-time transmission of river and lake inspection images based on unmanned aerial vehicle (UAV) swarms, thereby resolving the issues raised in the background section.

[0005] The technical solution of this invention to solve the above-mentioned technical problems is as follows: specifically, it includes the following steps:

[0006] Step S1: During the inspection mission of the UAV swarm, collect historical channel state information between each UAV node in the swarm; based on the collected historical channel state information, construct an input tensor containing time and space dimensions, and input it into a pre-trained first neural network model; the first neural network model uses its internal model weight parameters to process the input tensor and outputs the predicted channel gain matrix of each potential communication link within a specified future time window.

[0007] Step S2: Using the predicted channel gain matrix obtained in step S1 as the environmental state input, the UAV swarm is modeled as a distributed multi-agent system; by executing the second decision algorithm, each UAV agent makes collaborative decisions based on its internal policy network parameters to generate a set of multi-hop optimal transmission paths for image data transmission with spatial and frequency diversity effects.

[0008] Step S3: On the set of multi-hop optimal transmission paths determined in step S2, obtain the real-time measured signal-to-noise ratio (SNR) value for each hop transmission link; fuse the real-time measured SNR value with the predicted channel gain value of the corresponding link in the predicted channel gain matrix obtained in step S1, and determine the final modulation and coding scheme level for the current data transmission through an adaptive selection mechanism.

[0009] Step S4: When data transmission fails and retransmission is required, a cooperative retransmission protocol is initiated on the multi-hop optimal transmission path set determined in step S2. The relay UAV participating in the retransmission performs linear network coding on the stored data packets and sends incremental redundant packets. At the same time, acknowledgment signals and updated channel measurement values ​​are collected during the transmission process. The acknowledgment signals and updated channel measurement values ​​are used as online learning data and fed back to the first neural network model in step S1 and the second decision algorithm in step S2, respectively, to update the model weight parameters of the first neural network model and the policy network parameters of the UAV agent in the second decision algorithm.

[0010] In a preferred embodiment, in step S1, the historical channel state information includes received signal strength indication, root mean square delay spread, carrier frequency offset, three-dimensional spatial coordinates of the source UAV node and the three-dimensional spatial coordinates of the destination UAV node; for any communication link, the historical channel state information at the time of acquisition constitutes a five-dimensional feature vector.

[0011] In a preferred embodiment, the process of constructing the input tensor containing time and space dimensions is as follows:

[0012] The total number of nodes in the UAV swarm is set as a first quantity, and the window length for maintaining historical moment data is set as a second quantity. At each decision moment, a four-dimensional tensor is constructed as the input tensor. The first dimension of this four-dimensional tensor has a length of the second quantity, corresponding to a continuous sequence of historical moments. The second and third dimensions both have a length of the first quantity, corresponding to the UAV node index as the signal transmitter and the UAV node index as the signal receiver, respectively. The fourth dimension has a length of five, corresponding to the five elements of the five-dimensional feature vector. For each element of this four-dimensional tensor, it is determined by the first, second, and third dimension indices. This element is a five-dimensional vector, and its value is the five-dimensional feature vector composed of historical channel state information collected from the transmitting node indicated by the second dimension index to the receiving node indicated by the third dimension index at the historical moment indicated by the first dimension index before the decision moment.

[0013] In a preferred embodiment, the specific process by which the first neural network model processes the input tensor using its internal model weight parameters is as follows:

[0014] The processing steps of the first neural network model include spatial attention encoding, spatial feature aggregation, temporal convolution prediction, and link gain calculation. The model weight parameters include a trainable weight matrix and a trainable attention parameter vector shared in spatial attention encoding and spatial feature aggregation, as well as a trainable weight vector and bias parameters inside the fully connected layer in the link gain calculation.

[0015] The spatial attention encoding utilizes a trainable weight matrix and attention parameter vector to calculate and normalize normalized attention weights that reflect the importance between nodes. The spatial feature aggregation performs a weighted summation of the features of neighboring nodes based on the normalized attention weights to obtain the spatial feature representation of each node. The temporal convolutional network processes the spatial feature representation of each node in the time series and outputs its predicted state feature vector. The link gain calculation concatenates the predicted state feature vectors of any node pair and maps them through a fully connected layer to the predicted channel gain value of the link between the node pair. The predicted channel gain values ​​of all node pairs constitute the predicted channel gain matrix.

[0016] In a preferred embodiment, the specific process of modeling the drone swarm as a distributed multi-agent system in step S2 is as follows:

[0017] Each UAV node in the UAV swarm is defined as an independent agent. A local observation state is constructed for each agent. The local observation state consists of: the predicted channel gain matrix obtained in step S1, the real-time remaining energy of the UAV node corresponding to the agent, the real-time data queue length of the UAV node, the real-time three-dimensional spatial coordinates of the UAV node, and partial state information of other neighboring UAV nodes within the communication range of the agent. The partial state information includes at least the three-dimensional spatial coordinates of the neighboring UAV nodes and their corresponding predicted channel gain values ​​in the predicted channel gain matrix between them and the UAV node corresponding to the current agent. The action space of each agent is the set of all its communicable next-hop nodes, which includes other UAV nodes and ground data aggregation nodes. All agents share a parameterized policy network. The policy network includes an encoding layer, an attention module, and an output layer. The encoding layer is used to map the local observation state of each agent to an observation feature vector. The attention module is used to generate interaction weights based on the observation feature vectors between agents. The output layer is used to output the next-hop selection probability distribution vector corresponding to the agent based on the agent's own observation feature vector and the information of other agents aggregated by the interaction weights.

[0018] In a preferred embodiment, the specific process of generating a set of multi-hop optimal transmission paths for image data transmission with spatial and frequency diversity effects by executing the second decision algorithm and making collaborative decisions based on their internal policy network parameters is as follows:

[0019] The policy network adopts a centralized training and distributed execution architecture. The policy network parameters include trainable parameters in the encoding layer, attention module, and output layer. The attention module specifically includes a first trainable linear transformation matrix, a second trainable linear transformation matrix, and a trainable attention parameter vector.

[0020] The process of the attention module generating interaction weights includes the following steps:

[0021] First, the local observation state of each agent is mapped to an observation feature vector through an encoding layer;

[0022] Next, for any pair of agents that need to calculate interaction weights, they are defined as the current agent and the target agent, respectively. The observation feature vector of the current agent is projected through the first trainable linear transformation matrix to obtain the first projected feature vector. At the same time, the observation feature vector of the target agent is projected through the second trainable linear transformation matrix to obtain the second projected feature vector.

[0023] Then, the first projection feature vector and the second projection feature vector are concatenated to form a combined feature vector;

[0024] Subsequently, the dot product of the trainable attention parameter vector and the combined feature vector is calculated to obtain an original attention score.

[0025] Next, the original attention score is input into a LeakyReLU nonlinear activation function for processing to obtain the processed attention score;

[0026] Finally, taking the current agent as the benchmark, the processed attention scores calculated by the agent, all neighboring agents, and itself are normalized so that the sum of all scores is one. The normalized value is the interaction weight of the current agent to the target agent.

[0027] The output layer of the policy network uses interaction weights to weighted aggregate the observed feature vectors of all agents, and calculates the action value estimate of each agent's available actions in its action space based on the aggregated information and its own features. The probability of an agent choosing a specified action in its action space is determined by the probability distribution of its action value estimate after temperature coefficient adjustment.

[0028] The collaborative decision-making of the agents is guided by a global reward function, which calculates a scalar reward value at each decision moment. This scalar reward value is a weighted sum of four sub-rewards. The first sub-reward is the product of the first weighting coefficient and the logarithm of the predicted channel gain of all communication links that successfully completed data transmission in the current decision moment plus one. The second sub-reward is the negative of the product of the second weighting coefficient and the sum of the end-to-end delays of all data packets transmitted by all agents. The third sub-reward is the negative of the product of the third weighting coefficient and the sum of the energy consumption of all UAV nodes in the current decision period. The fourth sub-reward is the product of the fourth weighting coefficient and the redundancy evaluation value calculated based on the structural characteristics of the generated set of optimal multi-hop transmission paths.

[0029] By optimizing the policy network parameters to maximize the long-term accumulated global reward function value, the agents can collaboratively generate a set of multi-hop optimal transmission paths.

[0030] In a preferred embodiment, the specific process of fusing the real-time measured signal-to-noise ratio value with the predicted channel gain value of the corresponding link in the predicted channel gain matrix obtained in step S1 is as follows:

[0031] For each hop transmission link in the multi-hop optimal transmission path set, perform the following operations:

[0032] First, obtain the real-time measured signal-to-noise ratio value of the current transmission link;

[0033] At the same time, query the predicted channel gain value corresponding to the current transmission link from the predicted channel gain matrix obtained in step S1;

[0034] Next, the predicted channel gain value obtained from the query is converted into a predicted signal-to-noise ratio value;

[0035] Then, an equivalent signal-to-noise ratio (SNR) value for decision-making is calculated. The calculation process is as follows: a weighted sum is performed on the real-time measured SNR value of the current transmission link and the predicted SNR value of the link to obtain a basic fusion value. Based on this, a trend bias value is introduced. The calculation process of this trend bias value includes the following steps: First, the slope of the predicted SNR change is calculated and denoted as the slope value. The slope value is calculated as follows: a time interval is determined, the predicted SNR value at the current moment and the predicted SNR value at a specified future moment are obtained, and the difference between these two predicted SNR values ​​is divided by the time interval. The result is... The first step is to determine the slope value; the second step is to determine the sign function value; if the slope value is positive, the sign function value is positive one; if the slope value is negative, the sign function value is negative one; if the slope value is zero, the sign function value is zero; the third step is to calculate the trend bias magnitude; the trend bias magnitude is the minimum of the absolute value of the slope value, a positive real parameter called the decision look-ahead time, and a positive real parameter called the maximum bias magnitude; the fourth step is to calculate the final trend bias value; the trend bias value is obtained by multiplying the sign function value, a second adjustable parameter called the trend term gain coefficient, and the trend bias magnitude.

[0036] The base fusion value is added to the trend bias value to obtain the equivalent signal-to-noise ratio value.

[0037] In a preferred embodiment, the specific process of determining the final modulation and coding scheme level for the current data transmission is as follows:

[0038] A set of modulation and coding scheme levels is preset, and a minimum signal-to-noise ratio (SNR) threshold value is set for each level. After obtaining the equivalent SNR value of each hop transmission link, the final modulation and coding scheme level is determined based on the equivalent SNR value and the set of SNR threshold values. The determination process is as follows: from all preset modulation and coding scheme levels, levels whose corresponding SNR threshold values ​​are less than or equal to the current equivalent SNR value are selected to form a candidate level set. Then, from the candidate level set, the level with the largest value is selected as the final determined modulation and coding scheme level. Through this mechanism, a modulation and coding scheme level that matches the channel quality is determined for each active link in the multi-hop optimal transmission path set.

[0039] In a preferred embodiment, the specific process of initiating the cooperative retransmission protocol in step S4 is as follows:

[0040] When data transmission fails and retransmission is required, based on the set of optimal multi-hop transmission paths determined in step S2, a set of relay UAV nodes suitable for collaborative retransmission is identified. For each relay UAV node in this set, multiple data packets related to the data block to be retransmitted are selected from its buffer. Each participating relay UAV node independently and randomly generates a number of random coding coefficients equal to the number of selected data packets. Each random coding coefficient is uniformly and randomly selected from a predefined finite field of sufficient size. Then, these random coding coefficients are used to linearly combine the selected data packets. Specifically, the linear combination involves multiplying the first selected data packet by the first random coding coefficient, multiplying the second selected data packet by the second random coding coefficient, and so on, until the last data packet is multiplied by the last random coding coefficient. Then, all the results of the multiplications are processed... The process involves summing the results to generate an incremental redundancy packet. Each relay UAV node encapsulates its generated incremental redundancy packet and all its corresponding random coding coefficients, and then transmits it along the path specified by the multi-hop optimal transmission path set determined in step S2, from the node to the destination node. During transmission, the modulation and coding scheme level determined for each hop of the path in step S3 is used, or a more robust modulation and coding scheme level is selected for retransmission. Upon receiving a new incremental redundancy packet, the receiving end attempts to jointly decode it with the packet and its carried random coding coefficients, along with previously received incremental redundancy packets and coding coefficients, to recover the original data block. If decoding is successful, a positive acknowledgment signal is sent to the source node and related relay nodes. If decoding is unsuccessful after reaching the preset maximum number of retransmissions or time limit, the retransmission is terminated.

[0041] In a preferred embodiment, the specific process of collecting acknowledgment signals and updated channel measurements during transmission, and feeding them back as online learning data to the first neural network model in step S1 and the second decision algorithm in step S2, is as follows:

[0042] In each data transmission attempt, including the initial transmission and any retransmission, two types of online learning data are collected: the first is acknowledgment signals, which are positive or negative acknowledgment signals that indicate the success or failure of this transmission attempt, and the corresponding transmission time and the link identifier involved are recorded; the second is updated channel measurements, which are a set of real-time channel state information measured and recorded at each data packet transmission or reception time, including at least the following measurement data arranged in order: the real-time measured signal-to-noise ratio, the real-time measured root mean square delay spread, the real-time measured carrier frequency offset, and the three-dimensional spatial coordinates of the transmitting UAV node and the receiving UAV node at the measurement time.

[0043] For updating the first neural network model in step S1, the collected real-time channel state information arranged in chronological order is combined with the historical channel state information used to construct the input tensor in step S1 to form a training sample set; an online time-series learning algorithm is adopted to improve the accuracy of channel prediction and adjust the model weight parameters of the first neural network model.

[0044] For updating the second decision algorithm in step S2, each complete transmission attempt is considered as a decision round, and an experience data tuple is constructed. This experience data tuple includes: the environment state before the transmission begins, which includes the predicted channel gain matrix output in step S1; the action selected by each UAV agent according to the policy network, i.e., the next-hop node; the immediate reward calculated based on the actual transmission result, which is calculated based on whether the transmission attempt was successful, the end-to-end delay generated by the transmission, and the energy consumption of all participating nodes during the transmission; and the new environment state after the transmission ends. Multiple such experience data tuples are stored in a distributed experience replay buffer, and a distributed reinforcement learning algorithm is used to calculate the policy gradient using the experience data tuples, thereby updating the policy network parameters of the UAV agents in the second decision algorithm.

[0045] The beneficial effects of this invention are as follows: By combining forward-looking channel prediction with intelligent path planning, the reliability of image transmission by UAV swarms in complex river and lake environments is significantly improved. Its adaptive modulation and coding mechanism, which integrates real-time measurement and prediction information, effectively avoids transmission interruptions caused by rapid channel fading, ensuring the continuity and clarity of video images. At the same time, the collaborative network coding retransmission and online learning closed loop based on actual transmission feedback enable the entire system to continuously optimize itself and enhance its adaptability to dynamic environments, thereby achieving efficient and robust real-time transmission of inspection image data while ensuring low latency. Attached Figure Description

[0046] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0047] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0048] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0049] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application. Example

[0050] This embodiment provides, for example Figure 1 The method for real-time transmission of river and lake inspection images based on UAV swarms is shown, and specifically includes the following steps:

[0051] Step S1: During the inspection mission of the UAV swarm, collect historical channel state information between each UAV node in the swarm; based on the collected historical channel state information, construct an input tensor containing time and space dimensions, and input it into a pre-trained first neural network model; the first neural network model uses its internal model weight parameters to process the input tensor and outputs the predicted channel gain matrix of each potential communication link within a specified future time window.

[0052] Step S2: Using the predicted channel gain matrix obtained in step S1 as the environmental state input, the UAV swarm is modeled as a distributed multi-agent system; by executing the second decision algorithm, each UAV agent makes collaborative decisions based on its internal policy network parameters to generate a set of multi-hop optimal transmission paths for image data transmission with spatial and frequency diversity effects.

[0053] Step S3: On the set of multi-hop optimal transmission paths determined in step S2, obtain the real-time measured signal-to-noise ratio (SNR) value for each hop transmission link; fuse the real-time measured SNR value with the predicted channel gain value of the corresponding link in the predicted channel gain matrix obtained in step S1, and determine the final modulation and coding scheme level for the current data transmission through an adaptive selection mechanism.

[0054] Step S4: When data transmission fails and retransmission is required, a cooperative retransmission protocol is initiated on the multi-hop optimal transmission path set determined in step S2. The relay UAV participating in the retransmission performs linear network coding on the stored data packets and sends incremental redundant packets. At the same time, acknowledgment signals and updated channel measurements during the transmission process are collected and used as online learning data. These acknowledgment signals and updated channel measurements are fed back to the first neural network model in step S1 and the second decision algorithm in step S2, respectively, to update the model weight parameters of the first neural network model and the policy network parameters of the UAV agent in the second decision algorithm.

[0055] In this embodiment, it should be specifically noted that the historical channel state information collected in step S1 includes: received signal strength indication, root mean square delay spread, carrier frequency offset, and the three-dimensional spatial coordinates of the UAV node. Specifically, for any two UAV nodes, the historical channel state information collected at the time of collection constitutes a five-dimensional feature vector. This five-dimensional feature vector is composed of a first element, a second element, a third element, a fourth element, and a fifth element in sequence. The first element is the received signal strength indication, used to characterize the signal power attenuation of the communication link; the second element is the root mean square delay spread, used to characterize the signal time-domain dispersion caused by multipath propagation; the third element is the carrier frequency offset, used to characterize the carrier frequency change caused by the relative motion of the UAVs; the fourth element is the three-dimensional spatial coordinates of the source UAV node, used to determine the node's position in space; and the fifth element is the three-dimensional spatial coordinates of the destination UAV node, used to determine the node's position in space.

[0056] In practical implementation, received signal strength indication is usually collected and recorded in decibels and milliwatts; root mean square delay spread is in nanoseconds and is obtained by statistical analysis of channel impulse response; carrier frequency offset is in hertz and can be estimated by the phase difference between the received signal and the local oscillator; the three-dimensional spatial coordinates of the UAV node can be obtained by fusion positioning of the global satellite navigation system and the inertial measurement unit, and is usually expressed in meters to represent its position in the Earth coordinate system. After normalizing the above five physical quantities to a similar numerical range, they together constitute a five-dimensional feature vector characterizing the link status.

[0057] The process of constructing an input tensor that includes time and space dimensions is as follows:

[0058] The total number of nodes in the drone swarm is set as the first quantity, and the window length for maintaining historical time data is set as the second quantity. For example, in a typical inspection formation, the first quantity can be set to 4 to 10 drones, and the window length for maintaining historical time data is the second quantity. The second quantity can be set according to the channel change rate, for example, to 10 to 20 consecutive sampling times, with each sampling interval being 100 milliseconds to 1 second. At each decision time, a four-dimensional tensor is constructed as the input tensor. The first dimension of this four-dimensional tensor has the length of the second quantity, corresponding to the continuous historical time sequence; the second and third dimensions both have the length of the first quantity, corresponding to the drone node index as the signal transmitter and the drone node index as the signal receiver, respectively; the fourth dimension has the length of five, corresponding to the five elements of the five-dimensional feature vector; for each of the four-dimensional tensors... An element, determined by its first, second, and third dimension indices, is a five-dimensional vector. Its value is a five-dimensional feature vector composed of historical channel state information collected from the transmitting node indicated by the second dimension index to the receiving node indicated by the third dimension index at the historical time indicated by the first dimension index before the decision time. The unit of received signal strength is decibel-milliwatt, the unit of root mean square delay spread is nanosecond, and the unit of carrier frequency offset is Hertz. The three-dimensional spatial coordinates are represented by the geodetic coordinate system or local rectangular coordinate system, and the unit is usually meters. Through this structured organization, the input tensor associates multiple dimensions such as time, transmitting node, receiving node, and channel characteristics to form a complete spatiotemporal data block, providing a uniform and complete input for subsequent neural network models.

[0059] The specific process by which the first neural network model processes the input tensor using its internal model weight parameters is as follows:

[0060] The processing steps of the first neural network model include spatial attention encoding, spatial feature aggregation, temporal convolution prediction, and link gain calculation. The model weight parameters include a trainable weight matrix and a trainable attention parameter vector shared in spatial attention encoding and spatial feature aggregation, as well as a trainable weight vector and bias parameters inside the fully connected layer in the link gain calculation.

[0061] The shared trainable weight matrix is ​​used to project the five-dimensional feature vector into a higher-dimensional hidden feature space. The dimension can be set according to the model complexity, such as mapping the five-dimensional features to 16-dimensional or 32-dimensional hidden features. The trainable attention parameter vector is used to calculate the importance weights between nodes, and its dimension is twice that of the hidden feature dimension. The trainable weight vector and bias parameters inside the fully connected layer are used to map the concatenated joint feature vector into the final predicted channel gain scalar value.

[0062] In the spatial attention encoding process, for each historical moment indicated by the first-dimensional index in the input tensor, the first neural network model treats the UAV swarm communication network at the current moment as a graph. Nodes in the graph represent UAVs, and directed edges from sending nodes to receiving nodes represent communication links. The feature of each directed edge is a five-dimensional feature vector determined from the input tensor by the first-dimensional index of that historical moment, the second-dimensional index of the sending node, and the third-dimensional index of the receiving node. For any UAV acting as a receiving node and any sending node acting as its neighbor, the first neural network model calculates a spatial attention coefficient. When calculating the spatial attention coefficient, firstly, using a shared trainable weight matrix, linear transformations are performed on the five-dimensional feature vectors corresponding to the receiving node and the sending node, respectively, to obtain the transformed features of the receiving node and the transformed features of the sending node. Next, the transformed features of the receiving node and the transformed features of the sending node are concatenated to form a combined feature vector. Then, the dot product of the trainable attention parameter vector and the combined feature vector is calculated to obtain the original attention score. Finally, the original attention score is input into a LeakyReLU activation function, which multiplies negative inputs by a fixed small slope (e.g., 0.01) while keeping positive values ​​unchanged to introduce non-linearity. The output value is the spatial attention coefficient between the receiving node and the sending node. After calculating the spatial attention coefficients with the same UAV node as the receiving node and all other UAV nodes as the sending nodes, all spatial attention coefficients are normalized. Specifically, the softmax function is used for normalization so that the sum of the normalized attention weights of all nodes with the same node as the receiving node is 1, thus obtaining the normalized attention weights corresponding to each sending node.

[0063] During the spatial feature aggregation process, for each UAV acting as a receiving node in the graph, the first neural network model performs a weighted summation of the features of all neighboring sending nodes pointing to that receiving node, based on all normalized attention weights obtained during the spatial attention encoding process for that node as the receiving node. During the weighted summation, the contribution value of each neighboring sending node is the result of multiplying the normalized attention weight corresponding to that neighboring sending node by the shared trainable weight matrix and the five-dimensional feature vector corresponding to that neighboring sending node obtained from the input tensor. A non-linear activation function, such as the ReLU function or the Sigmoid function, is applied to the weighted summation result to obtain the spatial feature representation of the receiving node aggregating neighborhood information at the current historical moment. The dimension of the spatial feature representation is the same as the dimension of the hidden features, for example, 16 or 32 dimensions.

[0064] In the temporal convolutional prediction process, for each UAV node, the first neural network model represents the spatial features of the node at all historical moments, arranged in chronological order to form the node's time series. The first neural network model uses a temporal convolutional network to process the time series of each node separately. The temporal convolutional network captures multi-scale temporal dependencies in the time series through causal convolution operations with dilated structures. The temporal convolutional network can contain multiple convolutional layers, and the dilation factor of each layer can grow exponentially (e.g., 1, 2, 4, 8) to expand the receptive field and capture long-term dependencies, and outputs the predicted state feature vector of each node within a specified future time window. The dimension of the predicted state feature vector can be set according to design needs, such as 8-dimensional or 16-dimensional, which integrates the state evolution information of the node in the temporal dimension.

[0065] During the link gain calculation process, for any pair of nodes for which channel gain needs to be predicted, the first neural network model concatenates the predicted state feature vector of the sending node with the predicted state feature vector of the receiving node to form a joint feature vector. The joint feature vector is then input into a fully connected layer. The fully connected layer performs a linear transformation on the joint feature vector through its internal trainable weight vector and bias parameters, and finally outputs a scalar value. This scalar value is a dimensionless relative power gain estimate, which can be directly used for signal-to-noise ratio conversion. This scalar value is the predicted channel gain value of the future link from the sending node to the receiving node. The predicted channel gain values ​​calculated between all node pairs together constitute the predicted channel gain matrix. The predicted channel gain matrix is ​​an N×N square matrix, where N is the total number of UAV nodes, and the element in the i-th row and j-th column represents the predicted channel gain from node i to node j.

[0066] In this embodiment, the specific process of modeling the drone swarm as a distributed multi-agent system in step S2 is as follows:

[0067] Each UAV node in the UAV swarm is defined as an independent agent. A local observation state is constructed for each agent. The local observation state consists of: the predicted channel gain matrix obtained in step S1, the real-time remaining energy of the UAV node corresponding to the agent, the real-time data queue length of the UAV node, the real-time three-dimensional spatial coordinates of the UAV node, and partial state information of other neighboring UAV nodes within the communication range of the agent. The partial state information includes at least the three-dimensional spatial coordinates of the neighboring UAV nodes and their corresponding predicted channel gain values ​​in the predicted channel gain matrix between them and the UAV node corresponding to the current agent. The action space of each agent is the set of all its communicable next-hop nodes, which includes other UAV nodes and ground data aggregation nodes. All agents share a parameterized policy network. The policy network includes an encoding layer, an attention module, and an output layer. The encoding layer is used to convert each agent's input into output. The local observation state is mapped to an observation feature vector. The encoding layer can be a multilayer perceptron, whose input dimension matches the dimension of the local observation state. The output dimension, i.e., the length of the observation feature vector, can be set to 32 or 64 dimensions to carry sufficient state information. The attention module is used to generate interaction weights based on the observation feature vectors between agents. The interaction weights are used to quantify the importance of the state information of other agents to this agent during decision-making. The output layer is used to output the next-hop selection probability distribution vector corresponding to the agent based on the agent's own observation feature vector and the information of other agents aggregated by the interaction weights. Each probability value in the next-hop selection probability distribution vector is associated with a node in the action space, which represents the probability that the agent chooses the associated node as the next hop for data packet forwarding. The next-hop selection probability distribution vector is obtained by normalizing the original logical value calculated by the output layer through the softmax function to ensure that the sum of all probability values ​​is one.

[0068] The specific process by which each UAV agent collaboratively makes decisions based on its internal policy network parameters to generate a set of multi-hop optimal transmission paths for image data transmission, possessing spatial and frequency diversity effects, through the execution of the second decision algorithm is as follows:

[0069] The policy network employs a centralized training and distributed execution architecture. Its training phase relies on a centralized critic network. This centralized critic network takes the aggregated information of the observed feature vectors of all agents and the global state as input, and outputs an estimate of the current global state value, used to calculate the advantage function to guide the policy network's parameter updates. The policy network parameters include trainable parameters in the encoding layer, attention module, and output layer. The attention module specifically includes a first trainable linear transformation matrix, a second trainable linear transformation matrix, and a trainable attention parameter vector. The first and second trainable linear transformation matrices project the observed feature vectors onto a common subspace for attention computation. The projected dimension, i.e., the length of the first and second projected feature vectors, is typically set to 16 dimensions or the same dimension as the observed feature vectors. The length of the trainable attention parameter vector is equal to the length of the combined feature vector obtained by concatenating the first and second projected feature vectors, i.e., twice the projected dimension.

[0070] The process of the attention module generating interaction weights includes the following steps:

[0071] First, the local observation state of each agent is mapped to an observation feature vector through an encoding layer;

[0072] Next, for any pair of agents that need to calculate interaction weights, they are defined as the current agent and the target agent, respectively. The observation feature vector of the current agent is projected through the first trainable linear transformation matrix to obtain the first projected feature vector. At the same time, the observation feature vector of the target agent is projected through the second trainable linear transformation matrix to obtain the second projected feature vector.

[0073] Then, the first projection feature vector and the second projection feature vector are concatenated to form a combined feature vector;

[0074] Subsequently, the dot product of the trainable attention parameter vector and the combined feature vector is calculated to obtain an original attention score.

[0075] Next, the original attention score is input into a LeakyReLU nonlinear activation function for processing to obtain the processed attention score. The fixed slope coefficient of the LeakyReLU activation function for negative inputs can usually be set between 0.01 and 0.2, for example, 0.1, to avoid gradient vanishing while introducing nonlinearity.

[0076] Finally, taking the current agent as the benchmark, the processed attention scores calculated by the agent, all neighboring agents, and itself are normalized so that the sum of all scores is one. The normalized value is the interaction weight of the current agent to the target agent. The normalization process uses the softmax function, which is based on the natural constant e. It performs an exponential operation on the processed attention scores and normalizes them, ensuring the non-negativity of the interaction weights and the property that the sum is 1. Thus, it can be regarded as a weighted probability distribution.

[0077] The output layer of the policy network uses interaction weights to weighted aggregate the observed feature vectors of all agents, and calculates the action value estimate of each agent's available actions in its action space based on the aggregated information and its own features. The probability of an agent selecting a specific action in its action space is determined by the action value estimate corresponding to that action. Specifically, the calculation process for this probability is as follows: First, the action value estimates corresponding to all available actions of the current agent are divided by a positive real number parameter called the temperature coefficient, and the resulting quotients are then subjected to an exponential operation with the natural constant as the base; then, the currently selected action... The probability of choosing an action is the ratio obtained by dividing the result of the exponential operation by the sum of the results of the exponential operations of all possible actions. The temperature coefficient is used to control the degree of randomness in the agent's decision-making. It is set to a large value in the early stage of training to encourage exploration, and a small value in the later stage of training to make the policy tend to stabilize. The temperature coefficient can be set in the range of 1.0 to 5.0 in the early stage of training. As the number of training steps increases, it can be gradually reduced to the range of 0.1 to 0.5 by using linear decay or exponential decay to reduce randomness and make the policy tend to select high-value actions deterministically.

[0078] The collaborative decision-making of the agents is guided by a global reward function, which calculates a scalar reward value at each decision moment. This scalar reward value is a weighted sum of four sub-rewards. The first sub-reward is the product of the first weighting coefficient and the logarithm of the predicted channel gain of all communication links that successfully complete data transmission in the current decision moment plus one. The second sub-reward is the negative of the product of the second weighting coefficient and the sum of the end-to-end delays of all data packets transmitted by the agents. The third sub-reward is the negative of the product of the third weighting coefficient and the sum of the energy consumption of all UAV nodes in the current decision period. The fourth sub-reward is the product of the fourth weighting coefficient and the redundancy evaluation value calculated based on the structural characteristics of the generated multi-hop optimal transmission path set. One way to calculate the redundancy evaluation value is to count the maximum number of paths that exist between any source UAV node and the ground data aggregation node in the multi-hop optimal transmission path set, without sharing any intermediate UAV nodes. The quantity is used as the path redundancy of the source node. The path redundancy of all source nodes is then summed or averaged to obtain the redundancy evaluation value. The larger the value, the stronger the network's survivability under single-point or single-link failure. Among them, the first weighting coefficient, the second weighting coefficient, the third weighting coefficient, and the fourth weighting coefficient are all non-negative real numbers, used to adjust the relative importance of each sub-reward in the total reward. In specific implementation, these weighting coefficients are adjusted manually through grid search or based on experience. For example, the first weighting coefficient can be set between 0.5 and 2.0 to emphasize the use of high-channel-quality links, the second weighting coefficient can be set between 0.01 and 0.1 to moderately penalize latency, the third weighting coefficient can be set between 0.001 and 0.01 to slightly penalize energy consumption, and the fourth weighting coefficient can be set between 0.1 and 1.0 to encourage a certain degree of path redundancy. The specific ratios of these coefficients reflect the trade-offs between different objectives such as transmission reliability, real-time performance, energy efficiency, and robustness in the system design.

[0079] By optimizing the policy network parameters to maximize the long-term accumulated global reward function value, agents collaboratively generate a set of multi-hop optimal transmission paths. The optimization process typically employs reinforcement learning algorithms based on policy gradients, such as the near-end policy optimization algorithm or the deep deterministic policy gradient algorithm. In the centralized training phase, the policy gradient is calculated and the parameters of the policy network are updated using the collected interaction data sequence (state, action, reward, next state) and the critic network. After training, the policy network parameters are fixed. In the distributed execution phase, each agent makes independent decisions based only on its local observed state and the policy network, without interacting with the critic network, thereby achieving low-latency online path planning.

[0080] In this embodiment, it is specifically necessary to explain the process in step S3 of fusing the real-time measured signal-to-noise ratio value with the predicted channel gain value of the corresponding link in the predicted channel gain matrix obtained in step S1 as follows:

[0081] For each hop transmission link in the multi-hop optimal transmission path set, perform the following operations:

[0082] First, obtain the real-time measured signal-to-noise ratio value of the current transmission link;

[0083] At the same time, query the predicted channel gain value corresponding to the current transmission link from the predicted channel gain matrix obtained in step S1;

[0084] Next, the predicted channel gain value is converted into a predicted signal-to-noise ratio (SNR) value. The conversion process is as follows: the result of applying a base-10 logarithmic function to the predicted channel gain value is calculated, multiplied by 10, and then a fixed constant offset determined by the system transmit power and noise power spectral density is added. Finally, the predicted SNR value of the link in decibels is obtained. The fixed constant offset is pre-calculated and determined during system deployment based on the transmitter's nominal power, receiver noise figure, and path loss reference model to ensure that the predicted SNR value and the measured SNR value are within the same dimension and comparable range.

[0085] Then, an equivalent signal-to-noise ratio (SNR) value for decision-making is calculated. The calculation process involves weighted summation of the real-time measured SNR value of the current transmission link and the predicted SNR value of that link, yielding a basic fusion value. The weighting factor used in the weighted summation is a first adjustable parameter, with a value ranging from zero to one. This first adjustable parameter is used to balance the confidence level between the current instantaneous channel state and the future predicted state. For example, in scenarios where channel changes are relatively slow, this parameter can be set to 0.7 to 0.9 to rely more on real-time measurements; in scenarios where the channel changes rapidly or the prediction model has high confidence, this parameter can be set to 0.3 to 0.6 to give the prediction... Information is given greater weight; based on this, a trend bias value is introduced; the calculation process of this trend bias value includes the following steps: First, calculate the slope of the predicted signal-to-noise ratio change, denoted as the slope value; the slope value is calculated as follows: determine a time interval, obtain the predicted signal-to-noise ratio value at the current moment and the predicted signal-to-noise ratio value at a specified future moment, divide the difference between these two predicted signal-to-noise ratio values ​​by the time interval, and the result is the slope value, the unit of which is decibels per second; the time interval is usually set to an integer multiple of a transmission time interval, such as 10 milliseconds to 50 milliseconds, to capture the short-term changing trend of the channel; Second, determine the symbol function value; if the slope value is positive, the symbol function value is positive one; if the slope value is negative, the symbol function value is negative. If the rate value is negative, the sign function value is negative one; if the slope value is zero, the sign function value is zero. The third step is to calculate the trend bias amplitude. The trend bias amplitude is the minimum of the absolute value of the slope, a positive real-valued parameter called the decision look-ahead time, and a positive real-valued parameter called the maximum bias amplitude. The decision look-ahead time represents the length of time for forward prediction during decision-making, measured in seconds. The decision look-ahead time is typically set to the length of 1 to 3 data transmission slots in the future, for example, 100 milliseconds to 300 milliseconds, allowing the decision to react in advance to channel changes over a future period. The maximum bias amplitude is used to limit the adjustment range of the trend bias value, measured in decibels. The signal-to-noise ratio tolerance between adjacent levels of the link adaptive modulation and coding scheme can be set, for example, to 3 to 6 dB, to avoid excessive level jumps due to over-adjustment; the fourth step is to calculate the final trend offset value; the trend offset value is obtained by multiplying the symbol function value, a second adjustable parameter called the trend term gain coefficient, and the trend offset amplitude; the trend term gain coefficient ranges from zero to one; the trend term gain coefficient is used to control the strength of the trend effect, and can be set to about 0.5 in the initial stage, and fine-tuned according to the system's performance in the real environment; when a regular rapid deep fading is detected in the channel, this coefficient can be appropriately increased to enhance the forward avoidance capability;

[0086] The base fusion value is added to the trend bias value to obtain the equivalent signal-to-noise ratio (SNR). This equivalent SNR is a comprehensive quality indicator that integrates current measurements, future predictions, and trends. Its core innovation lies in the fact that when the prediction indicates that the channel is about to deteriorate (the slope is negative), the trend bias value is negative, and the equivalent SNR will be lower than the current fusion value. This drives the system to select a more robust low-order modulation and coding scheme in advance, effectively avoiding continuous transmission failures caused by using a high-order scheme when fading occurs. Conversely, when the predicted channel improves, it can create conditions for upgrading to a high-order scheme in advance, thereby improving the average spectral efficiency.

[0087] The specific process for determining the final modulation and coding scheme level for the current data transmission is as follows:

[0088] A set of modulation and coding scheme levels is preset, and a minimum signal-to-noise ratio (SNR) threshold value is set for each level. The higher the level, the higher the required SNR threshold value. The set of modulation and coding scheme levels can be selected or customized based on communication standards (such as the definitions in Wi-Fi, LTE, and 5G NR), typically including various combinations ranging from low-order BPSK or QPSK with low code rates to high-order 64-QAM or 256-QAM with high code rates. The SNR threshold value for each level is obtained through offline link-level simulation or theoretical calculation, with a certain safety margin, such as 1 to 2 dB, reserved to cope with the deviation between the actual channel and the ideal model. After obtaining the equivalent SNR value of each hop transmission link, the final modulation and coding scheme level is determined based on the equivalent SNR value and a set of SNR threshold values. The determination process is as follows: from all preset modulation and coding scheme levels, levels whose corresponding SNR threshold values ​​are less than or equal to the current equivalent SNR value are selected to form a set of modulation and coding scheme levels. A candidate level set is generated; then, the level with the largest value is selected from the candidate level set as the final determined modulation and coding scheme level. This "maximum feasible level" selection principle ensures that the instantaneous data transmission rate of the link is maximized while satisfying transmission reliability (equivalent signal-to-noise ratio not lower than the threshold). Through this mechanism, a modulation and coding scheme level matching its channel quality is determined for each active link in the multi-hop optimal transmission path set. The level decision result will be directly used to configure the physical layer coding and modulation of the data packets transmitted on the corresponding link. This design realizes a refined, link quality-aware adaptive modulation and coding. Compared with the traditional method that only relies on real-time signal-to-noise ratio, this invention introduces predictive information, making the adjustment of physical layer parameters more forward-looking and robust, which can smooth the performance jitter caused by rapid channel fluctuations, thereby providing a solid physical layer guarantee for the real-time, reliable, and efficient transmission of image data throughout the multi-hop path.

[0089] In this embodiment, the specific process of initiating the cooperative retransmission protocol in step S4 is as follows:

[0090] When data transmission fails and retransmission is required, a set of relay UAV nodes is identified based on the multi-hop optimal transmission path set determined in step S2. This set of relay UAV nodes consists of UAV nodes located upstream of the link where the transmission failure occurred, which have successfully cached the data block to be retransmitted or related encoded packets. For each relay UAV node in this set, multiple data packets related to the data block to be retransmitted are selected from its cache. Each participating relay UAV node independently and randomly generates a number of random encoding coefficients equal to the number of selected data packets. Each random encoding coefficient is generated from a predefined... The random coding coefficients are uniformly and randomly selected from a sufficiently large finite field. Then, these random coding coefficients are used to linearly combine the selected data packets. Linear combination means multiplying the first selected data packet by the first random coding coefficient, multiplying the second selected data packet by the second random coding coefficient, and so on, until the last data packet is multiplied by the last random coding coefficient. Then, all the multiplication results are summed to generate an incremental redundancy packet. Each relay UAV node encapsulates its generated incremental redundancy packet and all its corresponding random coding coefficients, and then, according to the multi-hop optimal transmission path set determined in step S2, follows the path from the node to the destination node, defined in step S2. The transmission is performed along the paths defined by the determined set of optimal multi-hop transmission paths. During transmission, the modulation and coding scheme level determined for each hop of the path in step S3 is used, or a more robust modulation and coding scheme level is selected for retransmission. A "more robust modulation and coding scheme level" usually refers to reducing the original level by one or two levels, using a lower modulation order and / or a lower coding rate, such as downgrading from 64-QAM to 16-QAM, sacrificing some spectral efficiency for a higher retransmission success rate. The preset maximum number of retransmissions is usually set to 3 to 5 times, and the time limit is set according to the real-time requirements of the image data, such as 300 milliseconds to 1 second. 000 milliseconds; After receiving a new incremental redundancy packet, the receiving end uses the packet and its carried random coding coefficients to perform joint decoding with previously received incremental redundancy packets and coding coefficients to recover the original data block. Joint decoding usually uses Gaussian elimination or iterative belief propagation algorithm. At the receiving end, a system of linear equations with random coding coefficients is constructed. When the number of received linearly independent equations is equal to or greater than the number of data packets contained in the original data block, decoding is successful. If decoding is successful, a positive acknowledgment signal is sent to the source node and related relay nodes. If decoding is still unsuccessful after reaching the preset maximum number of retransmissions or time limit, the current retransmission is terminated.

[0091] The specific process of collecting acknowledgment signals and updated channel measurements during transmission, and feeding them back as online learning data to the first neural network model in step S1 and the second decision algorithm in step S2 is as follows:

[0092] In each data transmission attempt, including the initial transmission and any retransmission, two types of online learning data are collected: the first is an acknowledgment signal, which is a positive or negative acknowledgment signal that indicates the success or failure of this transmission attempt, and the corresponding transmission time and the link identifier involved are recorded; the second is updated channel measurement values, which are a set of real-time channel state information measured and recorded at each data packet transmission or reception time. The types of physical quantities contained in this real-time channel state information are the same as those in the historical channel state information collected in step S1, and include at least the following measurement data arranged in sequence: the real-time measured signal-to-noise ratio value, the real-time measured root mean square delay spread, the real-time measured carrier frequency offset, and the three-dimensional spatial coordinates of the transmitting UAV node and the receiving UAV node at the measurement time.

[0093] For updating the first neural network model in step S1, the collected real-time channel state information (as updated channel measurements) arranged in chronological order, together with the historical channel state information used to construct the input tensor in step S1, constitute the training sample set. An online time-series learning algorithm is employed to improve the accuracy of channel prediction by adjusting the model weight parameters of the first neural network model, enabling it to better fit the latest observed channel dynamics. Specifically, the online time-series learning algorithm can employ mini-batch stochastic gradient descent based on time-series data, or a strategy combining sliding window empirical replay and model fine-tuning. During the update process, the learning rate is typically set to a small value, such as 0.001 to 0.01, to avoid catastrophic forgetting of learned knowledge. The batch size of training samples used in each update can be set to 32 to 128 time-series samples.

[0094] For updating the second decision algorithm in step S2, each complete transmission attempt is considered as a decision round, and an experience data tuple is constructed. This experience data tuple includes: the environmental state before the transmission begins, which includes the predicted channel gain matrix output in step S1; the action selected by each UAV agent according to the policy network, i.e., the next-hop node; the immediate reward calculated based on the actual transmission result, which is calculated based on whether the transmission attempt was successful, the end-to-end delay generated by the transmission, and the energy consumption of all participating nodes during the transmission. The specific calculation formula for the immediate reward can be designed as follows: the reward value equals the first coefficient multiplied by the transmission success flag (1 for success, 0 for failure), minus the second coefficient multiplied by the end-to-end delay (in seconds), and then minus the third coefficient multiplied by the total energy consumption (in joules). Among them, the first, second, and third coefficients are preset positive real numbers, for example, they can be set to 10.0, 0.1, and 0.01 respectively to balance the weights of different optimization objectives; and the new environmental state after the transmission ends. Multiple such experience data tuples are stored in a distributed experience replay buffer. The distributed experience replay buffer can be configured with a total capacity of 10,000 to 50,000 experience tuples, managed using a first-in-first-out (FIFO) strategy. A distributed reinforcement learning algorithm is employed to calculate policy gradients using these experience data tuples, thereby updating the policy network parameters of the UAV agent in the second decision-making algorithm. This optimizes the agent's action selection strategy when facing similar situations in the future. Specifically, the distributed reinforcement learning algorithm can be the asynchronous dominant actor-commentator algorithm or a variant thereof. During the update process, multiple UAV nodes or computing nodes sample small batches of experience data from the experience replay buffer in parallel, independently calculate gradients, and aggregate the gradients asynchronously or synchronously to update the parameters of the central policy network. The learning rate of the policy network is typically set to 0.0001 to 0.001, while the learning rate of the commentator network can be slightly higher. Through this online learning and feedback mechanism, the entire transmission system can continuously optimize its channel prediction model and routing decision strategy using actual transmission experience, forming a complete intelligent closed loop from perception, decision-making, execution to learning. This enables adaptive improvement and long-term stability of image data transmission performance in the dynamic and complex river and lake inspection environment.

[0095] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0096] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0097] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0098] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0099] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0100] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0101] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for real-time transmission of river and lake inspection images based on unmanned aerial vehicle (UAV) swarms, characterized in that, Specifically, the steps include the following: Step S1: During the inspection mission of the UAV swarm, collect historical channel state information between each UAV node in the swarm; based on the collected historical channel state information, construct an input tensor containing time and space dimensions, and input it into a pre-trained first neural network model; the first neural network model uses its internal model weight parameters to process the input tensor and outputs the predicted channel gain matrix of each potential communication link within a specified future time window. Step S2: Using the predicted channel gain matrix obtained in step S1 as the environmental state input, the UAV swarm is modeled as a distributed multi-agent system, specifically as follows: Each UAV node in the UAV swarm is defined as an independent agent. A local observation state is constructed for each agent. The local observation state consists of: the predicted channel gain matrix obtained in step S1, the real-time remaining energy of the UAV node corresponding to the agent, the real-time data queue length of the UAV node, the real-time three-dimensional spatial coordinates of the UAV node, and partial state information of other neighboring UAV nodes within the communication range of the agent. The partial state information includes at least the three-dimensional spatial coordinates of the neighboring UAV nodes and their corresponding predicted channel gain values ​​in the predicted channel gain matrix between them and the UAV node corresponding to the current agent. The action space of each agent is the set of all its communicable next-hop nodes, which includes other UAV nodes and ground data aggregation nodes. All agents share a parameterized policy network. The policy network includes an encoding layer, an attention module, and an output layer. The encoding layer maps the local observation state of each agent to an observation feature vector. The attention module generates interaction weights based on the observation feature vectors between agents. The output layer outputs the next-hop selection probability distribution vector corresponding to the agent based on the agent's own observation feature vector and the information of other agents aggregated by the interaction weights. By executing the second decision-making algorithm, each UAV agent collaboratively makes decisions based on its internal policy network parameters to generate a set of multi-hop optimal transmission paths for image data transmission that possess spatial and frequency diversity effects, specifically: The policy network adopts a centralized training and distributed execution architecture. The policy network parameters include trainable parameters in the encoding layer, attention module, and output layer. The attention module specifically includes a first trainable linear transformation matrix, a second trainable linear transformation matrix, and a trainable attention parameter vector. The process of the attention module generating interaction weights includes the following steps: First, the local observation state of each agent is mapped to an observation feature vector through an encoding layer; Next, for any pair of agents that need to calculate interaction weights, they are defined as the current agent and the target agent, respectively. The observation feature vector of the current agent is projected through the first trainable linear transformation matrix to obtain the first projected feature vector. At the same time, the observation feature vector of the target agent is projected through the second trainable linear transformation matrix to obtain the second projected feature vector. Then, the first projection feature vector and the second projection feature vector are concatenated to form a combined feature vector; Subsequently, the dot product of the trainable attention parameter vector and the combined feature vector is calculated to obtain an original attention score. Next, the original attention score is input into a LeakyReLU nonlinear activation function for processing to obtain the processed attention score; Finally, taking the current agent as the benchmark, the processed attention scores calculated by the agent, all neighboring agents, and itself are normalized so that the sum of all scores is one. The normalized value is the interaction weight of the current agent to the target agent. The output layer of the policy network uses interaction weights to weighted aggregate the observed feature vectors of all agents, and calculates the action value estimate of each agent's available actions in its action space based on the aggregated information and its own features. The probability of an agent choosing a specified action in its action space is determined by the probability distribution of its action value estimate after temperature coefficient adjustment. The collaborative decision-making of the agents is guided by a global reward function, which calculates a scalar reward value at each decision moment. This scalar reward value is a weighted sum of four sub-rewards. The first sub-reward is the product of the first weighting coefficient and the logarithm of the predicted channel gain of all communication links that successfully completed data transmission in the current decision moment plus one. The second sub-reward is the negative of the product of the second weighting coefficient and the sum of the end-to-end delays of all data packets transmitted by all agents. The third sub-reward is the negative of the product of the third weighting coefficient and the sum of the energy consumption of all UAV nodes in the current decision period. The fourth sub-reward is the product of the fourth weighting coefficient and the redundancy evaluation value calculated based on the structural characteristics of the generated set of optimal multi-hop transmission paths. By optimizing the policy network parameters to maximize the long-term accumulated global reward function value, the agents can collaboratively generate a set of multi-hop optimal transmission paths. Step S3: On the set of multi-hop optimal transmission paths determined in step S2, obtain the real-time measured signal-to-noise ratio (SNR) value for each hop transmission link; fuse the real-time measured SNR value with the predicted channel gain value of the corresponding link in the predicted channel gain matrix obtained in step S1, and determine the final modulation and coding scheme level for the current data transmission through an adaptive selection mechanism. Step S4: When data transmission fails and retransmission is required, a cooperative retransmission protocol is initiated on the multi-hop optimal transmission path set determined in step S2. The relay UAV participating in the retransmission performs linear network coding on the stored data packets and sends incremental redundant packets. At the same time, acknowledgment signals and updated channel measurements during the transmission process are collected and used as online learning data. These acknowledgment signals and updated channel measurements are fed back to the first neural network model in step S1 and the second decision algorithm in step S2, respectively, to update the model weight parameters of the first neural network model and the policy network parameters of the UAV agent in the second decision algorithm.

2. The method for real-time transmission of river and lake inspection images based on UAV swarms according to claim 1, characterized in that: In step S1, the historical channel state information includes received signal strength indication, root mean square delay spread, carrier frequency offset, three-dimensional spatial coordinates of the source UAV node and the three-dimensional spatial coordinates of the destination UAV node; for any communication link, the historical channel state information at the time of acquisition constitutes a five-dimensional feature vector.

3. The method for real-time transmission of river and lake inspection images based on UAV swarms according to claim 2, characterized in that: The process of constructing the input tensor, which includes time and space dimensions, is as follows: The total number of nodes in the UAV swarm is set as a first quantity, and the window length for maintaining historical moment data is set as a second quantity. At each decision moment, a four-dimensional tensor is constructed as the input tensor. The first dimension of this four-dimensional tensor has a length of the second quantity, corresponding to a continuous sequence of historical moments. The second and third dimensions both have a length of the first quantity, corresponding to the UAV node index as the signal transmitter and the UAV node index as the signal receiver, respectively. The fourth dimension has a length of five, corresponding to the five elements of the five-dimensional feature vector. For each element of this four-dimensional tensor, it is determined by the first, second, and third dimension indices. This element is a five-dimensional vector, and its value is the five-dimensional feature vector composed of historical channel state information collected from the transmitting node indicated by the second dimension index to the receiving node indicated by the third dimension index at the historical moment indicated by the first dimension index before the decision moment.

4. The method for real-time transmission of river and lake inspection images based on UAV swarms according to claim 3, characterized in that: The specific process by which the first neural network model processes the input tensor using its internal model weight parameters is as follows: The processing steps of the first neural network model include spatial attention encoding, spatial feature aggregation, temporal convolution prediction, and link gain calculation. The model weight parameters include a trainable weight matrix and a trainable attention parameter vector shared in spatial attention encoding and spatial feature aggregation, as well as a trainable weight vector and bias parameters inside the fully connected layer in the link gain calculation. Spatial attention encoding uses a trainable weight matrix and attention parameter vector to calculate and normalize normalized attention weights that reflect the importance between nodes; Spatial feature aggregation is performed by weighting and summing the features of neighboring nodes based on normalized attention weights to obtain the spatial feature representation of each node. The temporal convolutional network processes the spatial feature representation of each node in the time series and outputs its predicted state feature vector. The link gain calculation concatenates the predicted state feature vectors of any node pair and maps them to the predicted channel gain value of the link between the node pair through a fully connected layer. The predicted channel gain values ​​of all node pairs constitute the predicted channel gain matrix.

5. The method for real-time transmission of river and lake inspection images based on UAV swarms according to claim 4, characterized in that: In step S3, the specific process of fusing the real-time measured signal-to-noise ratio value with the predicted channel gain value of the corresponding link in the predicted channel gain matrix obtained in step S1 is as follows: For each hop transmission link in the multi-hop optimal transmission path set, perform the following operations: First, obtain the real-time measured signal-to-noise ratio value of the current transmission link; At the same time, query the predicted channel gain value corresponding to the current transmission link from the predicted channel gain matrix obtained in step S1; Next, the predicted channel gain value obtained from the query is converted into a predicted signal-to-noise ratio value; Then, an equivalent signal-to-noise ratio (SNR) value for decision-making is calculated. The calculation process is as follows: a weighted sum is performed on the real-time measured SNR value of the current transmission link and the predicted SNR value of the link to obtain a basic fusion value. Based on this, a trend bias value is introduced. The calculation process of this trend bias value includes the following steps: First, the slope of the predicted SNR change is calculated and denoted as the slope value. The slope value is calculated as follows: a time interval is determined, the predicted SNR value at the current moment and the predicted SNR value at a specified future moment are obtained, and the difference between these two predicted SNR values ​​is divided by the time interval. The result is... The first step is to determine the slope value; the second step is to determine the sign function value; if the slope value is positive, the sign function value is positive one; if the slope value is negative, the sign function value is negative one; if the slope value is zero, the sign function value is zero; the third step is to calculate the trend bias magnitude; the trend bias magnitude is the minimum of the absolute value of the slope value, a positive real parameter called the decision look-ahead time, and a positive real parameter called the maximum bias magnitude; the fourth step is to calculate the final trend bias value; the trend bias value is obtained by multiplying the sign function value, a second adjustable parameter called the trend term gain coefficient, and the trend bias magnitude. The base fusion value is added to the trend bias value to obtain the equivalent signal-to-noise ratio value.

6. The method for real-time transmission of river and lake inspection images based on UAV swarms according to claim 5, characterized in that: The specific process for determining the final modulation and coding scheme level for the current data transmission is as follows: A set of modulation and coding scheme levels is preset, and a minimum signal-to-noise ratio (SNR) threshold value is set for each level. After obtaining the equivalent SNR value of each hop transmission link, the final modulation and coding scheme level is determined based on the equivalent SNR value and the set of SNR threshold values. The determination process is as follows: from all preset modulation and coding scheme levels, levels whose corresponding SNR threshold values ​​are less than or equal to the current equivalent SNR value are selected to form a candidate level set. Then, from the candidate level set, the level with the largest value is selected as the final determined modulation and coding scheme level. Through this mechanism, a modulation and coding scheme level that matches the channel quality is determined for each active link in the multi-hop optimal transmission path set.

7. The method for real-time transmission of river and lake inspection images based on UAV swarms according to claim 6, characterized in that: In step S4, the specific process of initiating the cooperative retransmission protocol is as follows: When data transmission fails and retransmission is required, based on the set of optimal multi-hop transmission paths determined in step S2, a set of relay UAV nodes suitable for collaborative retransmission is identified. For each relay UAV node in this set, multiple data packets related to the data block to be retransmitted are selected from its buffer. Each participating relay UAV node independently and randomly generates a number of random coding coefficients equal to the number of selected data packets. Each random coding coefficient is uniformly and randomly selected from a predefined finite field of sufficient size. Then, these random coding coefficients are used to linearly combine the selected data packets. Specifically, the linear combination involves multiplying the first selected data packet by the first random coding coefficient, multiplying the second selected data packet by the second random coding coefficient, and so on, until the last data packet is multiplied by the last random coding coefficient. Then, all the results of the multiplications are processed... The process involves summing the results to generate an incremental redundancy packet. Each relay UAV node encapsulates its generated incremental redundancy packet and all its corresponding random coding coefficients, and then transmits it along the path specified by the multi-hop optimal transmission path set determined in step S2, from the node to the destination node. During transmission, the modulation and coding scheme level determined for each hop of the path in step S3 is used, or a more robust modulation and coding scheme level is selected for retransmission. Upon receiving a new incremental redundancy packet, the receiving end attempts to jointly decode it with the packet and its carried random coding coefficients, along with previously received incremental redundancy packets and coding coefficients, to recover the original data block. If decoding is successful, a positive acknowledgment signal is sent to the source node and related relay nodes. If decoding is unsuccessful after reaching the preset maximum number of retransmissions or time limit, the retransmission is terminated.

8. The method for real-time transmission of river and lake inspection images based on UAV swarms according to claim 7, characterized in that: The specific process of collecting confirmation signals and updated channel measurements during the transmission process, and feeding them back as online learning data to the first neural network model in step S1 and the second decision algorithm in step S2 is as follows: In each data transmission attempt, including the initial transmission and any retransmission, two types of online learning data are collected: the first is acknowledgment signals, which are positive or negative acknowledgment signals that indicate the success or failure of this transmission attempt, and the corresponding transmission time and the link identifier involved are recorded; the second is updated channel measurements, which are a set of real-time channel state information measured and recorded at each data packet transmission or reception time, including at least the following measurement data arranged in order: the real-time measured signal-to-noise ratio, the real-time measured root mean square delay spread, the real-time measured carrier frequency offset, and the three-dimensional spatial coordinates of the transmitting UAV node and the receiving UAV node at the measurement time. For updating the first neural network model in step S1, the collected real-time channel state information arranged in chronological order is combined with the historical channel state information used to construct the input tensor in step S1 to form a training sample set; an online time-series learning algorithm is adopted to improve the accuracy of channel prediction and adjust the model weight parameters of the first neural network model. For updating the second decision algorithm in step S2, each complete transmission attempt is regarded as a decision round, and an experience data tuple is constructed. The experience data tuple includes: the environmental state before the transmission starts, which includes the predicted channel gain matrix output in step S1; the action selected by each UAV agent according to the policy network, i.e., the next hop node; and the instant reward calculated based on the actual transmission result. The calculation basis of the instant reward includes whether the transmission attempt was successful, the end-to-end delay generated by the transmission, and the energy consumption of all participating nodes during the transmission. And the new environmental state after the transmission ends; Multiple such experience data tuples are stored in a distributed experience replay buffer, and a distributed reinforcement learning algorithm is used to calculate the policy gradient using the experience data tuples, thereby updating the policy network parameters of the UAV agent in the second decision algorithm.