Efficient unmanned aerial vehicle cooperative sensing method based on semantic communication
Through multi-scale feature coding and fine reconstruction modules based on semantic communication, the problems of low data transmission efficiency and insufficient semantic information extraction capabilities in the drone group intelligence perception system are solved, efficient perceived data sharing and reconstruction are realized, and the coordinated perception performance of the system is improved.
Patent Information
- Application Number
- CN202510273447.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-08
AI Technical Summary
The existing drone group intelligence perception systems have limited data transmission efficiency, limited semantic information extraction capabilities, high communication bandwidth requirements and lack of intelligent semantic processing capabilities, resulting in limited system efficiency and reliability.
The efficient collaborative perception method of drone based on semantic communication is adopted, and the semantic features of perceived data are extracted through the multi-scale feature encoding module, and the multi-scale expansion attention mechanism is used to perform feature fusion. Combined with simulated transmission and fine reconstruction modules to transmit and recover data between drones, a joint learning framework for multi-objective optimization is built to balance transmission efficiency and task accuracy.
It improves the data transmission efficiency and semantic information extraction capabilities of the drone group intelligence perception system, ensures high-quality perceived data reconstruction, and realizes the coordinated perception and information sharing of drones in complex environments.
Smart Images

Figure CN120279439A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicles, and particularly to an efficient collaborative perception method for unmanned aerial vehicles based on semantic communication. Background Art
[0002] In related technologies, multi-task learning, as a major research direction in the field of machine learning, aims to improve overall efficiency and performance by concurrently optimizing multiple interrelated tasks. This field has been widely applied in areas such as computer vision, natural language processing, speech recognition, and speech enhancement. A multi-task model typically includes a shared module and multiple task-specific prediction heads, concentrating a large number of parameters in the shared module to achieve more efficient inference and enhance the generalization of feature extraction. However, in actual deployment, directly executing multiple tasks on a unified model often encounters training conflicts between tasks, which can lead to a performance degradation compared to training each task separately.
[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present invention, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The present invention provides an efficient collaborative perception method for unmanned aerial vehicles based on semantic communication, a storage medium, a computer program product, and an unmanned aerial vehicle device, which can effectively improve the efficiency and reliability of the unmanned aerial vehicle cooperation system, and thus can overcome the defects existing in the prior art to a certain extent.
[0005] Other features and advantages of the present invention will become apparent through the following detailed description, or be learned in part through the practice of the present invention.
[0006] According to a first aspect of the present invention, there is provided an efficient collaborative perception method for unmanned aerial vehicles based on semantic communication, the method comprising:
[0007] A target unmanned aerial vehicle collects perception data within a target area;
[0008] The perception data is input into a multi-scale feature encoding module, and a multi-scale expansion attention module is used to extract features from the perception data to obtain multi-scale feature data; wherein, the multi-scale expansion attention module includes a plurality of convolutional layers, and each convolutional layer is configured with a convolutional kernel having a different dilation rate; and
[0009] Feature fusion processing is performed on the multi-scale feature data to obtain multi-scale low-dimensional semantic features;
[0010] Generate a service signal according to the multi-scale low-dimensional semantic features and send it to the neighboring drones; wherein, the target drone and the neighboring drones belong to a set of drones; the neighboring drones include at least one drone.
[0011] In some exemplary embodiments, the performing feature fusion processing on the multi-scale feature data to obtain multi-scale low-dimensional semantic features includes:
[0012] Determine the feature weight data of each feature channel based on a channel attention mechanism;
[0013] Perform feature fusion on the multi-scale feature data by combining the feature weight data through a spatial attention mechanism to obtain multi-scale low-dimensional semantic features.
[0014] In some exemplary embodiments, the method further includes:
[0015] The target drone receives the service signal sent by the neighboring drones;
[0016] Input the service signal into a decoder, and use a refined diffusion model based on a neural network to perform noise estimation on the service signal to obtain a noise prediction result; wherein, the mean and covariance of the noise prediction result follow a Gaussian distribution;
[0017] Perform step-by-step denoising processing on the service signal based on the noise estimation result to obtain the reconstructed perceptual data.
[0018] In some exemplary embodiments, the service signal includes channel noise;
[0019] The method further includes: predefined total number of diffusion steps for step-by-step denoising processing;
[0020] Configure the noise adjustment degree corresponding to each step based on the mean and covariance of the noise prediction result, so as to perform step-by-step denoising on the service signal in combination with the noise prediction result and the noise adjustment degree corresponding to each step.
[0021] In some exemplary embodiments, performing step-by-step denoising on the service signal in combination with the noise prediction result and the noise adjustment degree corresponding to each step, the formula includes:
[0022]
[0023] wherein, α t is a predefined noise parameter; ε θ (x t , t) is the predicted noise; σ t is a parameter for adjusting the noise amplitude; z is a noise vector sampled from the standard normal distribution N(0, I); x t-1 is the perceptual data without noise; is the cumulative noise parameter.
[0024] In some exemplary embodiments, the method further includes: pre-training an efficient drone collaborative perception model based on semantic communication, including:
[0025] Constructing a drone collaborative perception communication system model based on semantic communication and defining a target area for the deployment of a drone set; wherein, the drone set includes multiple drones, each drone is used to perform environmental perception tasks and collect perception data in the target area; the system model includes: a multi-scale feature encoding module for performing multi-scale feature extraction on the collected perception data to obtain multi-scale low-dimensional semantic features; a simulated transmission module for superimposing channel noise on the multi-scale low-dimensional semantic features and transmitting the feature data with superimposed channel noise between drones; a fine reconstruction module for decoding the feature data with superimposed channel noise and gradually denoising to eliminate the channel noise and reconstruct the perception data;
[0026] Initializing the communication model and model parameters; wherein, the model parameters include at least one of a noise level, a compression ratio, and the amount of data of training samples;
[0027] Inputting the training samples into the system model, performing multi-scale feature extraction on the training samples using the multi-scale low-dimensional semantic features; adding noise data to the multi-scale features using the simulated transmission module; denoising and reconstructing the perception data for the multi-scale feature data with added noise data using the fine reconstruction module;
[0028] Calculating the model loss using the reconstructed perception data and the training samples, and updating the model parameters of the system model based on the model loss using backpropagation to obtain the trained system model.
[0029] In some exemplary embodiments, the method further includes:
[0030] Constructing a joint learning framework for multi-objective optimization for the system model, and defining the target tasks of the joint school framework including: a classification task and an image reconstruction task;
[0031] Configuring an integrated loss function for the system model; wherein, the integrated loss function includes: a cross-entropy loss function for the classification task and a mean squared error loss function for the reconstruction of the perception data.
[0032] According to a second aspect of the present invention, there is provided a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the above-mentioned efficient drone collaborative perception method based on semantic communication is implemented.
[0033] According to a third aspect of the present invention, there is provided a computer program product having a computer program stored thereon, and when the computer program is executed by a processor, the above-mentioned efficient drone collaborative perception method based on semantic communication is implemented.
[0034] According to a fourth aspect of the present invention, there is provided a drone device, including a drone body; provided on the drone body are:
[0035] a processor; and
[0036] a memory for storing executable instructions of the processor;
[0037] wherein the processor is configured to implement the above-mentioned efficient drone collaborative perception method based on semantic communication when executing the executable instructions.
[0038] The efficient drone collaborative perception method provided by the embodiments of the present invention can efficiently extract semantic features in the perception data by performing multi-scale feature encoding on the perception data collected by the drone; by introducing a multi-scale dilation fusion attention mechanism to fuse the feature information captured at multiple scales, the diversity of feature expression and the accuracy of information extraction are improved. Furthermore, it is ensured that the receiving end can reconstruct the original perception data with high quality, enabling the drone to more quickly and accurately share and utilize perception information in the crowd-sourced perception task.
[0039] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0041] Figure 1 A schematic diagram schematically showing a method for performing multi-scale encoding on perception data in an exemplary embodiment of the present invention;
[0042] Figure 2 A schematic diagram schematically showing a method for decoding a service signal and reconstructing perception data in an exemplary embodiment of the present invention;
[0043] Figure 3 A schematic diagram schematically showing a schematic architecture of an efficient drone collaborative perception model based on semantic communication in an exemplary embodiment of the present invention;
[0044] Figure 4 Schematically shows a schematic diagram of the composition of an electronic device in an exemplary embodiment of the present invention. Detailed implementation manners
[0045] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this invention will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.
[0046] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0047] In the related art, there are still obvious technical bottlenecks in the existing UAV swarm intelligence perception methods in many aspects. First, the data transmission efficiency is low. Currently, it mainly relies on traditional compression methods. Although it can ensure pixel-level accuracy, it is significantly insufficient in the face of large-scale and real-time data transmission in complex environments. Second, the semantic information extraction ability is limited. Due to the fixed receptive field of traditional convolutional neural networks, it is difficult to effectively capture the features of multi-scale and irregular targets, resulting in the inability to extract the core semantic information related to the task, thereby increasing the redundancy in data transmission. In addition, the communication bandwidth requirement is high, and the semantic information is not fully utilized to optimize the transmission, resulting in a large amount of redundant data occupying limited communication resources, which restricts the efficiency and reliability of the overall system. Finally, in the multi-UAV cooperation scenario, traditional methods mostly regard UAVs as simple data collection tools, lacking task-related intelligent semantic processing and dynamic adaptation capabilities, and failing to fully utilize the collaborative advantages of the multi-UAV system in complex environments. These problems seriously limit the performance improvement of the UAV swarm intelligence perception system.
[0048] Aiming at the shortcomings and deficiencies of the existing technology, an efficient UAV collaborative perception method based on semantic communication is provided in this example embodiment. Refer to Figure 1 As shown, the method includes:
[0049] Step S11, the target UAV collects the perception data in the target area;
[0050] Step S12: Input the sensed data into the multi-scale feature encoding module, and use the multi-scale dilated attention module to extract features from the sensed data to obtain multi-scale feature data. Among them, the multi-scale dilated attention module includes multiple convolutional layers, and each convolutional layer is configured with a convolutional kernel with a different dilation rate; and
[0051] Step S13: Perform feature fusion processing on the multi-scale feature data to obtain multi-scale low-dimensional semantic features;
[0052] Step S14: Generate a service signal according to the multi-scale low-dimensional semantic features and send it to nearby drones. Among them, the target drone and nearby drones belong to a drone set; the nearby drones include at least one drone.
[0053] Next, each step of the efficient drone collaborative sensing method based on semantic communication in this exemplary embodiment will be described in more detail with reference to the accompanying drawings and embodiments.
[0054] In this exemplary embodiment, drone swarm collaborative sensing refers to sharing observation information among multiple drones, improving their own sensing capabilities by fusing local observations and shared observations, and accordingly making collaborative decisions on flight trajectories to achieve collaborative detection and tracking of multiple targets. For drone swarm collaborative sensing, flexible collaborative sensing strategies and efficient information transmission methods are the keys to improving collaborative sensing performance.
[0055] Exemplarily, the method further includes: pre-training an efficient drone collaborative sensing model based on semantic communication, including:
[0056] Step S101: Construct a drone collaborative sensing communication system model based on semantic communication and define the target area where the drone set is deployed. Among them, the drone set includes multiple drones, and each drone is used to perform environmental sensing tasks and collect sensed data in the target area. The system model includes: a multi-scale feature encoding module, which is used to perform multi-scale feature extraction on the collected sensed data to obtain multi-scale low-dimensional semantic features; a simulated transmission module, which is used to superimpose channel noise on the multi-scale low-dimensional semantic features and transmit the feature data with superimposed channel noise between drones; a fine reconstruction module, which is used to decode the feature data with superimposed channel noise and gradually denoise it to eliminate the channel noise and reconstruct the sensed data;
[0057] Step S102: Initialize the communication model and model parameters. Among them, the model parameters include at least one of the noise level, compression rate, and data volume of training samples;
[0058] Step S103: Input the training samples into the system model, extract multi-scale features of the training samples using multi-scale low-dimensional semantic features; add noise data to the multi-scale features using the simulation transmission module; denoise and reconstruct the perceptual data from the multi-scale feature data with added noise data using the fine reconstruction module.
[0059] Step S104: Calculate the model loss using the reconstructed perceptual data and the training samples, and update the model parameters of the system model based on backpropagation using the model loss to obtain the trained system model.
[0060] Specifically, for the application scenario of UAV swarm collaborative perception, a UAV collaborative perception communication system model based on semantic communication can be trained first. Correspondingly, for the system model, the target area can be set as a rectangular area, a circular area, or an irregular polygon area. Several UAVs are deployed in this area and fly at a fixed altitude to perform crowd-sourcing perception tasks. The set of UAVs is denoted as {U1, U2,..., U N}; where each UAV is configured with limited communication bandwidth and battery power, and is equipped with sensors to collect environmental data; for example, the UAV can be equipped with devices such as radar and cameras to sense environmental information. These UAVs fly in the specified area to perform environmental perception tasks. The perceptual data collected by each UAV can be sent to other adjacent UAVs, and at the same time, it can receive the perceptual data sent by adjacent UAVs, and use the received data and the data collected by itself to perform tasks such as target recognition and target tracking.
[0061] For example, a UAV in the UAV swarm can initiate a collaboration request to other UAVs in the swarm. After receiving the collaboration request, the other UAVs can respond to the collaboration request, determine one or more collaborating UAVs corresponding to the current UAV, and establish a communication link for transmitting perceptual data.
[0062] Exemplarily, based on the above application scenario, a system model can be constructed, including: a multi-scale feature encoding module, a simulation transmission module, and a fine reconstruction module.
[0063] Among them, the multi-scale feature encoding module is used to extract multi-scale features of the collected perceptual data to obtain multi-scale low-dimensional semantic features. The simulation transmission module is used to superimpose channel noise on the multi-scale low-dimensional semantic features and transmit the feature data with superimposed channel noise between UAVs. The fine reconstruction module is used to decode the feature data with superimposed channel noise and gradually denoise it to eliminate the channel noise and reconstruct the perceptual data.
[0064] Specifically, for the multi-scale feature encoding module, a multi-scale dilation attention module can be used to extract feature data from the perceptual data to obtain multi-scale feature data; among them, the multi-scale dilation attention module includes multiple convolutional layers, and each convolutional layer is configured with a convolutional kernel with a different dilation rate; then, the multi-scale feature data is subjected to feature fusion processing to obtain multi-scale low-dimensional semantic features.
[0065] By combining convolutional operations and attention mechanisms, the diversity of feature representation and the accuracy of information extraction are improved. By pre-configuring convolutional kernels of multiple scales and different sizes, the model can capture finer-grained features, especially in dynamic environments, and can better adapt to complex and changing perceptual tasks.
[0066] For example, to extract multi-scale features through convolutional kernels with different dilation rates, the corresponding formulas include:
[0067]
[0068] Among them, f2 is the input feature, k is the convolutional kernel size, d is the dilation rate, and i is different scales.
[0069] When performing feature fusion on the features of each scale, the corresponding formulas include:
[0070]
[0071] Exemplarily, performing feature fusion processing on the multi-scale feature data to obtain multi-scale low-dimensional semantic features includes: determining feature weight data for each feature channel based on the channel attention mechanism; performing feature fusion on the multi-scale feature data by combining the feature weight data through the spatial attention mechanism to obtain multi-scale low-dimensional semantic features.
[0072] Specifically, calculating the weight of each feature channel through the channel attention mechanism, the corresponding formulas include:
[0073] f channel = f concat ·σ(FC(GAP(f concat )))
[0074] Performing feature fusion on the multi-scale feature data by combining the feature weight data through the spatial attention mechanism, the corresponding formulas include:
[0075] f spatial = f channel ·Sigmoid(Conv(f channel ))
[0076] Among them, GAP is the global average pooling operation, σ is the Sigmoid activation function, and FC is the fully connected layer for performing fully connected calculations.
[0077] By using a spatial attention mechanism during the encoding process, the model can selectively focus on specific regions or channels, which helps to more effectively retain and enhance key information during the reconstruction of perceptual data.
[0078] Exemplarily, during the model training process, to simulate the communication environment between drones, a simulation transmission module is configured to impose channel noise.
[0079] Specifically, the feature vectors are transmitted between drones through a channel, and the channel is represented by an additive white Gaussian noise (AWGN) model; the noise n follows a Gaussian distribution with a mean of 0 and a variance of The generated noise n is added to the quantized feature vector z quan t to obtain the received signal r.
[0080] The received signal is defined as:
[0081]
[0082] where z quant is the feature vector after quantizing the original feature vector, n represents the noise, and SNR is the signal-to-noise ratio.
[0083] Exemplarily, the fine reconstruction module can be used to decode and denoise the received signal, and recover and reconstruct the perceptual data.
[0084] Specifically, the simulation transmission module can add noise to the encoded data output by the multi-scale feature encoding module and then send it to the fine reconstruction module.
[0085] Specifically, the signal r received by the fine reconstruction module can be the quantized feature vector z quant superimposed with channel noise to obtain. The received signal r is used as the initial state x T of the fine reconstruction module.
[0086] According to the received channel noise variance and the predefined total diffusion steps, the noise adjustment parameter for each step is dynamically calculated. Among them, the parameter σ t for calculating the adjusted noise amplitude is used to control the influence degree of the noise during each denoising process, and the corresponding formula is
[0087]
[0088] Meanwhile, according to the requirement of noise variance matching, the noise adjustment parameter α of the diffusion model is adjusted t , making the cumulative noise variance in the forward process consistent with the variance of the noise added by the analog transmission module, that is, satisfying
[0089] Starting from t = T, the denoising operation is carried out step by step according to the time steps. At each step t, the neural network θ is used to process the current noisy data x t and the time step t. From the perspective of probability distribution, assuming that given the noisy perceptual data x t , the conditional probability distribution of the cleaner data x t-1 at the previous moment follows a Gaussian distribution with mean μ θ (x t , t) and covariance ∑ θ (x t , t). Assume that the conditional probability distribution p θ (x t-1 |x t ) follows a Gaussian distribution, and its mean and covariance are predicted by the neural network. The specific formula is as follows:
[0090]
[0091] where μ θ (x t , t) and ∑ θ (x t , t) are the parameterized mean and covariance matrices respectively, t represents the current time step, x t represents the noisy perceptual data, and θ represents the neural network.
[0092] The neural network θ predicts the mean and covariance based on the current noisy perceptual data x t and the time step t, so as to infer x t-1 closer to the original clean data.
[0093] According to the prediction of the mean and covariance by p θ (x t-1 |x t ) by the neural network, the predicted noise ε θ (x t , t) is obtained. According to the predicted noise, the data x t is gradually updated to obtain the noise-free data x t-1 , and the formula is as follows:
[0094]
[0095] where α t is a predefined noise parameter that controls the amount of noise added in each step, is the cumulative noise parameter, and ε θ (xt , t) is the noise predicted by the neural network, and σ t is a parameter for adjusting the noise amplitude, and z is a noise vector sampled from the standard normal distribution N(0, I).
[0096] Exemplarily, the method further includes: constructing a joint learning framework for multi-objective optimization for the system model, and defining that the joint school framework target tasks include: a classification task and an image reconstruction task;
[0097] Configuring the comprehensive loss function of the system model; wherein, the comprehensive loss function includes: the cross-entropy loss function of the classification task and the mean square error loss function for the reconstruction of the perceptual data.
[0098] Specifically, a joint learning framework for multi-objective optimization is constructed for the noise channel environment. This framework aims at the Pareto optimality of the classification accuracy and the image reconstruction quality, and realizes the balanced optimization of the transmission efficiency and the task accuracy through a dynamic trade-off mechanism. Specifically, the formula for the reconstruction error of the perceptual data may include:
[0099]
[0100] wherein, x and are the original data and the reconstructed data respectively. The task objective is defined as minimizing the reconstruction error to ensure the recovery of high-quality perceptual data.
[0101] Define the classification accuracy The goal is to achieve high classification accuracy through the data recovered at the decoding end to support the object detection or recognition task.
[0102] By defining the comprehensive loss function, balancing the classification accuracy and the reconstruction quality, and combining the dynamic environmental characteristics of the actual UAV swarm intelligence perception task, modular training is carried out to obtain optimized model parameters, so as to realize the efficient semantic information transmission and data reconstruction ability.
[0103] To balance the data reconstruction quality and the classification accuracy, define the comprehensive loss function:
[0104]
[0105] wherein, L CE is the cross-entropy loss of the classification task, which is used to measure the accuracy of the reconstructed data in the classification task; L MSE is the mean square error of the reconstruction error, which is used to evaluate the quality of the reconstructed data.
[0106] wherein,
[0107]
[0108] where y and are the true label and the predicted label respectively; x and are the original data (original sensed data) and the reconstructed data (reconstructed sensed data) respectively; N is the number of samples; λ1 and λ2 are weight coefficients used to control the trade-off between classification accuracy and reconstruction quality.
[0109] Exemplarily, referring to Figure 3 the model architecture shown, the training process of the model may specifically include: First, build key modules for the UAV swarm intelligence sensing communication system model: a multi-scale feature encoder, an analog transmission module, and a fine reconstruction module; and they can be trained independently, and the parameters can be carefully adjusted to optimize the functional performance of each module. Subsequently, the algorithm takes the training dataset and the test dataset as inputs, and at the same time sets the model parameters, the noise level, and initializes the compression ratio. In the training stage, first initialize the UAV semantic communication model and training parameters, including the noise level, the compression ratio, and the batch size. For each data batch in the training set, use the multi-scale encoding module to encode the input data, and then use the analog transmission module to superimpose noise on the encoding according to the set noise level to simulate the real transmission environment, and send it to the fine reconstruction module. In the fine reconstruction module, gradually denoise and recover the data. After decoding, calculate the overall loss based on the comprehensive classification cross-entropy loss and the reconstruction loss (MSE), and use the backpropagation mechanism to update the model parameters. At the same time, the algorithm will accumulate and record the training loss and accuracy of each batch. After the training is completed, the algorithm saves the optimized model and its performance metrics, and verifies the task accuracy and peak signal-to-noise ratio (PSNR) of the model on the test dataset. Finally, output the trained model, the task accuracy, and the PSNR to ensure an optimal balance between transmission efficiency and task accuracy, and achieve efficient semantic information transmission and data reconstruction capabilities. The trained UAV cooperative sensing communication system model based on semantic communication can be deployed on each UAV in the UAV set, so that each UAV can perform multi-scale encoding on the collected sensing data; decode the received service signal to realize the reconstruction of the sensing data.
[0110] In this exemplary embodiment, the process of performing multi-scale encoding on the sensing data collected by the UAV is described. Referring to Figure 1 shown, the method further includes:
[0111] In step S11, the target UAV collects the sensing data in the target area;
[0112] In step S12, the perception data is input into the multi-scale feature encoding module, and the multi-scale dilated attention module is used to extract features from the perception data to obtain multi-scale feature data; wherein, the multi-scale dilated attention module includes a plurality of convolutional layers, and each convolutional layer is configured with a convolutional kernel with a different dilation rate; and
[0113] In step S13, the multi-scale feature data is subjected to feature fusion processing to obtain multi-scale low-dimensional semantic features;
[0114] In step S14, a service signal is generated according to the multi-scale low-dimensional semantic features and sent to nearby drones; wherein, the target drone and the nearby drones belong to a drone set; the nearby drones include at least one drone.
[0115] Exemplarily, the perception data can be image data. The drone set is currently performing a target recognition task in the target area. After the target drone in the set confirms the nearby drones that need to share the perception data currently, it uses the assembled camera component to collect the perception data and uses it as input data, and inputs it into the efficient drone collaborative perception model based on semantic communication. The multi-scale encoding module is used to extract features from the currently collected perception data by using the multi-scale dilated attention module to obtain multi-scale feature data, and by configuring multiple scales and convolutional kernels of different sizes, semantic features of different scales are extracted.
[0116] Exemplarily, in step S13, the processing of performing feature fusion processing on the multi-scale feature data to obtain multi-scale low-dimensional semantic features includes: determining feature weight data for each feature channel based on a channel attention mechanism; and performing feature fusion on the multi-scale feature data by combining the feature weight data through a spatial attention mechanism to obtain multi-scale low-dimensional semantic features.
[0117] Specifically, the weight of each feature channel is calculated through the channel attention mechanism, and feature fusion calculation is performed on the multi-scale feature data based on the weight data to obtain multi-scale low-dimensional semantic features, and the feature data is used as a generated service signal and sent to nearby drones.
[0118] In the embodiment of the present example, the processing process of decoding the received service signal is described.
[0119] Refer to Figure 2 As shown, the method further includes:
[0120] Step S21, the target drone receives the service signal sent by the nearby drone;
[0121] Step S22: Input the service signal into a decoder, and use a refined diffusion model based on a neural network to estimate the noise of the service signal to obtain a noise prediction result; wherein, the mean and covariance of the noise prediction result follow a Gaussian distribution.
[0122] Step S23: Perform step-by-step denoising processing on the service signal based on the noise estimation result to obtain the reconstructed perceptual data.
[0123] Specifically, during the classification process, the target UAV can receive service signals sent by neighboring UAVs. The received service signal data can be input into a decoder, and a trained refined diffusion model based on a neural network is used to estimate the noise of the service signal to obtain a noise prediction result. The noise can include channel noise and transmission loss during the service signal transmission process. Based on the noise estimation result, step-by-step denoising processing is performed on the service signal to achieve fine reconstruction of the perceptual data.
[0124] Exemplarily, the service signal includes channel noise. The method further includes: predefined total diffusion steps for step-by-step denoising processing;
[0125] Configure the noise adjustment degree corresponding to each step based on the mean and covariance of the noise prediction result for step-by-step denoising of the service signal in combination with the noise prediction result and the noise adjustment degree corresponding to each step.
[0126] Exemplarily, for step-by-step denoising of the service signal in combination with the noise prediction result and the noise adjustment degree corresponding to each step, the formula includes:
[0127]
[0128] where α t is a predefined noise parameter; ε θ (x t ,t) is the predicted noise; σ t is a parameter for adjusting the noise amplitude; z is a noise vector sampled from the standard normal distribution N(0, I); x t-1 is the perceptual data without noise; is the cumulative noise parameter.
[0129] Specifically, when configuring the currently executed task for a set of drones, the total number of diffusion steps for progressive denoising can be configured according to the environmental type of the target area, the type of the current task, and the communication environment. For example, if the environment of the current target area has fewer obstacles and less communication interference, a smaller total number of diffusion steps can be configured. Or, if the target area has more obstacles, such as a jungle, forest, or urban environment, with more communication interference and tasks such as tracking and identifying moving people, a larger total number of diffusion steps can be configured; thereby ensuring the processing efficiency of data and the accuracy of the reconstructed data.
[0130] Exemplarily, the noise adjustment parameter σ for each step can be dynamically calculated according to the variance of the noise prediction result and the predefined total number of diffusion steps t , which is used to control the influence degree of noise in each denoising process.
[0131] Specifically, starting from t = T, denoising operations are gradually performed according to the time steps. At each step t, the neural network θ processes the current noisy data x t and the time step t. From the perspective of probability distribution, assuming that given the noisy perception data x t conditions, the conditional probability distribution of the cleaner data x t-1 at the previous moment follows a Gaussian distribution with a mean of μ θ (x t , t) and a covariance of ∑ θ (x t , t). Assume that the conditional probability distribution p θ (x t-1 | x t ) follows a Gaussian distribution, and its mean and covariance are predicted by the neural network. The formula includes:
[0132] p θ (x t-1 | x t ) = N(x t-1 ; μ θ (x t , t), ∑ θ (x t , t))
[0133] Among them, μ θ (x t , t) and ∑ θ (x t , t) are respectively the parameterized mean and covariance matrices. t represents the current time step, x t represents the noisy perception data, and θ represents the neural network. The neural network θ predicts the mean and covariance based on the current noisy data x t and the time step t, thereby inferring x closer to the original clean datat-1 。
[0134] According to the prediction of the mean and covariance by the p θ (x t-1 |x t ) neural network, the predicted noise ε θ (x t , t) is obtained. Based on the predicted noise, the data is gradually denoised to obtain the noise-free data x t-1 , and the formula includes:
[0135]
[0136] where α t is a predefined noise parameter that controls the amount of noise added in each step, is the cumulative noise parameter, ε θ (x t , t) is the noise predicted by the neural network, σ t is a parameter for adjusting the noise amplitude, and z is a noise vector sampled from the standard normal distribution N(0, I).
[0137] The method provided by the embodiments of the present invention can efficiently extract semantic features from the perception data of drones through the designed multi-scale feature encoding module. The introduction of the multi-scale dilated fusion attention mechanism (MDFA) enables the model to capture and fuse feature information from multiple scales, improving the diversity of feature expression and the accuracy of information extraction. At the same time, the fine reconstruction module uses the fine diffusion model to gradually denoise and restore the data, ensuring that the receiving end can reconstruct the original perception data with high quality. This efficient semantic information transmission method enables drones to share and utilize perception information more quickly and accurately in the crowd-sensing task. During the model training process, by adding a simulated transmission module, the actual communication environment is simulated, and the impact of additive white Gaussian noise (AWGN) during the communication between drones is considered. This simulation enables the model to learn how to effectively transmit and reconstruct data in a noisy channel environment during the training process. By constructing a joint learning framework for multi-objective optimization, with the Pareto optimality of the classification accuracy and the image reconstruction quality as the goal, the balance optimization of the transmission efficiency and the task accuracy is achieved through a dynamic trade-off mechanism. This multi-objective optimization ability enables the model to ensure high-quality perception data recovery while pursuing high classification accuracy.
[0138] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for restrictive purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0139] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0140] Furthermore, in the embodiments of this example, a drone device is also provided, including a drone body; an electronic device is provided on the drone body. The electronic device includes: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to implement the above-mentioned efficient drone collaborative perception method based on semantic communication when executing the executable instructions.
[0141] Figure 4 A schematic diagram of an electronic device suitable for implementing the embodiments of the present invention is shown.
[0142] It should be noted that Figure 4 The shown electronic device 1000 is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0143] As Figure 4 shown, the electronic device 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage section 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for system operation are also stored. The CPU 1001, ROM 1002, and RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.
[0144] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as required. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 1010 as required so that a computer program read from thereon is installed into the storage section 1008 as required.
[0145] Specifically, according to an embodiment of the present invention, the processes described below with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product including a computer program carried on a storage medium, the computer program including program codes for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from the removable medium 1011. When the computer program is executed by a central processing unit (CPU) 1001, various functions defined in the system of the present application are executed.
[0146] It should be noted that the storage medium shown in the embodiments of the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any storage medium other than a computer-readable storage medium, and this storage medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the storage medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0148] The units involved in the embodiments of the present invention can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the units themselves in some cases.
[0149] It should be noted that, on the other hand, the present application also provides a storage medium, which can be included in an electronic device; or can exist alone without being assembled into the electronic device. The above storage medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device is caused to implement the methods described in the following embodiments. For example, the described electronic device can implement the steps of the method as shown in Figure 1 、 Figure 2 each step of the method shown.
[0150] In one embodiment, the present application provides a computer program product, including a computer program, which when executed by a processor implements the steps in the above method embodiments.
[0151] In addition, the above drawings are only schematic illustrations of the processes included in the methods according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0152] Those skilled in the art will readily think of other embodiments of the present invention after considering the specification and practicing the invention herein. The present application aims to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include the common general knowledge or conventional technical means in the technical field not disclosed in the present invention. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the claims.
[0153] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. An efficient drone collaborative sensing method based on semantic communication, characterized in that The method includes: The target UAV collects perception data within the target area; Input the perception data into the multi-scale feature encoding module, and use the multi-scale dilated attention module to extract features from the perception data to obtain multi-scale feature data; wherein, the multi-scale dilated attention module includes multiple convolutional layers, and each convolutional layer is configured with a convolutional kernel with a different dilated convolution rate; and Perform feature fusion processing on the multi-scale feature data to obtain multi-scale low-dimensional semantic features; Generate a service signal according to the multi-scale low-dimensional semantic features and send it to the neighboring UAVs; wherein, the target UAV and the neighboring UAVs belong to the UAV set; the neighboring UAVs include at least one UAV.
2. The method according to claim 1, characterized in that, The performing feature fusion processing on the multi-scale feature data to obtain multi-scale low-dimensional semantic features includes: Determine the feature weight data of each feature channel based on the channel attention mechanism; Perform feature fusion on the multi-scale feature data by combining the feature weight data through the spatial attention mechanism to obtain multi-scale low-dimensional semantic features.
3. The method according to claim 1, wherein The method further includes: The target UAV receives the service signal sent by the neighboring UAV; Input the service signal into the decoder, and use the fine diffusion model based on the neural network to perform noise estimation on the service signal to obtain a noise prediction result; wherein, the mean and covariance of the noise prediction result follow a Gaussian distribution; Perform step-by-step denoising processing on the service signal based on the noise estimation result to obtain the reconstructed perception data.
4. The method according to claim 3, wherein The service signal includes channel noise; The method further includes: predefined total number of diffusion steps for step-by-step denoising processing; Configure the noise adjustment degree corresponding to each step based on the mean and covariance of the noise prediction result, so as to perform step-by-step denoising on the service signal by combining the noise prediction result and the noise adjustment degree corresponding to each step.
5. The method according to claim 4, wherein The formula for performing step-by-step denoising on the service signal by combining the noise prediction result and the noise adjustment degree corresponding to each step includes: Among them, α t is a predefined noise parameter; ε θ (x t , t) is the predicted noise; σ t is a parameter for adjusting the noise amplitude; z is a noise vector sampled from the standard normal distribution N(0, I); x t-1 is the sensed data without noise; is the cumulative noise parameter.
6. The method according to claim 1, wherein The method further includes: pre-training an efficient UAV collaborative perception model based on semantic communication, including: Construct a UAV collaborative perception communication system model based on semantic communication, and define the target area where the UAV set is deployed; wherein, the UAV set includes multiple UAVs, and each UAV is used to perform environmental perception tasks and collect perception data in the target area; the system model includes: a multi-scale feature encoding module, which is used to perform multi-scale feature extraction on the collected perception data to obtain multi-scale low-dimensional semantic features; a simulation transmission module, which is used to superimpose channel noise on the multi-scale low-dimensional semantic features and transmit the feature data with superimposed channel noise between UAVs; a fine reconstruction module, which is used to decode and perform step-by-step denoising on the feature data with superimposed channel noise to eliminate channel noise and reconstruct perception data; Initialize the communication model and model parameters; wherein, the model parameters include at least one of the noise level, compression rate, and data volume of training samples; Input the training samples into the system model, and use multi-scale low-dimensional semantic features to perform multi-scale feature extraction on the training samples; use the simulation transmission module to add noise data to the multi-scale features; use the fine reconstruction module to denoise and reconstruct the perceptual data for the multi-scale feature data with added noise data; Calculate the model loss using the reconstructed perceptual data and the training samples, and update the model parameters of the system model based on the model loss using backpropagation to obtain the trained system model.
7. The method according to claim 6, wherein The method further includes: Construct a joint learning framework for multi-objective optimization for the system model, and define that the joint school framework target tasks include: classification tasks, image reconstruction tasks; Configure the comprehensive loss function of the system model; wherein, the comprehensive loss function includes: the cross-entropy loss function for the classification task, and the mean square error loss function for the reconstruction of the perceptual data.
8. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the efficient UAV cooperative perception method based on semantic communication according to any one of claims 1 to 7.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the efficient UAV cooperative perception method based on semantic communication according to any one of claims 1 to 7.
10. A drone device, characterized in that, It includes a UAV body; on the UAV body, there are provided: A processor; And A memory for storing the executable instructions of the processor; Wherein, the processor is configured to execute the efficient UAV cooperative perception method based on semantic communication according to any one of claims 1 to 7 by executing the executable instructions.