Unmanned aerial vehicle image semantic transmission method and system

By applying object detection and semantic segmentation models on the drone to process images and sending the results to edge devices through wireless channels, the problem that traditional communication technology in low-altitude tasks is difficult to efficiently transmit UAV detection results, and efficient communication and processing are achieved.

CN120107834APending Publication Date: 2025-06-06PENG CHENG LAB
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510284810.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In low-altitude tasks, traditional communication technology is difficult to efficiently transmit UAV detection results, resulting in difficult to meet real-time requirements.

Method used

The drone image semantic transmission method is used to process the original image through the object detection model and the semantic segmentation model, generate semantic segmentation diagrams and detection result text, and send them to the edge device through the wireless channel for processing.

Benefits of technology

It significantly improves the communication efficiency and processing speed of low-altitude tasks, reduces the demand for communication bandwidth and computing resources, and ensures efficient transmission of drone detection results and rapid processing of edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107834A_ABST
    Figure CN120107834A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle image semantic transmission method and system, and relates to the technical field of unmanned aerial vehicle communication, and the method comprises the steps: carrying out the collection of a target image under the condition of processing a low-altitude task, and obtaining an original image; inputting the original image into a target detection model to obtain a detection result graph and a detection result text; inputting the detection result graph into a semantic segmentation model to obtain a semantic segmentation graph; and sending the semantic segmentation map and the detection result text to an edge device through a wireless channel, so that the edge device performs task processing according to the semantic segmentation map and the detection result text. According to the invention, the communication efficiency and the processing speed of the low-altitude task are remarkably improved, and meanwhile, the requirements of the system for communication bandwidth and computing resources are reduced, so that the unmanned aerial vehicle detection result can be efficiently transmitted in the low-altitude task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of drone communication technology, and in particular to a method and system for transmitting semantic images of drones. Background Art

[0002] With the rapid development of communication technology and aviation technology, the low-altitude economy has rapidly emerged as a strategic emerging industry and has become an important direction of national economic work. The low-altitude economy covers typical application scenarios such as aerial tours, low-altitude logistics, and urban public governance, showing a development trend of convenience, green, intelligent, efficient, and safe. As one of the core tools of the low-altitude economy, drones are widely used in target detection tasks. They can collect low-altitude information in real time and guide downstream tasks such as traffic supervision, logistics distribution, and real-time navigation. However, the massive detection result images generated by drones in low-altitude tasks need to be efficiently transmitted to edge devices to complete downstream tasks, which puts higher requirements on communication technology.

[0003] At present, image transmission of drone target detection mainly relies on traditional communication technology, that is, directly transmitting the detection result image through a wireless channel. However, when transmitting drone detection results, traditional communication technology has a large amount of data and low transmission efficiency, which makes it difficult to meet the real-time requirements of low-altitude tasks. Secondly, although the semantic communication-based method can reduce the amount of data, the synchronous update of the transceiver model is a challenge. If the transceiver model is not synchronized, the transmitted feature vector will not be parsed, resulting in communication failure. In addition, the existing approach is not accurate enough when reconstructing the image, especially when the target area in the detection result image accounts for a very small proportion, and the target detection information is easily lost. Therefore, how to efficiently transmit drone detection results in low-altitude missions has become an urgent problem to be solved.

[0004] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0005] The purpose of this application is to provide a method and system for transmitting UAV image semantics, aiming to solve the technical problem of how to efficiently transmit UAV detection results in low-altitude missions.

[0006] To achieve the above purpose, the present application proposes a method for semantic transmission of drone images, which is applied to drones and includes:

[0007] In the case of processing low-altitude missions, target image acquisition is performed to obtain the original image;

[0008] Input the original image into the target detection model to obtain a detection result image and a detection result text;

[0009] Inputting the detection result graph into a semantic segmentation model to obtain a semantic segmentation graph;

[0010] The semantic segmentation map and the detection result text are sent to an edge device through a wireless channel, so that the edge device performs task processing according to the semantic segmentation map and the detection result text.

[0011] In one embodiment, the semantic segmentation model includes an encoder module, a decoder module, and an output module;

[0012] The step of inputting the detection result graph into a semantic segmentation model to obtain a semantic segmentation graph comprises:

[0013] Inputting the detection result graph into a semantic segmentation model, performing feature extraction and downsampling on the detection result graph through the encoder module to obtain a high-level semantic feature graph;

[0014] Upsampling and feature fusion of the high-level semantic feature map are performed by the decoder module to obtain a high-resolution feature map;

[0015] The high-resolution feature map is classified pixel by pixel through the output module to obtain a semantic segmentation map.

[0016] In addition, to achieve the above purpose, the present application also proposes a method for semantic transmission of drone images, which is applied to edge devices and includes:

[0017] Receive image data and text data sent by the drone through a wireless channel;

[0018] The text data is input into a detection text semantic conversion module to be restored into a detection result semantic text, so as to perform task processing according to the image data and the detection result semantic text.

[0019] In one embodiment, the module for detecting text semantic conversion includes a decoder and a text semantic conversion module;

[0020] The step of inputting the text data into a detection text semantic conversion module to restore the text data into a detection result semantic text, and performing task processing according to the image data and the detection result semantic text comprises:

[0021] Inputting the text data into the decoder to restore the text into the detection result text;

[0022] The detection result text is input into the text semantic conversion module to obtain the detection result semantic text, so as to perform task processing according to the image data and the detection result semantic text.

[0023] In one embodiment, the step of inputting the detection result text into the text semantic conversion module to obtain the detection result semantic text includes:

[0024] Input the detection result text into the text semantic conversion module, and use the detection result text as a keyword to query the text semantic correspondence table in the text semantic conversion module to obtain a text conversion relationship;

[0025] The detection result text is converted into an interpretable semantic text according to the text conversion relationship to obtain a detection result semantic text.

[0026] In one embodiment, after the step of inputting the text data into the detection text semantic conversion module to restore the detection result semantic text so as to perform task processing according to the image data and the detection result semantic text, the step further includes:

[0027] When the task is image reconstruction, calculating the signal-to-noise ratio of the wireless channel according to the average power of the image data and the text data;

[0028] Inputting the semantic text of the detection result into an initial encoding module to obtain an initial text encoding;

[0029] Inputting the initial text code into an optimized coding module to obtain an optimized text code;

[0030] The signal-to-noise ratio, the image data, the initial text encoding, and the optimized text encoding are input into a diffusion model based on an attention mechanism and text condition constraints to obtain a reconstructed detection result graph.

[0031] In one embodiment, after the step of inputting the signal-to-noise ratio, the image data, the initial text encoding, and the optimized text encoding into a diffusion model based on an attention mechanism and text condition constraints to obtain a reconstructed detection result graph, the step further includes:

[0032] Get the transmission time;

[0033] Calculating communication efficiency according to the reconstructed detection result graph and the transmission time;

[0034] Adjusting the diffusion model according to the communication efficiency, and training the adjusted diffusion model to obtain a target diffusion model;

[0035] The signal-to-noise ratio, the image data, the initial text encoding, and the optimized text encoding are input into the target diffusion model to obtain a target detection result map, and the target detection result map is used to replace the reconstructed detection result map.

[0036] In one embodiment, the step of adjusting the diffusion model according to the communication efficiency and training the adjusted diffusion model to obtain a target diffusion model includes:

[0037] Inputting the communication efficiency into a performance evaluation and feedback module to obtain a performance evaluation result;

[0038] Inputting the performance evaluation result into the diffusion model as feedback information to adjust the hyperparameters of the diffusion model to obtain an adjusted diffusion model;

[0039] The adjusted diffusion model is trained using the drone aerial photography data set, and the number of training times is recorded until the number of training times reaches a preset number of iterations to obtain a target diffusion model.

[0040] In one embodiment, the step of calculating the signal-to-noise ratio of the wireless channel according to the average power of the image data and the text data comprises:

[0041] Calculating average power of transmitting the image data and the text data through the wireless channel;

[0042] Acquire the noise power of the physical channel corresponding to the wireless channel;

[0043] A signal-to-noise ratio of the wireless channel is calculated according to the average power and the noise power.

[0044] In addition, to achieve the above-mentioned objectives, the present application also proposes a drone image semantic transmission system, the system comprising a drone and an edge device, the drone executing the drone image semantic transmission method as described above, and the edge device executing the drone image semantic transmission method as described above.

[0045] In addition, to achieve the above-mentioned purpose, the present application also proposes a drone image semantic transmission device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, and the computer program is configured to implement the steps of the drone image semantic transmission method as described above.

[0046] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the drone image semantic transmission method described above are implemented.

[0047] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the drone image semantic transmission method as described above.

[0048] One or more technical solutions proposed in this application have at least the following technical effects:

[0049] When processing low-altitude tasks, the drone first collects target images and obtains original images containing mission scene information. This step ensures that the drone can capture visual data in low-altitude tasks in real time and provide basic information for subsequent processing. Subsequently, the original image is input into the target detection model to obtain the detection result map and detection result text. The target detection model quickly identifies and locates the target object in the image and extracts key information, thereby streamlining the data volume and retaining valuable target data. Next, the detection result map is input into the semantic segmentation model to generate a semantic segmentation map. The semantic segmentation model further extracts the semantic information of the image, assigns each pixel in the image to a predefined semantic category, and generates a compact semantic segmentation map. This step significantly reduces the data volume while retaining the core semantic content and further optimizing the efficiency of data transmission. Finally, the semantic segmentation map and detection result text are sent to the edge device through the wireless channel. The edge device uses this information to process the task. This process uses wireless communication technology to efficiently transmit the processed data, avoiding the data redundancy and communication delay problems caused by directly transmitting the original image, ensuring that the edge device can quickly receive and process key information and complete downstream tasks. This application achieves full-process optimization from image acquisition to efficient transmission and task processing, significantly improving the communication efficiency and processing speed of low-altitude tasks, while reducing the system's demand for communication bandwidth and computing resources, thereby enabling efficient transmission of drone detection results in low-altitude missions. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0051] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0052] Figure 1 A flowchart diagram of a first embodiment of a method for semantic transmission of drone images applied to a drone is provided in the present application;

[0053] Figure 2 A flow chart of the second embodiment of the method for semantic transmission of drone images applied to edge devices of the present application;

[0054] Figure 3 A schematic diagram of a simplified flow chart of a method for transmitting semantic information of drone images provided in Embodiment 2 of the present application;

[0055] Figure 4A schematic diagram of the framework of the drone image semantic transmission method provided in Example 2 of the present application;

[0056] Figure 5 Schematic diagram of the device structure of the hardware operating environment involved in the drone image semantic transmission method in the embodiment of the present application.

[0057] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0058] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0059] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0060] With the development of communication and aviation technology, low-altitude economy, as an emerging industry, has shown a trend of convenience, greenness and intelligence in the fields of aerial sightseeing, logistics, urban governance, etc. Among them, drones are key tools and are widely used for target detection to support a variety of tasks. At present, drone image transmission mainly adopts the method of direct transmission through traditional wireless channels, but the data volume is large and the transmission efficiency is low, which makes it difficult to meet real-time needs, limiting the application efficiency of drones in the low-altitude economy.

[0061] The main solution of the embodiment of the present application is: first, the drone collects the original image containing the mission scene, and then identifies and locates the target object in the image through the target detection model, and generates the detection result map and text. Next, the image is converted into a compact semantic segmentation map using a semantic segmentation model to reduce the amount of data while retaining the core information. Finally, the semantic segmentation map and the detection result text are sent to the edge device for processing through a wireless channel. This method avoids the data redundancy and communication delay of directly transmitting the original image, ensures the rapid reception and processing of key information, and effectively supports the execution of downstream tasks.

[0062] It should be noted that the execution subject of the embodiment of the present application may be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a drone, an edge device, an image semantic transmission system, etc. that can realize the above functions. The following takes a drone as an example to illustrate this embodiment.

[0063] Based on this, the first embodiment of the present application provides a method for transmitting semantic information of drone images, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the UAV image semantic transmission method of the present application.

[0064] In this embodiment, the UAV image semantic transmission method is applied to a UAV, and the method includes steps S10 to S40:

[0065] Step S10, when processing a low-altitude mission, target image acquisition is performed to obtain an original image.

[0066] It should be noted that low-altitude missions refer to a series of low-altitude flight-related tasks performed by drones in low-altitude economic scenarios, mainly including low-altitude sightseeing, low-altitude logistics, urban public governance and other application scenarios. These tasks require drones to fly in low-altitude environments and complete specific target detection, information collection and data transmission. Target image acquisition refers to the process of drones taking images of targets of interest through their optical cameras or other imaging devices during the execution of low-altitude missions, such as real-time collection of image information of ground targets (such as vehicles, pedestrians, buildings, etc.). These image information is the basic data source for subsequent target detection and semantic communication. The original image refers to the unprocessed image data directly captured by the drone during the target image acquisition process. The image data collected by the drone through the optical camera is the original image.

[0067] It is understandable that, firstly, when the UAV performs a low-altitude mission, it autonomously flies to the target area according to the mission requirements, uses its optical camera to take real-time photos of ground targets, and obtains the original image containing target information. Secondly, by collecting the original image in this way, the real-time and integrity of the image data can be ensured, providing high-quality basic data for subsequent target detection and semantic analysis.

[0068] Step S20, inputting the original image into the target detection model to obtain a detection result image and a detection result text.

[0069] It should be noted that the target detection model refers to an algorithm framework based on deep learning, which is used to identify and locate specific target objects from input images. In this embodiment, the target detection model uses YOLOv5s, which is a lightweight and efficient detection model that can quickly and accurately identify targets in images and output the location information and category information of the targets.

[0070] The detection result graph refers to the image output generated by the target detection model after processing the original image. It is based on the original image and uses the model's detection algorithm to mark the detected target object with a bounding box, and can display the target's category information on the image. The main function of the detection result graph is to intuitively display the detection results of the target detection model, including information such as the location, size, and category of the target.

[0071] The detection result text refers to the textual description of the target information output by the target detection model. It contains the category, location coordinates (such as the center point coordinates, width and height of the bounding box) and other possible attribute information of each detected target. The detection result text is presented in a structured data form (usually a sequence of numbers or characters) to facilitate subsequent processing and transmission.

[0072] Input the original image p into the target detection model yolov5s to implement target detection and output the detection result image x and the detection result text t:

[0073] (x,t)=Det yolo (p)

[0074] Among them, Det yolo (.) indicates that target detection is completed through the yolov5s model.

[0075] It can be understood that, first, the drone takes the collected original image as input data and sends it to a pre-trained target detection model. This model is based on a deep learning algorithm and can quickly identify the target object in the image and locate its position; then, the model outputs a detection result map, which is a visual image with the target position and category marked on the basis of the original image, which is used to intuitively display the detection effect; at the same time, the model will also generate a detection result text, recording key information such as the category and location coordinates of each target in a structured text form for subsequent processing and transmission.

[0076] Step S30, inputting the detection result map into a semantic segmentation model to obtain a semantic segmentation map.

[0077] It should be noted that the semantic segmentation model is a deep learning model that is specifically used to assign each pixel in an image to a predefined category, thereby achieving pixel-by-pixel classification of the image content. In this embodiment, the semantic segmentation model uses U-Net, which is a classic convolutional neural network architecture that is widely used in image segmentation tasks.

[0078] The semantic segmentation map refers to the output image generated by the semantic segmentation model after processing the detection result map. It is an image with the same resolution as the input image, in which the value of each pixel is no longer the original RGB value, but represents the semantic category label to which the pixel belongs (such as "car", "pedestrian", "background", etc.). The semantic segmentation map expresses the semantic information in the image in a compact form, removes redundant pixel-level details, and only retains the outline and category information of the target object.

[0079] Pass the detection result map x through the semantic segmentation model U-net and output the semantic segmentation map x seg for:

[0080] xseg =Seg U-net (x)

[0081] Among them, Seg U-net (.) indicates semantic segmentation is completed through the U-net network.

[0082] It can be understood that, first, the drone uses the detection result map output by the target detection model as input data and sends it to the pre-trained semantic segmentation model; then, the semantic segmentation model analyzes the image pixel by pixel through its network architecture, assigns each pixel in the image to a predefined semantic category, and generates a semantic segmentation map; finally, through this processing, the originally complex detection result map is converted into a semantic and structured semantic segmentation map, which greatly reduces the amount of data and facilitates subsequent efficient transmission and processing, while retaining the core semantic information of the image, providing a basis for improving communication efficiency and image reconstruction in low-altitude missions.

[0083] As an example, the semantic segmentation model includes an encoder module, a decoder module and an output module; the step of inputting the detection result map into the semantic segmentation model to obtain the semantic segmentation map includes: inputting the detection result map into the semantic segmentation model, performing feature extraction and downsampling on the detection result map through the encoder module to obtain a high-level semantic feature map; upsampling and feature fusion on the high-level semantic feature map through the decoder module to obtain a high-resolution feature map; and performing pixel-by-pixel classification on the high-resolution feature map through the output module to obtain a semantic segmentation map.

[0084] The encoder module is part of the semantic segmentation model (such as U-Net), and its main function is to perform feature extraction and downsampling operations on the input detection result map. The encoder gradually extracts the semantic features of the image through a series of convolutional layers and pooling layers, and reduces the spatial resolution of the image. In this process, the encoder can capture important semantic information in the image, such as the shape, texture, and position of the target, while removing redundant detail information, providing a compact feature representation for subsequent feature processing.

[0085] The decoder module is another part of the semantic segmentation model. Its role is to upsample and fuse the high-level semantic feature maps output by the encoder. The decoder gradually restores the spatial resolution of the image through a series of deconvolution layers or upsampling layers, and combines the feature information from the encoder for feature fusion to generate a high-resolution feature map. This process aims to preserve the semantic information while restoring the detailed information of the image, providing a more accurate pixel-level classification basis for the final semantic segmentation.

[0086] The output module is the last part of the semantic segmentation model. Its function is to classify the high-resolution feature map generated by the decoder pixel by pixel. The output module usually contains a convolutional layer to map each pixel in the feature map to a predefined semantic category label. In the final output semantic segmentation map, the value of each pixel represents the semantic category to which it belongs, thereby achieving pixel-by-pixel semantic segmentation of the input image.

[0087] Feature extraction refers to the process of extracting important semantic features from an input image by processing it with a convolutional neural network. In the encoder module, the convolution layer captures local features in the image, such as edges, textures, and shapes, by performing convolution operations on the image. These features can reflect the semantic information of the image and provide a basis for subsequent semantic segmentation.

[0088] Downsampling refers to the operation of reducing the spatial resolution of an image through a pooling layer or a convolutional layer. In the encoder module, downsampling gradually reduces the spatial dimension of the image by reducing the width and height of the image while retaining important semantic features. This process helps remove redundant information, reduces the amount of computation, and enables the model to capture a wider range of contextual information.

[0089] High-level semantic feature maps refer to feature maps generated by the encoder module after feature extraction and downsampling. These feature maps have low spatial resolution but contain rich semantic information, which can reflect the high-level features of the target objects in the image, such as the overall shape of the target, category information, etc. High-level semantic feature maps are the basis for upsampling and feature fusion in the decoder module.

[0090] Upsampling refers to the operation of restoring the spatial resolution of feature maps through deconvolution layers or interpolation methods. In the decoder module, upsampling is used to gradually increase the width and height of the feature map, restore the detailed information of the image, and provide high-resolution feature representation for the final semantic segmentation.

[0091] Feature fusion refers to combining feature maps from different levels to retain more semantic and detail information. In the decoder module, feature fusion is usually implemented through skip connections, which concatenate or add the feature maps in the encoder with the feature maps in the decoder, thereby restoring the image resolution while retaining more contextual information and detail features.

[0092] High-resolution feature maps refer to feature maps generated by the decoder module after upsampling and feature fusion. These feature maps have high spatial resolution while retaining rich semantic information and detail information, which can provide more accurate feature representation for pixel-by-pixel classification in the output module.

[0093] Pixel-by-pixel classification refers to the process in which the output module classifies each pixel in the high-resolution feature map and assigns it to a predefined semantic category. Through convolution operations, the output module maps the features of each pixel to the corresponding category label and finally generates a semantic segmentation map.

[0094] First, the drone inputs the detection result map into the semantic segmentation model. The encoder module of the model extracts features from the image through multi-layer convolution operations, and uses the pooling layer to gradually reduce the spatial resolution of the image, thereby removing redundant details and extracting a high-level semantic feature map containing the target semantic information; then, the decoder module receives the high-level semantic feature map, gradually restores the spatial resolution of the image through deconvolution or upsampling operations, and performs feature fusion in combination with the feature information in the encoder stage to enhance details and contextual information, and finally generates a high-resolution feature map; finally, the output module classifies the high-resolution feature map pixel by pixel, assigns each pixel to a predefined semantic category, and thus generates a semantic segmentation map, realizing the semantic expression and compact representation of the detection result map, and providing a basis for subsequent efficient transmission and processing.

[0095] Step S40: sending the semantic segmentation map and the detection result text to an edge device through a wireless channel, so that the edge device performs task processing according to the semantic segmentation map and the detection result text.

[0096] It should be noted that the wireless channel refers to the wireless communication link between the drone and the edge device for transmitting data. In this embodiment, the wireless channel specifically adopts the additive white Gaussian noise channel (AWGN), which is a common wireless communication channel model used to describe the random noise interference to the signal during the transmission process. The reason for using the wireless channel is that the drone is usually in a mobile state during low-altitude missions, and the communication between it and the edge device needs to cross a certain spatial distance. Wireless communication can provide a flexible connection method to adapt to the dynamic flight needs of the drone.

[0097] Edge devices refer to terminal devices deployed close to data sources (such as drones) with strong computing and storage capabilities. The use of edge devices can avoid transmitting large amounts of data to the cloud or central servers, thereby reducing data transmission delays and bandwidth usage, and improving system response speed and efficiency.

[0098] It can be understood that, first, the UAV uses the semantic segmentation map obtained through semantic segmentation processing and the detection result text output by the target detection model as key information, and sends it to the edge device through a wireless channel (such as an additive white Gaussian noise channel). This process uses wireless communication technology to achieve efficient data transmission, ensuring that the processed information can be quickly transmitted to the edge device in low-altitude missions. Then, after receiving this data, the edge device uses the semantic information and target location information in the semantic segmentation map and the detection result text to perform subsequent task processing. In this way, the collaborative work between the UAV and the edge device can effectively reduce the amount of data transmission and improve communication efficiency, while ensuring that the edge device can quickly respond and complete the task based on the received semantic information, meeting the real-time and efficiency requirements of low-altitude missions.

[0099] This embodiment provides a method for semantic transmission of drone images. When processing low-altitude tasks, the drone first collects target images and obtains original images containing mission scene information. This step ensures that the drone can capture visual data in low-altitude tasks in real time and provide basic information for subsequent processing. Subsequently, the original image is input into the target detection model to obtain the detection result map and the detection result text. The target object in the image is quickly identified and located by the target detection model, and key information is extracted, thereby streamlining the amount of data and retaining valuable target data. Next, the detection result map is input into the semantic segmentation model to generate a semantic segmentation map. The semantic segmentation model further extracts the semantic information of the image, assigns each pixel in the image to a predefined semantic category, and generates a compact semantic segmentation map. This step significantly reduces the amount of data while retaining the core semantic content, further optimizing the efficiency of data transmission. Finally, the semantic segmentation map and the detection result text are sent to the edge device through the wireless channel, and the edge device uses this information to process the task. This process uses wireless communication technology to efficiently transmit the processed data, avoiding the data redundancy and communication delay problems caused by directly transmitting the original image, and ensuring that the edge device can quickly receive and process key information to complete downstream tasks. This embodiment achieves full-process optimization from image acquisition to efficient transmission and task processing, significantly improving the communication efficiency and processing speed of low-altitude tasks, while reducing the system's demand for communication bandwidth and computing resources, thereby enabling efficient transmission of drone detection results in low-altitude missions.

[0100] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above description, and no further description will be given later. On this basis, the following describes this embodiment with the edge device as the execution subject. Figure 2 , Figure 2 This is a flow chart of the second embodiment of the method for semantic transmission of drone images applied to edge devices in the present application.

[0101] In this embodiment, the drone image semantic transmission method is applied to an edge device, and the method includes steps S50 to S60:

[0102] Step S50: receiving image data and text data sent by the drone through a wireless channel.

[0103] It should be noted that image data refers to the semantic segmentation map collected and processed by the drone during the low-altitude mission and received by the edge device. However, due to the influence of the channel gain and channel noise of the wireless channel, the image data received by the edge device may be different from the data sent by the drone.

[0104] Text data refers to the detection result text generated by the drone during the target detection process and received by the edge device. During the transmission process, the text data will also be affected by the channel gain and channel noise of the wireless channel, resulting in the text data received by the edge device may not be completely consistent with the sent data.

[0105] The semantic segmentation map x seg The detection result text t is transmitted to the edge device through the wireless channel, and the received semantic segmentation map With the received test result text for:

[0106]

[0107] Where h(.) is the channel gain, n is the noise, and

[0108] It is understandable that the edge device turns on the receiving mode through the wireless communication module to receive the image data and text data from the drone. Due to the channel gain and noise interference in the wireless channel, the data sent by the drone will change during the transmission process, resulting in the data received by the edge device being inconsistent with the original data at the sender.

[0109] Step S60: input the text data into a detection text semantic conversion module to restore it into a detection result semantic text, so as to perform task processing according to the image data and the detection result semantic text.

[0110] It should be noted that the text semantic conversion module (D sem ) is a key module proposed in this embodiment, and its function is to convert the drone detection result text into semantic text so that it can be used for subsequent task processing, such as image reconstruction.

[0111] The received test result text By detecting the text semantic transformation module D sem Generate semantic text of detection results

[0112]

[0113] Among them, D sem (.) indicates passing through decoder D t The received test result text Restore the detection result text t, and then pass it through the text semantic conversion module TF sem , convert the detection result text t into the detection result semantic text

[0114] The semantic text of the detection result refers to the sem Text data processed by the module. It converts the data information in the drone detection result text (such as target category, coordinate information of the detection box, etc.) into a natural language description. This semantic text not only retains the core information of the detection result, but also enhances the semantic expression ability in the form of natural language, so that it can be used as a conditional constraint for image reconstruction of the diffusion model. For example, the "2, x, y, w, h" in the detection result text is converted into a natural language description of "This is a car (the definition of the target category comes from the setting in the yolo model), the center coordinates are (x, y), the width is w, and the height is h" (x, y, w, h are normalized values). This description method is closer to human understanding of image content, making the text information more semantically interpretable, and can provide more accurate semantic guidance for image reconstruction, thereby improving the quality and accuracy of the reconstructed image.

[0115] It can be understood that, first, the edge device receives text data from the drone, which contains key information for target detection. Then, these text data are input into the detection text semantic conversion module, and the original digital information is converted into text with semantic description through the mapping rules and natural language processing technology inside the module. Finally, the converted detection result semantic text is combined with the received image data for subsequent task processing, thereby improving the accuracy and efficiency of the processing and ensuring the effective use of data in low-altitude tasks.

[0116] As an example, the detection text semantic conversion module includes a decoder and a text semantic conversion module; the step of inputting the text data into the detection text semantic conversion module and restoring it into detection result semantic text so as to perform task processing according to the image data and the detection result semantic text includes: inputting the text data into the decoder and restoring it into detection result text; inputting the detection result text into the text semantic conversion module to obtain detection result semantic text so as to perform task processing according to the image data and the detection result semantic text.

[0117] The decoder is a key component in the detection text semantic conversion module. Its function is to decode the received text data and restore it to the original detection result text. Due to the influence of the channel gain and channel noise of the wireless channel, the text data sent by the drone may be distorted or erroneous during the transmission process, resulting in the data received by the edge device being inconsistent with the data sent. Therefore, the function of the decoder is to process these disturbed text data through traditional wireless communication decoding technology, restore the complete detection result text as much as possible, and provide accurate basic data for subsequent semantic conversion.

[0118] The text semantic conversion module is another core component in the detection text semantic conversion module. Its function is to convert the detection result text restored by the decoder (usually the target category, location coordinates and other information in digital form) into text with semantic description. This module converts digital information into natural language description through preset mapping rules and natural language processing technology. The detection result text refers to the text data restored by the decoder. It contains key information for target detection, such as the target category, location coordinates (such as the center point coordinates, width and height of the bounding box), etc. This information exists in digital form and can directly reflect the output results of the target detection model.

[0119] First, the edge device inputs the received text data affected by channel interference into the decoder, corrects the errors introduced during the transmission process through the decoding algorithm, and restores the accurate detection result text to ensure the integrity of the data; secondly, the restored detection result text is input into the text semantic conversion module, and the module converts the digital text into semantic text described in natural language, making it easier to understand and use for subsequent processing; finally, the semantic segmentation map and the detection result semantic text are combined for task processing, such as image reconstruction or target recognition, to improve processing efficiency and accuracy, and provide more reliable data support for low-altitude missions.

[0120] As an example, the step of inputting the detection result text into the text semantic conversion module to obtain the detection result semantic text includes: inputting the detection result text into the text semantic conversion module, and using the detection result text as a keyword to query the text semantic correspondence table in the text semantic conversion module to obtain a text conversion relationship; converting the detection result text into an interpretable semantic text according to the text conversion relationship to obtain the detection result semantic text.

[0121] Keywords refer to specific information in the detection result text as the basis for query, similar to the "key" in the dictionary, which is used to find matching semantic descriptions in the text semantic correspondence table in the text semantic conversion module to determine the text conversion relationship.

[0122] The text semantic correspondence table is a preset mapping table inside the text semantic conversion module, which stores the mapping relationship between keywords in the detection result text and natural language descriptions. The function of this table is to provide semantic conversion rules and templates for keywords, so that the module can convert the digital detection result text into semantic text described in natural language.

[0123] The text conversion relationship refers to a specific mapping relationship that converts keywords in the detection result text into natural language descriptions according to the preset rules in the text semantic correspondence table.

[0124] Interpretable semantic text refers to natural language description text that can be directly understood and used by humans after being processed by the text semantic conversion module. It converts the numerical information in the detection result text (such as target category and coordinates) into natural language description, making the text information more semantically interpretable.

[0125] First, the detection result text is input into the text semantic conversion module, and the correspondence table therein is used to query with the key information in the detection result text as keywords to find the matching semantic description, thereby determining the text conversion relationship; secondly, according to the queried conversion relationship, the digital information in the detection result text is converted into a semantic text described in natural language, for example, "2, x, y, w, h" is converted into "This is a car with a center coordinate of (x, y), a width of w, and a height of h"; finally, the detection result semantic text is obtained. This interpretable semantic text can be better combined with image data for subsequent task processing, thereby improving the accuracy and efficiency of processing and ensuring the effective use of data in low-altitude missions.

[0126] As an example, after the step of inputting the text data into the detection text semantic conversion module and restoring it to the detection result semantic text so as to perform task processing according to the image data and the detection result semantic text, it also includes: when the task is image reconstruction, calculating the signal-to-noise ratio of the wireless channel according to the average power of the image data and the text data; inputting the detection result semantic text into the initial encoding module to obtain the initial text encoding; inputting the initial text encoding into the optimization encoding module to obtain the optimized text encoding; inputting the signal-to-noise ratio, the image data, the initial text encoding and the optimized text encoding into a diffusion model based on the attention mechanism and text condition constraints to obtain a reconstructed detection result graph.

[0127] Image reconstruction refers to the process of restoring an image similar to the original detection result image through a specific model or algorithm based on the received image data (semantic segmentation map) and text data (semantic text of the detection result) at the receiving end (edge ​​device). In this embodiment, the purpose of image reconstruction is to generate high-quality images using limited transmission data (semantic segmentation map and text information) to meet the requirements for image accuracy in low-altitude missions.

[0128] The signal-to-noise ratio (SNR) refers to the ratio of signal power to noise power, and is usually used to measure the quality of a wireless channel. In this embodiment, the signal-to-noise ratio is obtained by calculating the ratio of the transmission signal power of the image data and text data to the channel noise power. The higher the signal-to-noise ratio, the better the quality of the signal transmission and the smaller the noise interference. The calculation of the signal-to-noise ratio is crucial to the subsequent image reconstruction process because it determines how to allocate resources (such as the ratio of channel coding and source coding) to optimize the quality of the reconstructed image.

[0129] The channel signal-to-noise ratio is given by the noise power σ 2 The noise power σ 2 Mainly caused by natural noise and interference in the physical channel, the channel signal-to-noise ratio γ is:

[0130]

[0131] in, Represents the transmitted semantic segmentation map x seg The average power of the detection result text t. Usually a power normalization operation is used to ensure is equal to a bounded value, e.g.

[0132] The initial encoding module is a processing unit whose function is to convert the semantic text of the detection result into the initial text encoding. The initial text encoding is an intermediate representation used to convert the semantic text described in natural language into a numerical form that can be processed by the model. The role of the initial encoding module is to extract the key information in the semantic text and encode it into the initial text encoding for subsequent optimization processing. The initial text encoding is the result of the output of the initial encoding module. It is a numerical representation used to describe the semantic information of the semantic text of the detection result. The initial text encoding is an intermediate step to convert the natural language text into a format that can be processed by the model. It retains the core information of the semantic text, but has not yet been optimized.

[0133] The detection result semantic text Generate the initial text encoding e through the initial encoding module Target Emb tgt :

[0134]

[0135] Among them, E tgt (.) indicates that the semantic text is passed through the initial encoding module Target Emb to generate the initial text encoding.

[0136] The optimized coding module is a unit that further processes the initial text coding, and its purpose is to improve the quality and efficiency of the text coding through the optimization algorithm. In the present embodiment, the optimized coding module processes the initial text coding through a specific algorithm (such as a neural network or an optimization algorithm) to generate a more efficient optimized text coding, and the optimized text coding can be better combined with the image data for the subsequent image reconstruction process. The optimized text coding is the output result of the optimized coding module, which is the text coding after optimization processing. The optimized text coding improves the efficiency and quality of the coding through the optimization algorithm on the basis of retaining the semantic information of the initial text coding, making it more suitable for the image reconstruction process. The optimized text coding can be better combined with the image data to provide more accurate semantic constraints for the diffusion model.

[0137] Encode the initial text tgt Generate optimized text encoding through the optimized encoding module Optimized Emb opt :

[0138] e opt =E opt (e tgt )

[0139] Among them, E opt (.) indicates that the initial text encoding is generated into an optimized text encoding through the optimized encoding module Optimized Emb.

[0140] The diffusion model based on attention mechanism and text conditional constraints is a deep learning model for image reconstruction. In this embodiment, the model combines the attention mechanism and text conditional constraints to generate high-quality reconstructed images. The role of the attention mechanism is to dynamically adjust the ratio of channel coding and source coding according to the signal-to-noise ratio to optimize the quality of image reconstruction. The text conditional constraints use the optimized text encoding as a semantic guide to ensure that the reconstructed image is semantically consistent with the original image. The model generates the final reconstruction detection result graph by combining image data, signal-to-noise ratio, initial text encoding and optimized text encoding.

[0141] The reconstructed detection result map (also called the final detection result map) refers to the image generated by the diffusion model based on the attention mechanism and text condition constraints. The image is semantically similar to the original detection result map, but may differ in details. The purpose of reconstructing the detection result map is to restore high-quality images using limited semantic information and text information while greatly reducing the amount of transmitted data, so as to meet the requirements for image accuracy in low-altitude missions. The quality of the reconstructed detection result map depends on the performance of the diffusion model, the signal-to-noise ratio, and the quality of the text encoding.

[0142] The signal-to-noise ratio γ and the received semantic segmentation map Initial text encoding tgt , optimize text encoding opt Input into the diffusion model based on text condition constraints using the attention mechanism to generate the final detection result map

[0143]

[0144] Among them, Diff(.) represents the generation of reconstructed images by diffusion model based on text conditional constraints using attention mechanism.

[0145] First, the edge device calculates the signal-to-noise ratio of the wireless channel based on the received image data and text data, and evaluates the channel quality by analyzing the signal-to-noise ratio to ensure that the image reconstruction effect can be optimized under different signal-to-noise ratio conditions; secondly, the semantic text of the detection result is input into the initial coding module, which is converted into the initial text coding, and the semantic information described in natural language is converted into a numerical form that can be processed by the model, providing a semantic basis for image reconstruction; then, the initial text coding is input into the optimized coding module to further improve the quality and efficiency of the coding, making it more suitable for image reconstruction and enhancing the expressiveness of semantic information; finally, the signal-to-noise ratio, image data, initial text coding and optimized text coding are input into the diffusion model based on the attention mechanism and text conditional constraints, and the generation ability of the model is combined with semantic information and channel signal-to-noise ratio SNR to generate a high-quality reconstructed detection result map, thereby achieving high-precision image reconstruction with a reduced amount of transmitted data, meeting the image accuracy requirements of low-altitude missions.

[0146] As an example, after the step of inputting the signal-to-noise ratio, the image data, the initial text encoding, and the optimized text encoding into a diffusion model based on an attention mechanism and text condition constraints to obtain a reconstructed detection result graph, the step further includes: obtaining transmission time; calculating communication efficiency based on the reconstructed detection result graph and the transmission time; adjusting the diffusion model based on the communication efficiency, and training the adjusted diffusion model to obtain a target diffusion model; inputting the signal-to-noise ratio, the image data, the initial text encoding, and the optimized text encoding into the target diffusion model to obtain a target detection result graph, and using the target detection result graph to replace the reconstructed detection result graph.

[0147] Transmission time refers to the time required for the entire process from the drone collecting the original image to the edge device completing the reconstruction of the detection result image. This time covers multiple links such as image acquisition, data processing, wireless transmission, and image reconstruction on the edge device.

[0148] Communication efficiency refers to the degree of match between the amount of task processing that the system can complete per unit time and the actual demand. In this embodiment, communication efficiency is measured by the ratio of the number of reconstructed detection result images to the transmission time, reflecting the efficiency of data transmission and processing in the entire process from image acquisition to image reconstruction when the system is processing low-altitude tasks.

[0149] Calculate the communication efficiency Eff of the new technology solution, that is, the detection result map that can be used for downstream tasks per unit time, and input it into the performance evaluation and feedback module. The communication efficiency Eff of the new technology solution is:

[0150]

[0151] Among them, Thr LPIPS Thr is the threshold value for LPIPS. PSNR is the threshold value of PSNR, T is the transmission time, represents the number of reconstructed final detection result images, Eval is the performance evaluation result, Thr is LPIPS Set the threshold for Eff.

[0152] The target diffusion model refers to a diffusion model that has been adjusted and trained, which can better adapt to the needs of low-altitude tasks and generate higher quality detection result maps.

[0153] The target detection result map refers to the final detection result map generated by the target diffusion model. It is the image output by the optimized and trained diffusion model. Compared with the initial reconstructed detection result map, the target detection result map has significantly improved semantic accuracy and image quality. The target detection result map can more accurately reflect the original image content collected by the drone, while meeting the requirements for image accuracy and real-time performance in low-altitude missions.

[0154] First, the drone records the timestamp of image acquisition when collecting images, and packages this timestamp as metadata together with the image data and text data and sends it to the edge device. In this way, after receiving the data, the edge device can calculate the transmission time from image acquisition to data reception by comparing the current time and the image acquisition timestamp. Secondly, the edge device calculates the communication efficiency, that is, the amount of image reconstruction tasks completed per unit time, based on the calculated transmission time and the number of reconstructed detection result graphs, to evaluate the overall performance of the system. Then, the diffusion model is adjusted according to the communication efficiency, such as modifying parameters to optimize the performance of the model. Next, the adjusted diffusion model is retrained to obtain the target diffusion model, which is more suitable for the high precision and fast response requirements of low-altitude tasks. Finally, the signal-to-noise ratio, image data, initial text encoding, and optimized text encoding are input into the target diffusion model to generate a target detection result graph, and the graph is used to replace the previous reconstructed detection result graph, so as to obtain a higher quality final detection result graph that better meets the task requirements, significantly improving the overall performance and practicality of the system, and ensuring that low-altitude tasks can be completed efficiently and accurately.

[0155] As an example, the steps of adjusting the diffusion model according to the communication efficiency and training the adjusted diffusion model to obtain a target diffusion model include: inputting the communication efficiency into a performance evaluation and feedback module to obtain a performance evaluation result; inputting the performance evaluation result into the diffusion model as feedback information to adjust the hyperparameters of the diffusion model to obtain an adjusted diffusion model; training the adjusted diffusion model using a drone aerial photography data set, and recording the number of training times until the number of training times reaches a preset number of iterations to obtain a target diffusion model.

[0156] The performance evaluation and feedback module is a key component in the system, which is used to evaluate the performance of the diffusion model under the current parameter settings. It generates performance evaluation results by analyzing communication efficiency (such as the amount of image reconstruction tasks completed per unit time) and other possible performance indicators (such as the quality of the reconstructed image, the similarity with the original image, etc.). The core function of this module is to provide quantitative feedback for model adjustment and help the system optimize the parameters of the diffusion model to better meet the needs of low-altitude missions.

[0157] The performance evaluation results refer to the evaluation indicators output by the performance evaluation and feedback module, which are used to measure whether the performance of the current diffusion model meets the expected goals. These results may include whether the communication efficiency meets the task requirements, whether the quality of the reconstructed image is high enough, and the convergence speed of the model, for example, measured by indicators such as PSNR (Peak Signal-to-Noise Ratio) or LPIPS (Learned Perceptual Image Patch Similarity). The performance evaluation results directly determine whether the hyperparameters of the diffusion model need to be adjusted and how to adjust them.

[0158] Hyperparameters are parameters that need to be manually set before model training. These parameters control the learning process and behavior of the model. In the diffusion model, hyperparameters may include learning rate, regularization parameters, weights of loss functions, etc. The setting of hyperparameters has an important impact on the performance and training effect of the model. Through the performance evaluation results provided by the performance evaluation and feedback module, the system can dynamically adjust these hyperparameters to optimize the performance of the model.

[0159] Eval is used as feedback information to input into the diffusion model to adjust the loss function and fine-tune the model. If Eval is 1, no change is made; if Eval is 0, the hyperparameters of the loss function are adjusted:

[0160] L total =λL diff +αL PSNR +βL LPIPS

[0161]

[0162] Among them, L total is the final loss function of the diffusion model, L diff is the loss function of the diffusion model, L PSNR is the loss function of PSNR, L LPIPS is the loss function of LPIPS. λ, α, and β represent the hyperparameters that control the three loss function terms. They represent the hyperparameters of the three adjusted loss function terms.

[0163] The drone aerial photography dataset refers to a collection of labeled data used to train and verify the diffusion model. These data usually include images collected by drones in low-altitude missions and their corresponding semantic information (such as semantic segmentation maps, detection result text, etc.). This dataset reflects the actual scenarios and requirements of low-altitude missions and is used to train the diffusion model to generate high-quality reconstructed images. The dataset may contain images of different scenes, different lighting conditions, and different target types to ensure the generalization ability of the model.

[0164] The preset number of iterations refers to the number of training rounds or training cycles that are preset when training the diffusion model, which is used to control the termination conditions of model training. For example, if the preset number of iterations is 100, the model will be trained for 100 rounds, and each round of training will adjust the model parameters based on the training data to optimize the model's performance. The setting of the preset number of iterations needs to be determined based on the task requirements and the convergence speed of the model. If the number of iterations is too small, the model may not be able to fully learn the features in the data; if the number of iterations is too large, it may cause the training time to be too long or the model to be overfitted.

[0165] First, the edge device inputs the communication efficiency into the performance evaluation and feedback module, and the module generates a performance evaluation result based on the communication efficiency. This result directly reflects whether the performance of the current diffusion model in the actual task meets the expected goal; secondly, the performance evaluation result is input into the diffusion model as feedback information, and the model adjusts its own hyperparameters based on the feedback information, such as modifying the learning rate or the loss function weight, to optimize the performance of the model, thereby obtaining the adjusted diffusion model; finally, the adjusted diffusion model is trained using the drone aerial photography dataset, and the number of training times is recorded. Each training session adjusts the model parameters based on the images and their semantic information in the dataset. When the number of training sessions reaches the preset number of iterations, the training ends and the target diffusion model is obtained. This process ensures that the model can efficiently and accurately complete image reconstruction tasks in low-altitude missions to meet the needs of actual applications.

[0166] As an example, the step of calculating the signal-to-noise ratio of the wireless channel based on the average power of the image data and the text data includes: calculating the average power of the wireless channel transmitting the image data and the text data; obtaining the noise power of the physical channel corresponding to the wireless channel; and calculating the signal-to-noise ratio of the wireless channel based on the average power and the noise power.

[0167] Average power refers to the average energy of a signal when transmitting image data and text data in a wireless channel. It reflects the strength of the signal during transmission. The calculation of average power usually involves the expected value of the square of the signal amplitude, which can characterize the overall energy level of the signal.

[0168] The physical channel refers to the medium for actually transmitting signals in a wireless communication system. The characteristics of the physical channel (such as attenuation, multipath effect, noise, etc.) directly affect the transmission quality and efficiency of the signal.

[0169] Noise power refers to the total energy of all non-signal components in the physical channel, which reflects the interference and noise level in the channel. The calculation of noise power usually involves the statistical characteristics of the noise in the channel, such as the mean and variance of the noise. In actual communication systems, noise power may include thermal noise, shot noise, external interference, etc.

[0170] First, the power of image data and text data transmitted in the wireless channel is measured, and its energy distribution during the transmission process is statistically analyzed to calculate the average power of the signal. This step is to quantify the signal strength for subsequent evaluation of the channel quality. Secondly, the noise power is estimated by measuring the background noise level in the physical channel or using a known noise model. This step is to evaluate the degree of interference in the channel and provide a basis for calculating the signal-to-noise ratio. Finally, the calculated average power is divided by the noise power to obtain the signal-to-noise ratio. The signal-to-noise ratio reflects the relative strength between the signal and the noise and is an important indicator for measuring the quality of the wireless channel. A high signal-to-noise ratio means better signal transmission quality, and more accurate image reconstruction and task processing can be performed, thereby improving the overall performance of the system.

[0171] This embodiment receives image data and text data sent by the drone through a wireless channel. This process ensures the real-time transmission of key information in low-altitude missions and provides basic data for subsequent processing; secondly, the received text data is input into the detection text semantic conversion module, and the module converts the key information in the data text into a detection result semantic text described in natural language. This step makes the text information more semantically interpretable and convenient for combining with image data; finally, task processing is performed based on the semantic segmentation map and the detection result semantic text, such as image reconstruction or target recognition, etc. This processing method combined with semantic information can significantly improve the accuracy of image reconstruction and the efficiency of task processing, ensuring the effective use of data in low-altitude missions and the high performance of the system.

[0172] For example, in order to help understand the implementation process of the drone image semantic transmission method obtained by combining this embodiment with the above-mentioned embodiment 1, please refer to Figure 3 , Figure 3 A brief flowchart of a method for semantic transmission of UAV images is provided. Specifically:

[0173] The figure shows the detailed process of drone image acquisition, processing, transmission and image reconstruction on edge devices. First, the drone collects the original image p through the optical camera (step 1). Then, the image p is input into the target detection model YOLOv5s, and the detection result image x and the detection result text t are output (step 2). Then, the detection result image x is passed through the semantic segmentation model U-net, and the semantic segmentation image x is output. seg (Step 3). Then, the semantic segmentation map x segThe detection result text t is transmitted to the edge device through the AWGN channel (step 4).

[0174] At the edge device, the semantic segmentation map is received With test result text And calculate the channel signal-to-noise ratio γ (step 5). By detecting the text semantic transformation module D sem Generate semantic text of detection results (Step 6). Generate the initial text encoding e through the initial encoding module Target Emb tgt (Step 7), then generate the optimized text code e through the optimized coding module Optimized Emb opt (Step 8).

[0175] The signal-to-noise ratio γ and the received semantic segmentation map Initial text encoding tgt and optimized text encoding opt Input the diffusion model based on the attention mechanism and text condition constraints to generate the final detection result map (Step 9). Calculate the communication efficiency Eff and input it into the performance evaluation and feedback module (Step 10). Input the performance evaluation result Eval as feedback information into the diffusion model to adjust the loss function and fine-tune the model (Step 11). Finally, execute step 9 again and generate the final detection result graph that meets the downstream task requirements based on the fine-tuned diffusion model. (Step 12) This process realizes the full automation of image acquisition, processing, transmission and final image reconstruction, improving the communication efficiency and image reconstruction quality of low-altitude missions.

[0176] Please refer to Figure 4 , Figure 4 The schematic diagram of the framework of the drone image semantic transmission method provided in the second embodiment of the present application is as follows: First, the drone collects the original image p through the optical camera on board. Then, the original image p is input into the target detection model YOLOv5s to generate the detection result text t and the detection result graph x. Secondly, the detection result graph x is input into the semantic segmentation model U-net to generate the semantic segmentation graph x. seg Then, the detection result graph x and the detection result text t are sent to the edge device through the wireless channel.

[0177] The edge device receives the detection result text And semantic segmentation map Through the decoder D t The test result text Restored to the detection result text t, and then passed through the text semantic conversion module TF sem Convert to detection result detection result semantic text

[0178] The initial encoding module Target Emb of the diffusion model will detect the semantic text of the result Convert to the original text encoding tgt Then, the optimized text encoding is generated through the optimized encoding module Optimized Emb opt .

[0179] The semantic segmentation map The calculated signal-to-noise ratio γ, the initial text encoding e tgt and optimized text encoding opt They are input into the diffusion model encoder and diffusion model decoder of the diffusion model diffusion to obtain the final detection result map

[0180] Finally, the performance evaluation and feedback module adjusts the diffusion model according to the performance evaluation result Eval, and obtains the target diffusion model by training the drone aerial photography data set to optimize the image reconstruction quality. The entire process realizes the full process automation from drone image acquisition, semantic information extraction, wireless transmission, semantic text conversion to final image reconstruction.

[0181] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the drone image semantic transmission method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0182] The present application provides a drone image semantic transmission device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the drone image semantic transmission method in the above-mentioned embodiment one.

[0183] Reference below Figure 5, which shows a schematic diagram of the structure of a drone image semantic transmission device suitable for implementing the embodiment of the present application. The drone image semantic transmission device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The drone image semantic transmission device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0184] like Figure 5 As shown, the drone image semantic transmission device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to the program stored in the ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 to the RAM (Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the drone image semantic transmission device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, an LCD (Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the drone image semantic transmission device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a drone image semantic transmission device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.

[0185] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0186] The drone image semantic transmission device provided by the present application adopts the drone image semantic transmission method in the above embodiment, which can solve the technical problem of how to efficiently transmit drone detection results in low-altitude missions. Compared with the prior art, the beneficial effects of the drone image semantic transmission device provided by the present application are the same as the beneficial effects of the drone image semantic transmission method provided by the above embodiment, and the other technical features in the drone image semantic transmission device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.

[0187] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0188] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0189] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the drone image semantic transmission method in the above-mentioned embodiment.

[0190] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory, Erasable Programmable Read Only Memory or Flash Memory), optical fiber, CD-ROM (CD-Read Only Memory, portable compact disk read-only memory), optical storage device, magnetic storage device, or any suitable combination of the above. In the present embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0191] The above-mentioned computer-readable storage medium may be included in the drone image semantic transmission device; or it may exist independently without being assembled into the drone image semantic transmission device.

[0192] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the drone image semantic transmission device, the drone image semantic transmission device: when processing low-altitude tasks, performs target image acquisition to obtain original images; inputs the original images into the target detection model to obtain detection result graphs and detection result texts; inputs the detection result graphs into the semantic segmentation model to obtain semantic segmentation graphs; and sends the semantic segmentation graphs and the detection result texts to the edge device via a wireless channel, so that the edge device performs task processing according to the semantic segmentation graphs and the detection result texts.

[0193] The computer program code for performing the operation of the present application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on the remote computer, or completely on the remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a LAN (Local Area Network) or a WAN (Wide Area Network), or it can be connected to an external computer (e.g., using an Internet service provider to connect through the Internet).

[0194] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0195] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0196] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned drone image semantic transmission method, and can solve the technical problem of how to efficiently transmit drone detection results in low-altitude missions. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the drone image semantic transmission method provided in the above-mentioned embodiment, and will not be repeated here.

[0197] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned drone image semantic transmission method.

[0198] The computer program product provided by this application can solve the technical problem of how to efficiently transmit drone detection results in low-altitude missions. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as the beneficial effects of the drone image semantic transmission method provided by the above embodiment, which will not be repeated here.

[0199] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for semantic transmission of drone images, characterized in that: The method is applied to a drone, and the method comprises: In the case of processing low-altitude missions, target image acquisition is performed to obtain the original image; Input the original image into the target detection model to obtain a detection result image and a detection result text; Inputting the detection result graph into a semantic segmentation model to obtain a semantic segmentation graph; The semantic segmentation map and the detection result text are sent to an edge device through a wireless channel, so that the edge device performs task processing according to the semantic segmentation map and the detection result text.

2. The method according to claim 1, characterized in that The semantic segmentation model includes an encoder module, a decoder module and an output module; The step of inputting the detection result graph into a semantic segmentation model to obtain a semantic segmentation graph comprises: Inputting the detection result graph into a semantic segmentation model, performing feature extraction and downsampling on the detection result graph through the encoder module to obtain a high-level semantic feature graph; Upsampling and feature fusion of the high-level semantic feature map are performed by the decoder module to obtain a high-resolution feature map; The high-resolution feature map is classified pixel by pixel through the output module to obtain a semantic segmentation map.

3. A method for semantic transmission of drone images, characterized in that: The method is applied to an edge device, and the method comprises: Receive image data and text data sent by the drone through a wireless channel; The text data is input into a detection text semantic conversion module to be restored into a detection result semantic text, so as to perform task processing according to the image data and the detection result semantic text.

4. The method according to claim 3, characterized in that The detection text semantic conversion module includes a decoder and a text semantic conversion module; The step of inputting the text data into a detection text semantic conversion module to restore the text data into a detection result semantic text, and performing task processing according to the image data and the detection result semantic text comprises: Inputting the text data into the decoder to restore the text into the detection result text; The detection result text is input into the text semantic conversion module to obtain the detection result semantic text, so as to perform task processing according to the image data and the detection result semantic text.

5. The method according to claim 4, characterized in that The step of inputting the detection result text into the text semantic conversion module to obtain the detection result semantic text comprises: Input the detection result text into the text semantic conversion module, and use the detection result text as a keyword to query the text semantic correspondence table in the text semantic conversion module to obtain a text conversion relationship; The detection result text is converted into an interpretable semantic text according to the text conversion relationship to obtain a detection result semantic text.

6. The method according to claim 3, characterized in that After the step of inputting the text data into the detection text semantic conversion module to restore the detection result semantic text so as to perform task processing according to the image data and the detection result semantic text, the step further includes: When the task is image reconstruction, calculating the signal-to-noise ratio of the wireless channel according to the average power of the image data and the text data; Inputting the semantic text of the detection result into an initial encoding module to obtain an initial text encoding; Inputting the initial text code into an optimized coding module to obtain an optimized text code; The signal-to-noise ratio, the image data, the initial text encoding, and the optimized text encoding are input into a diffusion model based on an attention mechanism and text condition constraints to obtain a reconstructed detection result graph.

7. The method according to claim 6, characterized in that After the step of inputting the signal-to-noise ratio, the image data, the initial text encoding, and the optimized text encoding into a diffusion model based on an attention mechanism and text condition constraints to obtain a reconstructed detection result graph, the method further includes: Get the transmission time; Calculating communication efficiency according to the reconstructed detection result graph and the transmission time; Adjusting the diffusion model according to the communication efficiency, and training the adjusted diffusion model to obtain a target diffusion model; The signal-to-noise ratio, the image data, the initial text encoding, and the optimized text encoding are input into the target diffusion model to obtain a target detection result map, and the target detection result map is used to replace the reconstructed detection result map.

8. The method according to claim 7, characterized in that The step of adjusting the diffusion model according to the communication efficiency and training the adjusted diffusion model to obtain a target diffusion model comprises: Inputting the communication efficiency into a performance evaluation and feedback module to obtain a performance evaluation result; Inputting the performance evaluation result into the diffusion model as feedback information to adjust the hyperparameters of the diffusion model to obtain an adjusted diffusion model; The adjusted diffusion model is trained using the drone aerial photography data set, and the number of training times is recorded until the number of training times reaches a preset number of iterations to obtain a target diffusion model.

9. The method according to any one of claims 6 to 8, characterized in that The step of calculating the signal-to-noise ratio of the wireless channel according to the average power of the image data and the text data comprises: Calculating average power of transmitting the image data and the text data through the wireless channel; Acquire the noise power of the physical channel corresponding to the wireless channel; A signal-to-noise ratio of the wireless channel is calculated according to the average power and the noise power.

10. A UAV image semantic transmission system, characterized in that: The drone image semantic transmission system includes a drone and an edge device. The drone executes the drone image semantic transmission method as described in claim 1 or 2 above, and the edge device executes the drone image semantic transmission method as described in any one of claims 3 to 9 above.

Citation Information

Cited By

  • Monitoring and end-side analysis method based on long-endurance sounding system

    CN120764861A

  • Image transmission method and device, image restoration method and device, electronic equipment and storage medium

    CN120935350A

  • Image processing method, electronic equipment and computer readable storage medium

    CN121099045A