Unmanned aerial vehicle image semantic transmission and question-answering system and method driven by large language model

Through the drone image semantic transmission and question-and-answer system driven by a large language model, combined with deep convolutional neural network and codebook mechanism, the problem of image information conversion decision content in drone rescue scenarios is solved, efficient image semantic communication and decision support is achieved, and the system interpretability and communication efficiency are improved.

CN120087484APending Publication Date: 2025-06-03NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510011196.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to effectively convert image information into decision-making content in drone rescue scenarios, and lacks deep fusion of large language models, resulting in insufficient interpretability and communication efficiency of the system in complex scenarios.

Method used

The image semantic transmission and question-and-answer system driven by a large language model is adopted, combined with deep convolutional neural network and codebook mechanism, and image semantic communication and natural language response are realized through modules such as image semantic extraction and quantization, image reconstruction, subtitle generation and large language model Q&A.

Benefits of technology

In complex and resource-constrained drone rescue scenarios, efficient image semantic communication and decision support are achieved, improving system interpretability and communication efficiency, and showing superior performance in low signal-to-noise ratio conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087484A_ABST
    Figure CN120087484A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle image semantic transmission and question-answering system and method driven by a large language model, and an image semantic extraction and quantification module carries out the semantic feature extraction of an original image captured by an unmanned aerial vehicle through a semantic encoder, and discretizes a feature vector through a codebook mechanism, and is used for wireless channel transmission. And the image reconstruction module recovers the transmission image from the semantic level according to the received discrete vector corresponding to the codebook shared by the receiving end and the transmitting end. And the image subtitle generation module takes the image recovered by the image reconstruction module as input, and generates text description information by using a BLIP subtitle generation system. And the cue word generation module generates a section of cue word according to the user task description and the text description information output by the image subtitle generation module. And the large language model question-answering module outputs natural language response to guide the next action of the unmanned aerial vehicle according to the cue word. According to the method, the robustness and efficiency of wireless image transmission under the condition of low signal-to-noise ratio are improved, the interpretability of a question answering system is improved, and the method has a certain practical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a large language model-driven drone image semantic transmission and question-answering system and method, belonging to the technical field of image communication. Background Art

[0002] With the advent of the era of all things intelligent interconnection, high-quality spectrum resources are becoming increasingly scarce, and the hardware cost and energy consumption are getting higher and higher. This poses a severe challenge to the existing communication systems to achieve low latency and high data transmission rate under limited spectrum resources. To solve these problems, semantic communication, as a new paradigm for the next-generation communication system, goes beyond bit-level accuracy and preferentially transmits meaningful information, triggering extensive exploration in the academic and industrial circles.

[0003] The visual modality occupies more data volume than the text modality. The huge image / video traffic often faces problems such as delay and network congestion during the transmission process, which are difficult to solve by traditional communication technologies. Therefore, it is particularly important to construct image semantic communication that can significantly reduce the data transmission volume without sacrificing the semantic fidelity of the image. At present, the solutions for image semantic communication are becoming increasingly rich. According to the core technical means adopted, the image semantic communication paradigm can be divided into three categories: deep learning-based semantic communication, generative adversarial network-based semantic communication, and semantic knowledge base-based semantic communication. Among them, the deep learning-based semantic communication system uses a deep neural network to perform semantic encoding and decoding on images, thereby reducing the amount of transmitted data. For example, in the paper "Deep Joint Source-Channel Coding for Wireless Image Transmission" published by E. Bourtsoulatze, D. B. Kurka, D. Gündüz et al., a source-channel joint coding technology based on a convolutional neural network is proposed for wireless image transmission. This technology directly maps the pixel values of the input image to complex-valued channel input symbols, omitting the traditional explicit compression and error correction coding steps. Through the end-to-end training mode, the encoder at the sending end and the decoder at the receiving end are jointly optimized to minimize the mean square error (MSE) between the reconstructed image and the original image. This paradigm has better image transmission performance than the traditional separate digital transmission scheme under low signal-to-noise ratio and limited channel bandwidth conditions. However, the optimization direction of the pixel-level loss function limits its ability in semantic feature extraction, making it difficult to balance high fidelity and high compression rate. Moreover, the end-to-end communication system is often regarded as a black box, and its performance is difficult to be explained or proven theoretically.

[0004] The semantic communication system based on generative adversarial networks uses generative adversarial networks to generate high-quality images or perform image reconstruction, further enhancing the generation and interpretability of image semantic information. In the article "Toward Semantic Communications: Deep Learning-Based Image Semantic Coding" published by D. Huang, F. Gao, X. Tao, Q. Du, etc., an image semantic coding method based on deep learning was proposed, and generative adversarial networks were combined to optimize image semantic communication. This scheme defines "semantic concepts" as image representation units, including categories, spatial distributions, and visual features, and extracts the above information through a convolutional semantic encoder. Secondly, a rate-semantic-perception loss function was designed, comprehensively considering the transmission bit rate, semantic fidelity, and perceptual quality to optimize the image coding process, and a reinforcement learning model was used to achieve semantic bit allocation, dynamically adjusting the quantization accuracy according to task requirements, thereby improving the transmission efficiency of key semantic content. At the receiving end, this paradigm uses a generative adversarial network combined with an attention mechanism for decoding operations, generating high-quality, semantically consistent images through local and global feature fusion, and performing well in low-bitrate and high-noise channels, providing an innovative solution for the field of semantic communication. However, this semantic communication system largely depends on the performance of generative adversarial networks and is particularly sensitive to the processing of refined semantic information, especially performing poorly on small targets or rare categories in complex scenarios.

[0005] Semantic communication systems based on semantic knowledge bases utilize the prior knowledge provided by the semantic knowledge bases to assist in the semantic encoding and decoding processes of images, improving transmission efficiency and accuracy. With the rapid development of artificial intelligence technology, large artificial intelligence models have gradually become the focus of research. Especially for semantic knowledge bases, aiming at problems such as limited knowledge representation, frequent knowledge updates, and insecure knowledge sharing in the construction of semantic knowledge bases, in the article "Large AI model-based semantic communications" published by F. Jiang, Y. Peng, L. Dong, K. Wang, K. Yang, C. Pan, etc., a large-scale semantic communication framework based on artificial intelligence models specifically for image data was proposed. This framework constructs a knowledge base using a segmented arbitrary model and divides the original image into different semantic segments through general semantic knowledge. It uses a semantic integration technology based on the attention mechanism to weight the semantic segments and integrate them into a semantic-aware image, and uses adaptive semantic compression coding to remove redundant information in the semantic features, achieving a reduction in communication overhead. The above three types of image semantic communications design semantic communication frameworks from the perspectives of reliability and effectiveness, but still face significant limitations, especially in dealing with complex and high-dimensional data representations. Therefore, in complex search and rescue scenarios, further exploration is still needed to achieve efficient image transmission under constrained bandwidth and noise interference.

[0006] In addition, with the rapid development of artificial intelligence technology (AI), large artificial intelligence models have gradually become the focus of research. Models such as GPT and BERT are pre-trained and fine-tuned based on large datasets, showing special capabilities in the fields of natural language understanding and generation. These models can not only accurately understand the needs of users but also generate smooth and logical answers, thus achieving great success in tasks such as natural language processing including machine translation, text summarization, and question answering systems. Flamingo proposed in the article "Flamingo: a visual language model for few-shot learning" published by J.-B. Alayrac, J. Donahue, P. Luc, etc. is a visual language model that connects powerful pre-trained visual models and language models through innovative designs and shows significant advantages in few-shot learning. It can quickly adapt to new tasks with the fewest samples and reach a new technical level in visual question answering, image captioning, and multiple-choice visual question answering.

[0007] Although the progress of advanced large-scale AI models and vision-language integration has provided powerful tools for communication, they are generally designed for general applications rather than specific task requirements. In the context of drone-based search and rescue scenarios, existing methods face challenges in directly converting image information into decision-making content, and the inference process based solely on images often lacks interpretability. In summary, the above-mentioned studies have failed to fully explore the potential application areas of image semantic communication, such as its specific applications in drone rescue scenarios, and also lack strategies for deep integration with large language models. Summary of the Invention

[0008] The object of the present invention is to address the defects or deficiencies of the above-mentioned existing technologies, and propose a large language model-driven drone image semantic transmission and question-answering system and method. The present invention combines the advantages of advanced technologies such as deep convolutional neural networks, codebook mechanisms, and large language models, explores the application scenarios of drone rescue in the image semantic communication framework, solves the problem of the lack of strategies for deep integration with large language models in existing research, and improves the interpretability of the system.

[0009] The technical solution adopted by the present invention to solve its technical problems is: a large language model-driven drone image semantic transmission and question-answering system, which includes an image semantic extraction and quantization module, an image reconstruction module, a caption generation module, a prompt generation module, and a large language model question-answering module. The image semantic extraction and quantization module uses a semantic encoder to extract feature vectors from the original images captured by the drone, and uses a codebook shared by the transmitter and receiver to map the latent feature vectors to discrete index sequences, which can significantly reduce the amount of data transmitted. This module is the basis for the subsequent modules. The image reconstruction module takes the index sequence as input, inverse maps the latent feature vectors according to the shared codebook, and then restores the original image through a semantic decoder. Semantic decoding is the inverse process of semantic encoding, and the original image can be restored based on the latent feature vectors. The caption generation module takes the reconstructed image output by the image reconstruction module as input and realizes the generation of text captions based on semantic images. The prompt generation module forms a prompt for question-answering according to a fixed question template, combines the text captions and user question information, and serves as the input for the subsequent large language model, which can improve the interpretability of the system. The large language model question-answering module is the output module of the system, and uses an advanced large language model to achieve high-quality natural language responses.

[0010] Furthermore, the system of the present invention is designed for complex drone rescue scenarios in the real world. By combining semantic communication and large language models, a framework for a large language model-driven drone image semantic transmission and question-answering system is constructed. For the complex and resource-constrained environment in drone search and rescue scenarios, the invention can achieve better transmission effects with fewer transmission bits at low signal-to-noise ratios by introducing a codebook mechanism into the semantic communication framework, while maintaining comparable results to other transmission schemes at high signal-to-noise ratios. And by designing a question-answering system that integrates image caption generation and language models, complex image information is converted into natural language descriptions, making it easier for users to understand the system's decisions and outputs, thereby improving the interpretability of the system.

[0011] The present invention also provides a method for implementing a large language model-driven drone image semantic transmission and question-answering system, which includes the following steps:

[0012] Step 1, dataset preprocessing and deployment of relevant large language models. Download the Open Images v6 dataset and select 2,000 training samples and 500 validation sample images. Retrieve 20 relevant pictures of rescue scenarios through a browser for drone search trajectory simulation, and locally deploy the BLIP and LLAMA large language models.

[0013] Step 2, extraction and quantization of latent feature vectors of drone images. Use a semantic encoder to extract the latent feature vectors of the original images, find similar vectors in the codebook through the nearest neighbor algorithm, replace the feature vectors with them, and generate a corresponding index sequence.

[0014] Step 3, image reconstruction based on feature vectors. Retrieve the corresponding feature vectors from the codebook according to the index, and use a semantic decoder to reconstruct the original image.

[0015] Step 4, image caption generation. Use the BLIP framework to automatically generate semantic descriptions of the reconstructed images.

[0016] Step 5, prompt generation and large language model question-answering. Generate prompts according to the structured prompt template, combining the image captions and user questions, and input the prompts into the large language model to obtain natural language responses for question-answering.

[0017] Beneficial effects:

[0018] 1. The present invention designs a drone image semantic transmission and question-answering system combined with a large language model, explores the actual application scenarios of drone rescue, and realizes efficient communication and guidance in complex and resource-constrained situations.

[0019] 2. The present invention extracts and quantifies the semantic information of images through a codebook mechanism, generates high-quality natural language responses while transmitting fewer bits, thereby significantly improving communication efficiency and showing superior performance under low signal-to-noise ratio conditions.

[0020] 3. The present invention utilizes a question-and-answer system that combines image caption generation with a large language model to convert complex image semantic information into natural language descriptions, enhancing the interpretability of the system and enabling users to more intuitively understand the system's decisions and outputs. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a system block diagram of the present invention.

[0022] Figure 2 It is a comparison chart of the performance of the present invention and other wireless image transmission schemes in learning the perceptual image block similarity of image restoration metrics under different signal-to-noise ratios.

[0023] Figure 3 It is a comparison chart of the question-and-answer accuracy performance of the present invention and other wireless image transmission schemes under different signal-to-noise ratios.

[0024] Figure 4 It is a comparison chart of the question-and-answer visualization effects of the present invention and other wireless image transmission schemes under different signal-to-noise ratios.

[0025] Figure 5 It is a simulation trajectory chart of a drone search and rescue scenario designed based on the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the following provides a more detailed description of the present invention in conjunction with the drawings in the specification.

[0027] It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0028] As Figure 1 shown, the present invention provides an implementation method for a large language model-driven drone image semantic transmission and question-and-answer system, and the method includes the following steps:

[0029] Step 1, preprocessing of the dataset and deployment of relevant large language models.

[0030] The publicly available dataset used in this invention is the Open Images v6 dataset, which contains millions of annotated images with high annotation accuracy, covering a variety of scenarios and object types. The original dataset contains approximately 1.7 million images, and the validation set contains 41,000 images. For the subsequent Q&A tasks, this invention selects 2,000 sample images and 500 sample images for training and validation respectively, retrieves 20 relevant pictures of rescue scenarios through the browser for the simulation of the drone search trajectory, and uniformly adjusts the image size in the dataset to 256×256. In addition, this invention deploys the BLIP model for image caption generation and the LLAMA model for Q&A tasks on the cloud server. Among them, BLIP is an advanced vision-language pre-training model, which realizes image feature extraction and caption generation tasks through its vision encoder and language decoder. LLAMA is a large language model developed by Meta, which is used in this invention to understand the input prompt words and generate corresponding natural language answers.

[0031] Step 2, extraction and quantization of the latent feature vectors of the drone images.

[0032] After completing the dataset preprocessing and the deployment of relevant large language models in the above Step 1, assume the input image is the source image captured by the drone where H and W represent the height and width of the image respectively, and "3" represents the RGB channels. This invention first needs to achieve the semantic transmission of this image. Specifically, this invention uses a semantic encoder designed based on a convolutional neural network (CNN) to perform semantic extraction on the original image x captured by the drone, and obtains the latent vector representation

[0033] This process can be expressed as:

[0034] z e =E α (x)

[0035] where E α (·) is the semantic encoding operation with parameter α, and z e is the latent feature vector. This semantic encoder consists of a 5×5 initial convolution, four downsampling modules and an output layer. The downsampling module contains a basic residual block designed based on ResNet and an optional attention block. Finally, the features are mapped to the latent vector representation through normalization and convolutional layers.

[0036] Then, this invention performs vector quantization on the extracted latent vectors through the codebook mechanism. Specifically, a discrete codebook is pre-trained At each encoding position of the feature vector output by the semantic encoder, the nearest vector is found in the codebook Z through the nearest neighbor algorithm and replaced with the quantized latent representation zq , this process can be expressed as:

[0037]

[0038] where i and j represent vectors At each position in the latent feature vector z e , q(σ) represents an element-wise quantization operation, and z q is the quantized latent representation, i.e., a sequence of feature indices.

[0039] Step 3, image reconstruction based on the feature vector.

[0040] The present invention transmits the latent representation obtained in the above step 2, i.e., z q through a wireless channel, and the channel is modeled as an additive white Gaussian noise (AWGN) channel. The transmission process can be expressed as:

[0041]

[0042] where is the received signal, represents independent and identically distributed (IID) channel noise samples, which follow a symmetric complex Gaussian distribution n ∼ CN(0, σ 2 I), and h is the channel gain. At the receiving end, the present invention retrieves the corresponding feature map from the codebook according to the received feature indices to obtain the feature vector z d , and inputs it into the semantic decoder for image reconstruction. This process can be expressed as:

[0043] x′ = D β (z d )

[0044] where D β (·) is a semantic decoding operation with parameter β, and x′ is the reconstructed image. The semantic decoder adopted by the present invention is a symmetric architecture, including a 5×5 head convolution, 4 upsampling modules, and a recoding module. The upsampling module uses residual blocks similar to the encoder and performs image upsampling using the nearest neighbor interpolation method. The recoding module consists of a 3×3 convolution and a Swish activation function. Then, end-to-end training is performed, and the loss function for the entire process is as follows:

[0045] L 1 (E, D, Z) = L rec + γL com + L reg

[0046] L rec = ||x - x′|| 2

[0047]

[0048] Among them, sg(·) represents the stop gradient operator, and L rec is the reconstruction loss, which is used to measure the difference between the reconstructed image and the original image; L com is the commitment loss, which measures the difference between the quantized latent vector and the output of the semantic encoder; L reg is used for regularizing the encoder output. To further improve the quality of the reconstructed image, the present invention introduces adversarial training and perceptual loss. The perceptual loss uses a pre-trained neural network to compare the high-level features of the original image and the reconstructed image, so that the generated image has a feature representation similar to that of the real image. The perceptual loss function is defined as follows:

[0049] L per = ||φ(x) - φ(x′)|| 2

[0050] where L per is the perceptual loss, φ(·) represents the features extracted by the pre-trained network, and x and x′ are the original image and the reconstructed image. The adversarial training objective is that the image generated by the semantic decoder is closer to the real image. The loss of the generative adversarial network (GAN) is expressed as:

[0051]

[0052] where represents the probability that the discriminator judges the real image x as true; represents the probability that the discriminator judges the reconstructed image x′ as true. Then, the overall loss function is as follows:

[0053] L total = L rec + γL com + μL 2 + θL per + L reg

[0054] where γ, μ, and θ are the balance factors of different loss functions. The present invention performs self-supervised training on the semantic encoder, decoder, codebook, and discriminator by optimizing this loss function.

[0055] Step 4, image caption generation.

[0056] According to Step 3 above, the reconstructed image, that is, x′, can be obtained. The present invention uses the BLIP model deployed in Step 1 to process the reconstructed image. First, the image is converted into visual features through its visual encoder, and then the visual features are input into the language decoder to generate the corresponding text description. The caption generation process can be expressed as:

[0057] z v = Ev (x')

[0058] T subtitle = D l (E v (x'))

[0059] where z v is the visual feature vector, E v (·) represents the visual encoder, D l (·) represents the language encoder, and T subtitle represents the generated text description. The generated text description then passes through the filtering system of the BLIP model, which evaluates the consistency between the input image features and the generated text description, ensures that the output text description is highly accurate and relevant, and finally generates a caption T containing the key details of the image (such as objects or people in the scene), providing an effective input for the question-and-answer tasks of the subsequent large language model.

[0060] Step 5, prompt generation and large language model question and answer.

[0061] According to the caption T generated by the image captioning system in the above step 4, the present invention combines the task description information provided by the user, forms a prompt according to the structured prompt template, and then inputs the generated prompt into the LLAMA large language model. Based on its powerful natural language understanding and generation capabilities, the model understands and infers the prompt, generates an effective answer or decision to complete the image-based question-and-answer task, and provides guidance for the next action of the drone.

[0062] The effects of the present invention will be further described below in conjunction with simulation experiments. Specifically, it includes:

[0063] 1) Simulation conditions and parameter settings

[0064] The simulation experiment of the present invention is carried out on a simulation platform of Python 3.8 and PyTorch 2.0.0. The computer CPU model is Intel Xeon Platinum 8474C, and the GPU model is NVIDIA GeForce RTX 4090. The dataset used in the present invention is Open Images v6, the learning rate is set to 0.0000045, and the batch size is set to 16.

[0065] 2) Simulation content

[0066] Figure 2 The performance comparison of the technical solution of the present invention and other image transmission methods in terms of the reconstructed image quality on the Kodak24 dataset is shown. Figure 2The horizontal axis represents different signal-to-noise ratios (SNRs), and the vertical axis represents the Learned Perceptual Image Patch Similarity (LPIPS). Since the image transmission method of the present invention is based on image semantic features rather than precise pixels, the above metrics evaluate the image reconstruction effect from the perceptual level and are more suitable than traditional image quality evaluation metrics (such as PSNR, SSIM). The solid square line represents the performance change curve of the present invention, the dashed diamond line represents the performance change curve of traditional WebP + LDPC, and the dashed circle line represents the performance change curve of the semantic communication method WITT. It can be seen from the comparison that the performance of the present invention is better than that of the baseline method, especially at low SNRs.

[0067] Figure 3 The figure shows the performance comparison of the technical solution of the present invention and other image transmission methods for image-based question-and-answer tasks. Figure 3 The horizontal axis represents different signal-to-noise ratios (SNRs), and the vertical axis represents the question-and-answer accuracy. The solid circle line represents the performance change curve of the present invention, the dashed upper triangle line represents the performance change curve of traditional JPEG, and the dashed diamond line represents the performance change curve of traditional WebP. It can be seen that the present invention shows higher answer accuracy under low SNR conditions (for example, between -5dB and 10dB) and has better noise resistance. When the SNR reaches 10dB and above, the answer accuracy of the present invention is comparable to that of the baseline method, with a value close to about 93%. Generally, this solution has good robustness and semantic information retention ability in the UAV image search and rescue question-and-answer scenario.

[0068] Figure 4 It is a visualization effect comparison chart of the present invention and other existing technologies for image transmission and question-and-answer under different SNR conditions. Figure 4 The first row shown is the question-and-answer results of the large language model based on the image transmission scheme of the present invention, the second row is the question-and-answer results of the large language model based on WITT, and the third and fourth rows are the question-and-answer results of the large language model based on traditional JPEG and WebP schemes respectively. It can be seen that even under low SNR conditions, the present invention can achieve the retention of meaningful content in the reconstructed image, that is, it has high semantic fidelity visually. When the SNR increases from 0dB to 20dB, the quality of the reconstructed image and the generated captions improves significantly, and the caption content corresponding to the original image is very close. Figure 5 The figure shows the simulation trajectory diagram of the UAV search and rescue scenario of the present invention. A 6×6 grid scenario is constructed, and scene pictures are placed at each grid position to simulate the search trajectory of the UAV. It is set that when the present invention recognizes that there is a person in the captured image, that is, when the large language model answers yes, the UAV will stop searching, otherwise it will continue to search.

[0069] It should be understood that the above specific embodiments of the present invention are only used for exemplary illustration or explanation of the principles of the present invention, and do not constitute a limitation on the present invention. Therefore, any modifications, equivalent substitutions, improvements, etc. made without departing from the spirit and scope of the present invention shall be included within the protection scope of the present invention. In addition, the appended claims of the present invention are intended to cover all variations and modifications that fall within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.

Claims

1. A large language model driven drone image semantic transmission and question answering system, characterized by: The system includes an image semantic extraction and quantification module, an image reconstruction module, a subtitle generation module, a prompt word generation module and a large language model question answering module; The image semantic extraction and quantization module uses a semantic encoder to extract feature vectors from the original images captured by the drone, and uses a codebook shared by the transceiver to map the potential feature vectors to discrete index sequences, which can greatly reduce the amount of transmitted data. This module is the basis for subsequent modules. The image reconstruction module takes the index sequence as input, inversely maps the latent feature vector according to the shared codebook, and then restores the original image through the semantic decoder. Semantic decoding is the inverse process of semantic encoding, and the original image can be restored according to the latent feature vector. The subtitle generation module takes the reconstructed image output by the image reconstruction module as input to realize text subtitle generation based on semantic images; The prompt word generation module generates a prompt word for question answering based on a fixed question template, combined with text subtitles and user question information. This serves as the input of the subsequent large language model, which can improve the interpretability of the system. The large language model question-answering module is the output module of the system, which uses the advanced LLAMA large language model to achieve high-quality natural language responses.

2. A method for implementing a large language model-driven UAV image semantic transmission and question-answering system, characterized in that: The method comprises the following steps: Step 1: Dataset preprocessing and related large language model deployment; Download the Open Images v6 dataset and select 2,000 training samples and 500 validation sample images. Use the browser to retrieve 20 images of rescue scenes for drone search trajectory simulation and deploy the BLIP and LLAMA large language models locally. Step 2: Extraction and quantification of potential feature vectors of drone images; The semantic encoder is used to extract the potential feature vector of the original image, and the nearest neighbor algorithm is used to find similar vectors in the codebook to replace the feature vector and generate a corresponding index sequence. Step 3: Image reconstruction based on feature vectors: retrieve the corresponding feature vector from the codebook according to the index, and reconstruct the original image using the semantic decoder; Step 4: Image caption generation, using the BLIP framework to automatically generate semantic descriptions of the reconstructed image; Step 5: Prompt word generation and large language model question and answer. Prompt words are generated based on the structured prompt word template, combined with image captions and user questions. The prompt words are input into the large language model to obtain natural language responses and realize question and answer.

3. The method for implementing a large language model driven drone image semantic transmission and question answering system according to claim 2 is characterized in that: Step 1 of the method includes: selecting 2000 samples and 500 sample images from the public Open Images v6 dataset for training and verification, and uniformly adjusting the image size in the dataset to 256×256, and deploying a BLIP model for image caption generation and a LLAMA model for question answering tasks on a cloud server, wherein BLIP is an advanced visual-language pre-training model that implements image feature extraction and caption generation tasks through its visual encoder and language decoder, and LLAMA is a large language model developed by Meta, which is used to understand input prompt words and generate corresponding natural language answers.

4. The method for implementing a large language model driven UAV image semantic transmission and question answering system according to claim 2 is characterized in that: Step 2 of the method includes: after completing the data set preprocessing and related large language model deployment in step 1 above, assuming that the input image is a source image captured by a drone Where H and W represent the height and width of the image respectively, and "3" represents the RGB channel. First, the semantic transmission of the image must be realized. The semantic encoder designed based on the convolutional neural network (CNN) is used to extract the semantics of the original image x captured by the drone to obtain the latent vector representation The process can be expressed as: z e =E α (x) Where E α (·) is a semantic encoding operation with parameter α, z e The semantic encoder is composed of a 5×5 initial convolution, four downsampling modules and an output layer. The downsampling module contains a basic residual block based on ResNet design and an optional attention block. Finally, the features are mapped to the latent vector representation through normalization and convolution layers. Then, the extracted latent vector is vector quantized through the codebook mechanism. Specifically, a discrete codebook is pre-trained in the language For each encoding position of the feature vector output by the encoder, the nearest vector is found in the codebook Z through the nearest neighbor algorithm and replaced with the quantized potential representation z q , the process can be expressed as: Where i and j represent vectors In the latent feature vector z e At each position in , q(·) represents the element-by-element quantization operation, z q is the quantized potential representation, that is, the feature index sequence.

5. The method for implementing a large language model driven drone image semantic transmission and question answering system according to claim 2, characterized in that: Step 3 of the method comprises: transforming the potential representation obtained in step 2 above, i.e., z q Transmission through wireless channels, the channel is modeled as an additive Gaussian white noise channel, and the transmission process can be expressed as: in is the received signal, n is the channel noise, which obeys the symmetric complex Gaussian distribution n~CN(0,σ 2 I), h is the channel gain. At the receiving end, the corresponding feature map is retrieved from the codebook according to the received feature index to obtain the feature vector z d , input to the semantic decoder for image reconstruction, the process can be expressed as: x′=D β (z d ) Where D β (·) is the semantic decoding operation with parameter β, x′ is the reconstructed image, and the semantic decoder used is a symmetric architecture, including 5×5 head convolution, four upsampling modules and a re-encoding module. The upsampling module uses a residual block similar to the encoder and uses the nearest neighbor interpolation method to upsample the image. The re-encoding module consists of 3×3 convolution and Swish activation function. Then, end-to-end training is performed, and the loss function of the whole process is as follows: L1(E,D,Z)=L rec +γL com +L reg L rec =||x-x′|| 2 Among them, sg(·) represents the stop gradient operator, L rec is the reconstruction loss, which is used to measure the difference between the reconstructed image and the original image; L com is the commitment loss, which measures the difference between the quantized latent vector and the output of the semantic encoder; L reg Used to regularize the encoder output. To further improve the quality of the reconstructed image, adversarial training and perceptual loss are introduced. The perceptual loss uses a pre-trained neural network to compare the high-level features of the original image and the reconstructed image, so that the generated image has a feature representation similar to the real image. The perceptual loss function is defined as follows: Among them, L per is the perceptual loss, φ(·) represents the features extracted by the pre-trained network, x and x′ are the original image and the reconstructed image, the adversarial training goal is to make the image generated by the semantic decoder closer to the real image, and the generative adversarial network (GAN) loss is expressed as: in, represents the probability that the discriminator judges the real image x to be true; represents the probability that the discriminator judges that the reconstructed image x′ is real, then the overall loss function is as follows: L total =L rec +γL com +μL 2 +θL per +L reg Among them, γ, μ, θ are the balancing factors of different loss functions. By optimizing this loss function, the semantic encoder, decoder, codebook and discriminator are self-supervised trained.

6. The method for implementing a large language model driven drone image semantic transmission and question answering system according to claim 2, characterized in that: Step 4 of the method comprises: obtaining a reconstructed image, namely x′, according to step 3 above, and processing the reconstructed image using the BLIP model deployed in step 1. First, the image is converted into visual features through its visual encoder, and then the visual features are input into the language decoder to generate corresponding text descriptions. The subtitle generation process can be expressed as: z v =E v (x′) T subtitle =D l (E v (x′)) Among them, z v is the visual feature vector, E v (·) represents the visual encoder, D l (·) represents the language encoder, and T subtitle The generated text description is then filtered through the BLIP model’s filtering system, which evaluates the consistency between the input image features and the generated text description to ensure that the output text description is highly accurate and relevant. Finally, a caption T containing key details of the image (such as objects or people in the scene) is generated, providing effective input for the subsequent large language model question-answering task.

7. The method for implementing a large language model driven UAV image semantic transmission and question answering system according to claim 2, characterized in that: Step 5 of the method comprises: according to the subtitles T generated by the image subtitle system in the above step 4, the present invention combines the task description information provided by the user, forms a prompt word according to the structured prompt word template, and then inputs the generated prompt word into the LLAMA large language model. Based on its powerful natural language understanding and generation capabilities, the model understands and infers the prompt word, generates an effective answer or decision, so as to complete the image-based question-answering task and provide guidance for the next action of the drone.

Citation Information

Cited By

  • Question and answer method for low-altitude multi-modal fusion retrieval

    CN121257755A

  • Compressed domain visual language large model reasoning-oriented system and method

    CN121860063A