Image transmission method and related equipment
Through cross-modal technology, the visual features of the image match the semantic knowledge base is extracted, and differentiated semantic information is generated for transmission, which solves the contradiction between image transmission efficiency and quality and realizes image reconstruction under efficient and low noise.
Patent Information
- Application Number
- CN202510155449.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-07-04
AI Technical Summary
While ensuring image quality, the existing image transmission technology has low transmission efficiency, especially when the wireless channel conditions are poor, it is prone to noise interference, and the image quality decreases under high compression rates, which is difficult to effectively solve the problem of traditional methods.
By extracting the visual features of the image and matching the background knowledge image in the shared semantic knowledge base, differential semantic information is generated for transmission. The receiving end reconstructs the image based on differential semantic information and background knowledge images, and uses cross-modal technology to achieve efficient transmission.
It significantly reduces the amount of data transmitted, improves image transmission efficiency, and ensures the visual quality of the image. It is suitable for communication environments with high latency and low signal-to-noise ratio.
Smart Images

Figure CN120263984A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present disclosure relate to the field of communication technologies, and in particular, to an image transmission method and related devices. Background Art
[0002] With the rapid increase in the intelligent application requirements of wireless communication, the amount of data to be transmitted in future communication networks will experience exponential explosive growth, which poses higher requirements on existing image transmission technologies. How to efficiently transmit image data has become an urgent problem to be solved.
[0003] Traditional image transmission technologies mainly rely on separate source coding and channel coding technical solutions, or achieve image transmission through source-channel joint coding technology. In the separate source coding and channel coding technical solutions, source coding technology is used to compress images to reduce the amount of data transmitted; channel coding is used to enhance the anti-interference ability and reliability during transmission. However, when the wireless channel conditions are poor, the traditional separate source coding and channel coding schemes have low transmission efficiency and are easily interfered by physical noise. Especially when the signal-to-noise ratio decreases, the so-called "cliff effect" will occur, resulting in a significant decline in transmission performance. The source-channel joint coding technology can effectively improve the transmission efficiency and reliability of image data by considering the characteristics of both the source and the channel during the communication process, thereby alleviating the "cliff effect" of the traditional separate source and channel coding schemes under poor channel conditions. However, at high compression rates, it may still lead to a decline in image quality. Especially in practical applications where high-quality image transmission is required, this quality decline is still a problem that cannot be ignored.
[0004] Related technologies have proposed an image transmission scheme that combines generative artificial intelligence and image transmission. This scheme extracts high-level semantic information of images and uses generative models to restore images, thereby significantly reducing the amount of data transmitted and ensuring visual quality. However, this technology still faces the problem of needing to transmit a large amount of semantic information, resulting in insufficient transmission efficiency. Therefore, how to further improve the transmission efficiency while ensuring image quality remains an urgent technical problem to be solved. Summary of the Invention
[0005] In view of this, the purpose of one or more embodiments of the present disclosure is to propose an image transmission method and related devices to solve the problems raised in the background art.
[0006] Based on the above purpose, one or more embodiments of the present disclosure provide an image transmission method. The method includes:
[0007] Performing visual feature extraction on the image to be transmitted to obtain the first visual feature of the image to be transmitted;
[0008] Match the first visual feature with a second visual feature in a preset semantic knowledge base. There is a corresponding background knowledge image for the second visual feature, and the semantic knowledge base is shared by the sender and the receiver.
[0009] Obtain differential semantic information through a preset semantic extraction module according to the first visual feature, the second visual feature, and a preset prompt template. The differential semantic information is a text description of the difference between the image to be transmitted and the background knowledge image.
[0010] Send the differential semantic information and the index corresponding to the second visual feature to the receiver, so that the receiver reconstructs the image to be transmitted based on the differential semantic information and the background knowledge image corresponding to the second visual feature. There is a unique correspondence among the background knowledge image, the index, and the second visual feature.
[0011] Optionally, the semantic knowledge base prestores background knowledge images, as well as the second visual features and indexes corresponding to the background knowledge images. The second visual feature and the first visual feature are obtained according to the same visual feature extraction method.
[0012] Optionally, the matching of the second visual feature in the preset semantic knowledge base according to the first visual feature includes:
[0013] Perform visual retrieval on the first visual feature based on the semantic knowledge base.
[0014] Determine the second visual feature with the highest similarity to the first visual feature through visual retrieval.
[0015] Use the second visual feature with the highest similarity as the target visual feature.
[0016] Optionally, it further includes:
[0017] Determine the background knowledge image from the semantic knowledge base through index retrieval according to the index of the second visual feature.
[0018] Obtain a reconstructed image according to the differential semantic information, the background knowledge image, and a preset visual reconstruction model.
[0019] Optionally, the obtaining of the reconstructed image according to the differential semantic information, the background knowledge image, and a preset visual reconstruction model includes:
[0020] Obtain a latent representation of the reconstructed image through a visual reconstruction sub-model according to the differential semantic information and the background knowledge image.
[0021] Obtain the reconstructed image through a visual decoder according to the latent representation of the reconstructed image.
[0022] Optionally, it further includes:
[0023] Based on the reconstructed image and the pre-trained image detection model, determine whether the reconstructed image meets the preset performance requirements;
[0024] In response to determining that the reconstructed image does not meet the preset performance requirements, feedback a retransmission request.
[0025] Based on the same inventive concept, one or more embodiments of the present disclosure further provide an image transmission device, including:
[0026] A visual feature extraction module, configured to perform visual feature extraction on the image to be transmitted to obtain the first visual feature of the image to be transmitted;
[0027] A visual retrieval module, configured to match the second visual feature in the preset semantic knowledge base according to the first visual feature, where there is a corresponding background knowledge image for the second visual feature, and the semantic knowledge base is shared by the sender and the receiver;
[0028] A semantic information extraction module, configured to obtain differential semantic information through a preset semantic extraction module according to the first visual feature, the second visual feature, and a preset prompt template, where the differential semantic information is a text description of the difference between the image to be transmitted and the background knowledge image;
[0029] A sending module, configured to send the differential semantic information and the index corresponding to the second visual feature to the receiving end, so that the receiving end reconstructs the image to be transmitted based on the differential semantic information and the background knowledge image corresponding to the second visual feature, and there is a unique correspondence between the background knowledge image, the index, and the second visual feature.
[0030] Optionally, it further includes:
[0031] An index retrieval module, configured to determine the background knowledge image from the semantic knowledge base through index retrieval according to the index of the second visual feature;
[0032] An image reconstruction module, configured to obtain a reconstructed image according to the differential semantic information, the background knowledge image, and a preset visual reconstruction model.
[0033] Based on the same inventive concept, one or more embodiments of the present disclosure further provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the image transmission method as described in any one of the above.
[0034] Based on the same inventive concept, one or more embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the image transmission method described in any one of the above.
[0035] As can be seen from the above, the image transmission method provided by one or more embodiments of the present disclosure extracts visual features of the image to be transmitted to obtain the first visual features of the image to be transmitted; matches the second visual features in a preset semantic knowledge base according to the first visual features, where there is a corresponding background knowledge image for the second visual features, and the semantic knowledge base is shared by the sender and the receiver; obtains differential semantic information through a preset semantic extraction module according to the first visual features, the second visual features, and a preset prompt template, where the differential semantic information is a text description of the difference between the image to be transmitted and the background knowledge image; and sends the differential semantic information and the index corresponding to the second visual features to the receiver so that the receiver reconstructs the image to be transmitted based on the differential semantic information and the background knowledge image corresponding to the second visual features, and there is a unique correspondence between the background knowledge image, the index, and the second visual features.
[0036] The technical solution of the present disclosure optimizes the image data transmission process by extracting the cross-modal core semantic information of the image. Specifically, the implementation of the present disclosure not only takes the visual features of the image as the transmission content, but also incorporates the differential semantic information (such as high-level abstract information such as objects, scenes, and situations) between the image and the background knowledge. In this way, the cross-modal core semantic information of the image becomes the basis for image reconstruction. Through the present disclosure, not only can the transmission quality of the image be effectively guaranteed to ensure that the receiver can restore relatively complete image details, but also the amount of data to be transmitted is significantly reduced, thus greatly improving the data transmission efficiency.
[0037] An image transmission device, an electronic device, and a computer-readable storage medium provided by the present disclosure can all implement the steps of the above image transmission method, and thus also have the beneficial effects of the above image transmission method. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in one or more embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only one or more embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1Schematic flowchart of the image transmission method according to one or more embodiments of the present disclosure;
[0040] Figure 2 Schematic flowchart of the image transmission method according to one or more embodiments of the present disclosure;
[0041] Figure 3 Schematic flowchart of the image transmission method according to one or more embodiments of the present disclosure;
[0042] Figure 4 Schematic structural diagram of the image transmission device according to one or more embodiments of the present disclosure;
[0043] Figure 5 Schematic structural diagram of the image transmission device according to one or more embodiments of the present disclosure;
[0044] Figure 6 Schematic diagram of the simulation experiment results according to one or more embodiments of the present disclosure;
[0045] Figure 7 Schematic diagram of the hardware structure of the electronic device according to one or more embodiments of the present disclosure. Detailed implementation manners
[0046] To make the objectives, technical solutions and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0047] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The "first", "second" and similar terms used in one or more embodiments of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0048] As described in the background art, how to further improve the transmission efficiency while ensuring the image quality is still a technical problem to be solved urgently.
[0049] Although the related art has proposed to transmit the semantic information of the image to improve the transmission efficiency, the image quality may still deteriorate at high compression ratios.
[0050] Cross-modal technology is a technology that realizes information interaction and processing by integrating different modal data (such as images, texts, audios, etc.). Its core lies in breaking the barriers between modalities and achieving efficient conversion and collaborative processing between multi-modal data through unified representation learning and semantic alignment. Through cross-modal representation learning and fusion mechanisms, this technology maps different modal data to a unified feature space, thereby achieving semantic alignment and information fusion between modalities. Cross-modal technology is widely applied in fields such as cross-modal retrieval, generation, multi-modal intelligent assistants, and autonomous driving, and has the advantages of efficient information fusion and enhanced robustness, and can effectively solve the problem of modal heterogeneity in multimedia data.
[0051] In the process of implementing the present disclosure, the applicant found that applying cross-modal technology to image transmission can significantly reduce the amount of transmitted data and remarkably improve the transmission efficiency. In addition, cross-modal technology combines the semantic associations of multi-modal data, can maintain the integrity and accuracy of image semantics in a dynamic environment, and effectively alleviates the performance degradation problem of traditional technologies under low signal-to-noise ratio conditions.
[0052] Therefore, the present disclosure proposes an image transmission method. By extracting the visual features of an image and matching them with the background knowledge images in a shared semantic knowledge base, differential semantic information is generated for transmission, so that the receiving end can reconstruct the image to be transmitted based on the differential semantic information and the shared background knowledge. This solution makes full use of the advantages of cross-modal technology, significantly reduces the amount of transmitted data, and at the same time ensures the visual quality of the image through the efficient extraction and transmission of semantic information.
[0053] Reference Figure 1 , the image transmission method according to one or more embodiments of the present disclosure includes the following steps:
[0054] Step S101: Extract visual features of the image to be transmitted to obtain the first visual features of the image to be transmitted;
[0055] Step S102: Match the second visual features in a preset semantic knowledge base according to the first visual features. There is a corresponding background knowledge image for the second visual features, and the semantic knowledge base is shared by the sending end and the receiving end;
[0056] Step S103: Obtain differential semantic information through a preset semantic extraction module according to the first visual features, the second visual features, and a preset prompt template. The differential semantic information is a text description of the difference between the image to be transmitted and the background knowledge image;
[0057] Step S104: Send the above-mentioned differential semantic information and the index corresponding to the second visual feature to the receiving end, so that the receiving end reconstructs the image to be transmitted based on the above-mentioned differential semantic information and the background knowledge image corresponding to the second visual feature. There is a unique correspondence among the background knowledge image, the index, and the second visual feature.
[0058] The first visual feature in step S101 represents the key information that can characterize the image content extracted from the image to be transmitted, and is used to describe the visual attributes of the image to be transmitted (such as color, texture, shape, spatial layout, etc.). In the present disclosure, the first visual feature has the functions of matching the semantic knowledge base and assisting in generating differential semantic information.
[0059] In some implementations of the present disclosure, a visual encoder can be used to map the image to be transmitted (i.e., the original image) I src ∈R C×W×H to the visual feature f v ∈R c×w×h for extracting the first visual feature. The above-mentioned visual encoder can select a suitable visual feature extraction model according to the actual application scenario. In some examples, the above-mentioned visual encoder can be a visual encoder based on a convolutional neural network, a visual encoder based on a self-attention mechanism neural network, a visual encoder in multi-modal fusion, or a generative visual encoder. The present disclosure does not make a specific limitation on the selection of the visual encoder.
[0060] According to the above content, the above-mentioned first visual feature can be expressed as:
[0061] where c << C, w << W, h << H.
[0062] The second visual feature in step S102 represents the key information that can characterize the image content extracted from the background knowledge image, and is used to describe the visual attributes of the background knowledge image. In the present disclosure, the visual feature has the functions of assisting in generating differential semantic information and supporting image reconstruction.
[0063] The above-mentioned background knowledge image is an image in the preset semantic knowledge base, and can provide known visual information and semantic background as a reference image. The above-mentioned background knowledge image is the extraction basis of the differential semantic information and also the reconstruction basis of the reconstructed image.
[0064] The second visual features and background knowledge images of the present disclosure are stored in the semantic knowledge base. Each background knowledge image in the semantic knowledge base corresponds to a unique second visual feature and index. That is to say, there is a unique correspondence among the background knowledge images, the second visual features, and the indexes in the semantic knowledge base. The semantic knowledge base is shared by the sending end and the receiving end, so that the semantic knowledge base can serve as the basis for semantic information extraction and image reconstruction, providing a consistent semantic reference and visual feature matching standard for both ends.
[0065] Specifically, in the technical solution of the present disclosure, the sending end matches the visual features of the image to be transmitted through the semantic knowledge base, extracts the differential semantic information between it and the background knowledge images, while the receiving end combines the corresponding background knowledge images in the shared semantic knowledge base to understand these differential descriptions and perform image reconstruction. Only when both ends share the same semantic knowledge base can the accurate transmission of semantic information and the correctness of image reconstruction be ensured, thereby achieving efficient and accurate image transmission and reconstruction.
[0066] It should be noted that the second visual features should be obtained by the same means as the first visual features. For example, if the first visual features are extracted by a visual encoder based on a convolutional neural network, the second visual features should also be extracted by a visual encoder based on a convolutional neural network. This is to ensure their comparability and consistency in visual content, so that they can be effectively matched and compared. By using the same feature extraction method, the visual features can be ensured to be consistent in the standardized feature space, improving the matching accuracy and reducing mis-matching or errors. In addition, using a consistent feature extraction method can ensure that the subsequent generated differential semantic information accurately reflects the differences between images, which is helpful for accurate image reconstruction and semantic information extraction.
[0067] In the embodiments of the present disclosure, the second visual features can be determined in the above-mentioned semantic knowledge base according to the first visual features of the image to be transmitted through a visual retrieval method. In the embodiments of the present disclosure, the above-mentioned visual retrieval method means determining the second visual feature with the highest similarity to the above-mentioned first visual feature from the semantic knowledge base. That is to say, the process of obtaining the visual features through the above-mentioned visual retrieval method can be expressed as:
[0068]
[0069] where, f k represents the second visual feature of the background knowledge image, f v represents the first visual feature of the image to be transmitted, represents the visual retrieval algorithm.
[0070] In the embodiments of the present disclosure, the retrieval method may include, but is not limited to, algorithms based on cosine similarity, algorithms based on other visual feature contents of images (such as texture features, local features, etc.), and cross-modal enhancement. Taking the retrieval method based on cosine similarity as an example, the retrieval process can be expressed as:
[0071]
[0072] Among them, represents a visual encoder based on the candidate background knowledge set K (a set composed of all candidate images in the semantic knowledge base that can be used as background knowledge images), and the candidate background knowledge set K includes several candidate background knowledge k, d cosine represents the cosine distance, and the calculation formula can be u and v respectively represent the visual feature vectors of the image to be transmitted and the candidate background knowledge image. The background knowledge image is determined by comparing the candidate background knowledge with the smallest cosine distance.
[0073] In the technical solution of the present disclosure, the differential semantic information between the image to be transmitted and the background knowledge image is obtained according to the first visual feature and the second visual feature. In order to provide clear task instructions and semantic directions for the multi-modal large language model, enabling it to focus on extracting the differential semantic information between the original image and the background knowledge, the technical solution of the present disclosure also introduces a prompt template. In this way, the multi-modal large prediction model can be guided to analyze and identify the differences in the input image features, so as to more accurately output the difference description, avoid the interference of irrelevant information, and enhance the pertinence and effectiveness of semantic extraction.
[0074] In some embodiments of the present disclosure, the above prompt template can be set to " <image1>and <image2>What are the differences between <image1>and <image2>)”. Among them, image1 represents the image to be transmitted or the background knowledge image, and image2 represents the background knowledge image or the image to be transmitted.
[0075] In other words, in the technical solution of the present disclosure, the first visual feature of the image to be transmitted, the second visual feature of the background knowledge image, and the preset prompt template are input into the trained multi-modal large language model, and the trained multi-modal large language model will output the differential semantic information between the image to be transmitted and the background knowledge image. This process can be expressed as:
[0076]
[0077] Among them, s represents the differential semantic information, f v represents the first visual feature of the image to be transmitted, f k represents the second visual feature of the background knowledge image, prompt represents the prompt template, represents including a semantic information extraction function.
[0078] After that, the above differential semantic information and the index of the second visual feature are sent to the receiving end, so that the receiving end can use the index of the second visual feature to determine the background knowledge image, and reconstruct the image according to the background knowledge image using the differential semantic information.
[0079] In the embodiments of the present disclosure, the Color Layout Descriptor (CLD) can be used as the index of the second visual feature. CLD is a visual descriptor defined in the MPEG-7 standard. By segmenting the image, extracting the average color value, and using the discrete cosine transform to extract the low-frequency components, a compact and representative feature vector is generated. This feature vector not only retains the main color information of the image, but also reflects the layout of the colors in space. Since the color layout of each image is unique, and CLD generates a unique feature identifier for each image through an accurate description of the color space distribution, target images similar to the query image can be quickly located and retrieved in a large-scale image database.
[0080] Compared with other indexing methods, CLD has the advantages of fast calculation and storage speed and high discrimination. Specifically, the feature vector of CLD can uniquely represent the color layout of the image, and can effectively distinguish even in the case of similar image content but different layouts.
[0081] In the embodiments of the present disclosure, a symbol with a unique identification function can also be used as the index, and the present disclosure does not limit the specific form of the index here.
[0082] Before sending the above differential semantic information and the above index to the receiving end, it is also necessary to encode the above differential semantic information.
[0083] Considering that the differential semantic information obtained in the present disclosure, which is the text description of the semantic difference between the image to be transmitted and the background knowledge image, can be expressed in natural language. Therefore, some embodiments of the present disclosure achieve efficient and reliable data transmission between the sending and receiving ends by means of source coding and channel coding for the text modality.
[0084] In the embodiments of the present disclosure, for the transmission of differential semantic information in the text modality, a separate source-channel coding can be adopted, or a source-channel joint coding can be adopted.
[0085] In the source-channel joint coding scheme of some embodiments of the present disclosure, the encoding process of the above differential semantic information at the sending end can be expressed as:
[0086]
[0087] Among them, z represents the semantic symbol transmitted through the channel, s represents the differential semantic information, represents the semantic encoder, represents the channel encoder. Through the joint coding of the semantic encoder and the channel encoder, the source characteristics and channel conditions can be considered simultaneously, and the coding strategy can be optimized, so as to improve the transmission efficiency and reliability under a limited code length.
[0088] In the embodiments of the present disclosure, the semantic encoder can be composed of multiple Transformer encoder layers to effectively extract the semantic associations of the differential description text, and the channel encoder can be composed of multiple linear layers with different dimensions to adjust the semantic symbols. In the embodiments of the present disclosure, power normalization can also be performed on the semantic symbols.
[0089] After the above semantic symbols are transmitted to the receiving end through the wireless fading channel, the semantic symbols received at the receiving end can be expressed as:
[0090] z′ = hz + n;
[0091] Among them, h represents the channel gain between the sending end and the receiving end, represents the Gaussian additive noise.
[0092] After channel decoding and semantic decoding of the above received semantic symbols, the restored semantic information can be obtained:
[0093]
[0094] Among them, s′ represents the restored semantic information, represents the channel decoder, Denote the semantic decoder. In an embodiment of the present disclosure, the above channel decoder can be constructed by multiple linear layers with different dimensions, and the semantic decoder can be composed of multiple Transformer decoder layers to achieve reliable data transmission of the original semantic information.
[0095] An embodiment of the present disclosure further includes: performing visual reconstruction based on the restored semantic information to obtain a reconstructed image. As Figure 2 shown, this process may specifically include:
[0096] Step S201: Determine the above background knowledge image from the above semantic knowledge base through index retrieval according to the index of the above second visual feature;
[0097] Step S202: Obtain a reconstructed image according to the above differential semantic information, the above background knowledge image, and a preset visual reconstruction model.
[0098] According to the above, the semantic knowledge bases of the receiving end and the sending end are synchronized, and in step S201, the background knowledge image can be retrieved from the above semantic knowledge base using the index sent by the sending end.
[0099] The process of determining the background knowledge image according to the index can be expressed as:
[0100]
[0101] where d k represents the index shared by the sending end and the receiving end, k represents the background knowledge image matched by the index retrieval method, and represents the index retrieval method. Determining the background knowledge image through index retrieval can ensure that there is a background knowledge image matching the extracted differential semantic information during the image reconstruction process.
[0102] It should be noted that in the present disclosure, the sending end uses a visual retrieval method to determine the background knowledge image, and the receiving end uses an index retrieval method to achieve fast and accurate image indexing and reconstruction. Specifically, the visual retrieval method can quickly locate the background knowledge image similar to the image to be transmitted through efficient feature extraction and matching techniques. The index retrieval method further improves the retrieval speed and accuracy by optimizing the index structure and retrieval algorithm. This optimization not only improves the efficiency of image retrieval but also provides a more accurate reference for image reconstruction at the receiving end, thereby achieving efficient and accurate image transmission and reconstruction under dynamic channel conditions.
[0103] After that, visual reconstruction can be performed relying on the received differential semantic information and background knowledge image to obtain a reconstructed image.
[0104] In an embodiment of the present disclosure, the process of reconstructing an image may include: First, according to the differential semantic information and the background knowledge image, obtain the latent representation of the reconstructed image through the visual reconstruction sub-model; according to the latent representation of the reconstructed image, obtain the reconstructed image through the visual decoder.
[0105] Specifically, the process of obtaining the latent representation can be expressed as:
[0106]
[0107] Among them, x T represents the latent representation of the reconstructed image, k represents the background knowledge image, s' represents the recovered differential semantic information, represents the conditional diffusion model.
[0108] In the present disclosure, the latent representation of the reconstructed image is generated through the reverse denoising process. Therefore, the conditional probability model of the reverse diffusion process in the present disclosure can be expressed as:
[0109]
[0110] Among them, T represents the number of iterations of the reverse denoising process.
[0111] In an embodiment of the present disclosure, a generative model such as styleGAN can also be used to obtain the latent representation of the reconstructed image.
[0112] After that, input the above latent representation into the visual decoder to obtain the reconstructed image I rec ∈R C×W×H , and the reconstructed image can be expressed as:
[0113]
[0114] An embodiment of the present disclosure may further include: judging whether the reconstructed image meets the preset performance requirements according to the reconstructed image and a preset image detection module; correspondingly, in response to determining that the reconstructed image does not meet the preset performance requirements, feedback a retransmission request.
[0115] Next, taking a specific embodiment of the present disclosure as an example, the technical solution of the present disclosure will be further described. The technical solution is as Figure 3 shown, including an image sending end and an image receiving end.
[0116] After receiving the captured image (i.e., the image to be transmitted), the image receiver first obtains the visual features of the captured image and the visual features of the background knowledge image through a visual encoder and a semantic knowledge base. According to the above-mentioned visual features of the captured image and the visual features of the background knowledge image, through a semantic extraction module based on a Multimodal Large Language Model (MLLM), differential semantic information can be obtained. The image sender jointly encodes the source and channel of the above differential semantic information through a semantic encoder and a channel encoder, and then transmits it to the image receiver together with the background knowledge image index.
[0117] After receiving the semantic symbols sent by the image sender, the image receiver first performs channel decoding through a channel decoder, and then obtains the restored differential semantic information through a semantic decoder. The process of further obtaining the reconstructed image may include performing reverse denoising using a conditional diffusion visual reconstruction sub-model to obtain the latent representation of the reconstructed image, and using a visual decoder to decode the above latent representation to obtain the reconstructed image. To ensure the quality of image transmission, after obtaining the reconstructed image, a detection module can also be used to check the reconstructed image. If the detection result is not good, feedback information can be sent to the image sender to request the image sender to re-transmit the image.
[0118] The technical solution of the present disclosure can be applied to cross-modal image transmission in various scenarios. For example, the technical solution of the present disclosure can be applied to cross-modal image transmission for satellite communication.
[0119] With the continuous development of satellite communication technology, the application of satellite platforms in the fields of remote sensing, meteorological monitoring, disaster warning, etc. has gradually increased. Traditional satellite image transmission methods rely on directly transmitting image data. However, due to the large amount of image data and limited transmission bandwidth, especially in environments with high latency and low signal-to-noise ratio, the efficiency and quality of image transmission are often affected. By adopting the image transmission method proposed in the present disclosure and using semantic information extraction and reconstruction technology, the satellite platform can transmit only the high-level semantic information of the image, while the ground station reconstructs high-quality images through a generative model. This solution can significantly improve the efficiency of satellite image transmission and is particularly suitable for satellite communication environments with high latency and limited bandwidth.
[0120] Specifically, during the implementation of the satellite communication scenario, the satellite platform and the ground station pre-share relevant background knowledge before image transmission, and respectively pre-deploy a trained semantic extraction module and a visual reconstruction module. The satellite platform collects image data of the ground scene or other spaces in real time through a high-resolution camera, and inputs the collected images into the semantic extraction module to extract high-level semantic information. The extracted semantic information includes text descriptions of differential information such as object categories, positions, and scene structures in the images. Then, the satellite platform encodes the extracted semantic information using source-channel joint coding technology and transmits it to the ground station through the satellite communication link. The ground station receives and decodes the restored semantic information. Then, the ground station uses the visual reconstruction module to perform image reconstruction based on the restored semantic information and the pre-shared background knowledge with the help of a generative model, and detects the quality of the generated image to ensure that the similarity between the generated image and the original image in terms of vision and semantics reaches a preset standard.
[0121] The technical solution of the present disclosure can also be applied to cross-modal image transmission for the surveillance scenario.
[0122] With the wide application of surveillance technologies, more and more surveillance scenarios require efficient and low-bandwidth image transmission solutions. Traditional video surveillance systems transmit high-definition video data continuously. Although they can provide real-time surveillance images, in environments with limited bandwidth or poor network conditions, they often face problems such as large transmission delays, high bandwidth occupancy, and increased costs. By adopting the image transmission method proposed in the present disclosure, the required transmission bandwidth can be significantly reduced, and the transmission efficiency can be improved while ensuring the image quality.
[0123] In the surveillance scenario, the surveillance device and the receiving-end system pre-share background knowledge. This background knowledge is static background pictures of the monitored scenario, including common object types, scene elements (such as buildings, roads, vehicles, etc.), to ensure semantic consistency of the surveillance scenario. The surveillance camera collects image data in the monitored area in real time and inputs it into the pre-deployed semantic extraction module to extract high-level semantic information, such as dynamic change information like vehicles passing by or crowded people flow. Then, the extracted semantic information is encoded and transmitted to the receiving end through a network (wireless network or wired connection). The receiving end receives and decodes the semantic information, and based on the decoded semantic information, combined with the pre-shared background knowledge, performs image reconstruction for subsequent task execution.
[0124] According to the above content, the image transmission method proposed in the present disclosure extracts the cross-modal core semantic information of the collected images, requires less data to be transmitted, can reduce the communication burden of the wireless channel, reduce the delay and energy consumption of data transmission, and is suitable for deployment in high-computing future communication environments with limited communication resources. Through the present disclosure, efficient image transmission in actual limited communication scenarios can be effectively achieved.
[0125] To verify the technical effects of the technical solution of the present disclosure, the applicant also conducted a simulation experiment. In this simulation experiment, the MagicBrush dataset was selected. The images in this dataset are all uniformly sized at 3×512×512, and the instruction score was defined as the retransmission performance threshold, which can be defined as:
[0126]
[0127] where C V and C T respectively represent the visual encoder and the text encoder of the CLIP model. I′, K, and s′ respectively represent the reconstructed image, the background knowledge image shared at the receiving end, and the restored semantic information. The retransmission performance threshold can be set to 0, and the retransmission basic image transmission scheme adopts an image transmission scheme based on joint source-channel coding.
[0128] In this simulation experiment, the semantic extraction module of the present technical solution was constructed using the llama3-llava-next-8b model, and the visual reconstruction module was constructed using the InstructPix2Pix model.
[0129] The results of the simulation experiment are as Figure 6 shown, where GC-SemCom represents the image transmission scheme proposed by the present disclosure technical solution, Deep JSCC represents the image transmission scheme of joint source-channel coding in the background art, and JEPG+LDPC represents a separate source-channel coding image transmission scheme that uses JEPG as the source coding and LDPC as the channel coding.
[0130] It can be seen from the simulation results that the proposed cross-modal image transmission method still maintains high-quality transmission effects under significant noise conditions compared to the image transmission scheme based on joint source-channel coding. At the same time, compared with the traditional separate source coding and channel coding schemes, the proposed method also effectively reduces the "cliff effect" brought about by the decrease in signal-to-noise ratio.
[0131] In addition, compared with the comparison scheme, the proposed method requires less data transmission, greatly reducing the communication cost. Specifically, for the acquired images with a dimension of 3×512×512, the proposed method only requires the sender to upload semantic symbols with a dimension of 1×30×16, accounting for 0.06% of the original image.
[0132] It can be understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities.
[0133] It should be noted that the method of one or more embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of one or more embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0134] It should be noted that the above specific embodiments of the present disclosure have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0135] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides an image transmission device. As Figure 4 shown, the device includes:
[0136] A visual feature extraction module 11, configured to perform visual feature extraction on the image to be transmitted to obtain the first visual feature of the image to be transmitted;
[0137] A visual retrieval module 12, configured to match the first visual feature with a second visual feature in a preset semantic knowledge base, where there is a corresponding background knowledge image for the second visual feature, and the semantic knowledge base is shared by the sender and the receiver;
[0138] A semantic information extraction module 13, configured to obtain differential semantic information through a preset semantic extraction module according to the first visual feature, the second visual feature, and a preset prompt template, where the differential semantic information is a text description of the difference between the image to be transmitted and the background knowledge image;
[0139] A sending module 14, configured to send the differential semantic information and the index corresponding to the second visual feature to the receiver, so that the receiver reconstructs the image to be transmitted based on the differential semantic information and the background knowledge image corresponding to the second visual feature.
[0140] As Figure 5 shown, the device further includes:
[0141] An index retrieval module 21, configured to determine the background knowledge image from the semantic knowledge base through index retrieval according to the index of the second visual feature;
[0142] The image reconstruction module 22 is configured to obtain a reconstructed image according to the differential semantic information, the background knowledge image, and a preset visual reconstruction model.
[0143] For convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing one or more embodiments of the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0144] The device in the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated herein.
[0145] Figure 7 FIG. shows a more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0146] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present disclosure.
[0147] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of the present disclosure through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0148] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0149] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. The communication module can achieve communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.).
[0150] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0151] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.
[0152] The electronic device of the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here.
[0153] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0154] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary, and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments of the present disclosure as described above, and they are not provided in detail for the sake of brevity.
[0155] In addition, for simplicity of explanation and discussion, and so as not to render one or more embodiments of the present disclosure difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid rendering one or more embodiments of the present disclosure difficult to understand, and this also takes into account the fact that details of the implementation of such block diagram devices are highly dependent on the platform on which one or more embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the understanding of those of ordinary skill in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those of ordinary skill in the art that one or more embodiments of the present disclosure may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions are to be regarded as illustrative rather than restrictive.
[0156] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations thereof will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0157] One or more embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of the present disclosure shall be included within the scope of protection of the present disclosure.
Claims
1. An image transmission method, characterized in that, Including: Performing visual feature extraction on the image to be transmitted to obtain the first visual feature of the image to be transmitted; Matching the second visual feature in the preset semantic knowledge base according to the first visual feature, where there is a corresponding background knowledge image for the second visual feature, and the semantic knowledge base is shared by the sender and the receiver; Obtaining differential semantic information through a preset semantic extraction module according to the first visual feature, the second visual feature, and a preset hint template, where the differential semantic information is a text description of the difference between the image to be transmitted and the background knowledge image; Sending the differential semantic information and the index corresponding to the second visual feature to the receiver so that the receiver reconstructs the image to be transmitted based on the differential semantic information and the background knowledge image corresponding to the second visual feature, and there is a unique correspondence among the background knowledge image, the index, and the second visual feature.
2. The method according to claim 1, wherein The semantic knowledge base prestores the background knowledge image, as well as the second visual feature and index corresponding to the background knowledge image, and the second visual feature and the first visual feature are obtained according to the same visual feature extraction method.
3. The method according to claim 2, characterized in that The matching of the second visual feature in the preset semantic knowledge base according to the first visual feature includes: Performing visual retrieval on the first visual feature based on the semantic knowledge base; Determining the second visual feature with the highest similarity to the first visual feature through visual retrieval; Taking the second visual feature with the highest similarity as the target visual feature.
4. The method according to claim 2, wherein Also including: Determining the background knowledge image from the semantic knowledge base through index retrieval according to the index of the second visual feature; Obtaining a reconstructed image according to the differential semantic information, the background knowledge image, and a preset visual reconstruction model.
5. The method according to claim 4, characterized in that The obtaining of the reconstructed image according to the differential semantic information, the background knowledge image, and a preset visual reconstruction model includes: Obtaining a latent representation of the reconstructed image through a visual reconstruction sub-model according to the differential semantic information and the background knowledge image; Obtaining the reconstructed image through a visual decoder according to the latent representation of the reconstructed image.
6. The method according to claim 5, characterized in that, Also including: Judging whether the reconstructed image meets the preset performance requirements according to the reconstructed image and a pre-trained image detection model; Responding to determining that the reconstructed image does not meet the preset performance requirements and feedbacking a retransmission request.
7. An image transmission device, characterized in that, Including: A visual feature extraction module configured to perform visual feature extraction on the image to be transmitted to obtain the first visual feature of the image to be transmitted; A visual retrieval module configured to match the second visual feature in the preset semantic knowledge base according to the first visual feature, where there is a corresponding background knowledge image for the second visual feature, and the semantic knowledge base is shared by the sender and the receiver; A semantic information extraction module configured to obtain differential semantic information through a preset semantic extraction module according to the first visual feature, the second visual feature, and a preset hint template, where the differential semantic information is a text description of the difference between the image to be transmitted and the background knowledge image; A sending module, configured to send the differential semantic information and the index corresponding to the second visual feature to a receiving end, so that the receiving end reconstructs the image to be transmitted based on the differential semantic information and the background knowledge image corresponding to the second visual feature, and there is a unique correspondence among the background knowledge image, the index, and the second visual feature.
8. The device according to claim 7, characterized in that, Further comprising: An index retrieval module, configured to determine the background knowledge image from the semantic knowledge base through index retrieval according to the index of the second visual feature; An image reconstruction module, configured to obtain a reconstructed image according to the differential semantic information, the background knowledge image, and a preset visual reconstruction model.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.