A semantic communication method, apparatus, device, and storage medium based on a generative AI large model.

By using generative AI large models for semantic segmentation and text generation, combined with channel coding technology, the problem of limited accuracy and reliability of semantic communication in existing technologies is solved, and efficient, reliable and high-quality image transmission is achieved.

CN120725928BActive Publication Date: 2025-12-02TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511141469.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-12-02
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

The accuracy and reliability of semantic communication in existing technologies are limited. Traditional communication systems fail to effectively distinguish the semantic importance of different regions of an image, resulting in bandwidth constraints and reduced transmission efficiency.

Method used

Semantic segmentation is performed using a generative AI model to distinguish between key and non-key regions of an image. Key regions are transmitted without distortion, while concise descriptive text is generated for non-key regions. Data compression and transmission are performed using channel coding techniques. The receiving end reconstructs non-key regions based on the descriptive text, ensuring the overall consistency of the reconstructed image.

Benefits of technology

It enables fine-grained understanding of image content, reduces transmission bandwidth, improves transmission efficiency, ensures high-fidelity transmission of critical areas and semantic integrity of non-critical areas, and solves the problems of accuracy and reliability of semantic communication in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725928B_ABST
    Figure CN120725928B_ABST
Patent Text Reader

Abstract

This application provides a semantic communication method, apparatus, device, and storage medium based on a generative AI large model, relating to the field of communication technology. By semantically segmenting the image to be transmitted, key image regions and non-key image regions are obtained. A lossless transmission strategy is adopted for the key image regions, while a text generation model is used to generate concise descriptive text for the non-key image regions. The receiving end reconstructs the non-key image regions based on the descriptive text, generating a reconstructed image including the original pixel data of the key image regions and the synthesized pixel data of the non-key image regions. This solves the technical problem of limited accuracy and reliability in existing semantic communication technologies, alleviating bandwidth resource constraints, improving transmission efficiency, and achieving high-fidelity transmission of key image regions and semantically complete transmission of non-key image regions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a semantic communication method, apparatus, device and storage medium based on a generative AI large model. Background Technology

[0002] With the rapid development of internet and mobile communication technologies, the demand for high-quality image transmission has increased dramatically. Traditional communication systems generally adopt a non-discriminatory transmission mode, striving for lossless reproduction of each pixel. However, due to the inherent differences in visual importance and information capacity among different regions of an image, a large amount of redundant data is inevitably generated, leading to bandwidth constraints and decreased transmission efficiency.

[0003] To alleviate bandwidth resource constraints and improve transmission efficiency, existing technologies have proposed semantic communication methods based on large AI models. These methods involve semantic extraction and transmission of the image to be transmitted. The receiving end, based on a knowledge base, fuses the received semantic information with prior knowledge to reconstruct the image to be transmitted. However, since the construction of the knowledge base relies on limited training data, it is difficult to comprehensively cover all semantic scenarios, leading to discrepancies between the reconstructed image and the image to be transmitted, thus reducing the accuracy and reliability of semantic communication. Summary of the Invention

[0004] This application provides a semantic communication method, apparatus, device, and storage medium based on a generative AI large model to solve the problem of limited accuracy and reliability of semantic communication in the prior art.

[0005] To address the aforementioned issues, this application discloses a semantic communication method based on a generative AI large model, comprising the following steps:

[0006] The sending end performs semantic segmentation on the image to be transmitted to obtain masking information, the original pixel data of the key image region of the image to be transmitted, and the original pixel data of the non-key image region of the image to be transmitted. The masking information is used to characterize the region that needs to be reconstructed by the receiving end.

[0007] The sending end uses the original pixel data of the non-critical image region to generate descriptive text for the non-critical image region through a text generation model. The descriptive text for the non-critical image region is used to describe the semantic information of the non-critical image region.

[0008] The sending end sends the masking information, the original pixel data of the key image region, and the descriptive text of the non-key image region to the receiving end;

[0009] The receiving end uses the masking information, the original pixel data of the key image region, and the descriptive text of the non-key image region to reconstruct the synthetic pixel data of the non-key image region through an image reconstruction model, and obtains the reconstructed image corresponding to the image to be transmitted based on the synthetic pixel data of the non-key image region and the original pixel data of the key image region.

[0010] Furthermore, the training process for the text generation model and the image reconstruction model includes:

[0011] The transmitting end performs semantic segmentation on the sample image to obtain sample masking information, original pixel data of the key image region of the sample that accounts for α proportion of the sample image, and original pixel data of the non-key image region of the sample image. The sample masking information is used to characterize the region that the receiving end needs to reconstruct the image.

[0012] The sending end uses the original pixel data of the non-critical image region of the sample to generate descriptive text of the non-critical image region of the sample through a pre-trained text generation model. The descriptive text of the non-critical image region of the sample is used to describe the semantic information of the non-critical image region of the sample.

[0013] The sending end sends the sample masking information, the original pixel data of the key image region of the sample, and the descriptive text of the non-key image region of the sample to the receiving end.

[0014] The receiving end uses the sample masking information, the original pixel data of the key image region of the sample, and the descriptive text of the non-key image region of the sample to reconstruct the synthetic pixel data of the non-key image region of the sample through a pre-trained image reconstruction model, and obtains the sample reconstruction image corresponding to the sample image based on the synthetic pixel data of the non-key image region of the sample and the original pixel data of the key image region of the sample.

[0015] The difference between the reconstructed image corresponding to the sample image and the sample image is less than [a certain value]. To achieve this, the model parameters of the pre-trained text generation model and the pre-trained image reconstruction model are updated to obtain the text generation model and the image reconstruction model, wherein, This is a preset constant.

[0016] Furthermore, the training dataset for the pre-trained text generation model is generated according to the following steps:

[0017] The original image is segmented into multiple local blocks, which are then input into the visual encoder of the base model to obtain the aggregated local features encoded in blocks. The base model has image understanding capabilities.

[0018] The original image is input into the visual encoder of the basic large model to obtain the global features of the original image scaled and encoded.

[0019] The aggregated local features from block encoding and the global features from the original image scaling encoding are fused to obtain the fused features;

[0020] Based on the fusion features, the text encoder of the basic large model generates the descriptive text of the original image;

[0021] Based on the difference between the descriptive text of the original image and the correct descriptive text of the original image, keep the parameters of the visual encoder of the basic large model unchanged, update the parameters of the text encoder of the basic large model, and obtain the trained text description large model.

[0022] Using the trained text description model, labeled text is added to the original image to obtain the training dataset of the pre-trained text generation model. The labeled text is used to describe the semantic information of non-critical regions of the original image.

[0023] Furthermore, it also includes:

[0024] The original image is input into multiple reference small models to obtain multi-level descriptive text of the original image. The multiple reference small models are used for at least one of the following: extracting global semantic information of the original image and generating a preliminary description, refining the object attributes of local image regions of the original image, identifying the text embedded in the original image, and segmenting and associating the spatial location of key image regions of the original image.

[0025] Based on the difference between the descriptive text of the original image and the correct descriptive text of the original image, while keeping the parameters of the visual encoder of the basic large model unchanged, the parameters of the text encoder of the basic large model are updated to obtain the trained text description large model, including:

[0026] Using the multi-level descriptive text of the original image as a contextual guide signal, and based on the difference between the descriptive text of the original image and the correct descriptive text of the original image, while keeping the parameters of the visual encoder of the basic large model unchanged, the text encoder of the basic large model is updated to obtain the trained text description large model.

[0027] Furthermore, it also includes:

[0028] A text description mini-model is selected, which has the ability to describe the semantic information of the image;

[0029] Using the training dataset of the pre-trained text generation model, with the goal of learning the ability to describe the semantic information of images from the trained large text description model, the small text description model is trained to improve the ability of the small text description model to describe the semantic information of images, thus obtaining the pre-trained text generation model.

[0030] Furthermore, it also includes:

[0031] The accuracy of the original model weights of the visual projection layer and the text encoder of the text description mini-model is preserved, and the accuracy of the original model weights of the remaining layers of the text description mini-model is quantized to obtain the quantized text description mini-model.

[0032] The text description mini-model is trained to improve its ability to describe the semantic information of images, resulting in the pre-trained text generation model, comprising:

[0033] The quantized text description mini-model is trained to improve its ability to describe the semantic information of images, thus obtaining the pre-trained text generation model.

[0034] Furthermore, the loss functions used to train the text description mini-model include an image-text alignment loss function and a distillation loss function. The image-text alignment loss function is used to measure the degree of matching between the text description output by the text description mini-model for an image and the semantic information of the image.

[0035] The distillation loss function is used to measure the degree of consistency between the text description output by the small text description model for the image and the text description output by the trained large text description model for the image.

[0036] To address the aforementioned technical problems, this application also discloses a semantic communication device based on a generative AI large model, comprising:

[0037] The segmentation module is used by the sending end to perform semantic segmentation on the image to be transmitted, and obtain masking information, original pixel data of key image regions of the image to be transmitted, and original pixel data of non-key image regions of the image to be transmitted. The masking information is used to characterize the regions that need to be reconstructed by the receiving end.

[0038] The generation module is used by the sending end to generate descriptive text for the non-critical image region using the original pixel data of the non-critical image region and a text generation model. The descriptive text for the non-critical image region is used to describe the semantic information of the non-critical image region.

[0039] The sending module is used by the sending end to send the masking information, the original pixel data of the key image region, and the descriptive text of the non-key image region to the receiving end;

[0040] The reconstruction module is used by the receiving end to reconstruct the synthetic pixel data of the non-critical image region using the masking information, the original pixel data of the key image region, and the descriptive text of the non-critical image region through an image reconstruction model, and to obtain the reconstructed image corresponding to the image to be transmitted based on the synthetic pixel data of the non-critical image region and the original pixel data of the key image region.

[0041] To address the aforementioned technical problems, this application also discloses an electronic device, comprising:

[0042] processor;

[0043] Memory for storing processor-executable instructions;

[0044] The processor is configured to execute the instructions to implement the semantic communication method based on a generative AI large model as described in any one of the claims.

[0045] To address the aforementioned technical problems, this application also discloses a computer-readable storage medium that, when the instructions in the computer-readable storage medium are executed by the processor of a terminal, enables the terminal to execute the semantic communication method based on a generative AI large model as described in any one of the claims.

[0046] Compared with the prior art, this application has the following advantages:

[0047] This application embodiment performs semantic segmentation on the image to be transmitted at the sending end to obtain key image regions and non-key image regions. A lossless transmission strategy is employed for the key image regions, while a text generation model is used to generate concise descriptive text for the non-key image regions. This reduces data volume while preserving the semantic information of the non-key image regions, reducing transmission bandwidth and alleviating transmission pressure. At the receiving end, the non-key image regions are reconstructed based on the descriptive text, generating a reconstructed image that includes the original pixel data of the key image regions and the synthesized pixel data of the non-key image regions. This ensures the overall consistency and visual effect between the reconstructed image and the image to be transmitted. By distinguishing between key and non-key image regions through semantic segmentation, the semantic structure of the image is effectively utilized for image transmission, achieving fine-grained understanding of the content of the image to be transmitted. This allows for accurate identification and processing of semantic information from different components within the image to be transmitted. It is suitable for image transmission tasks involving complex scenes or multiple important objects, and avoids bandwidth waste and information loss due to a lack of fine-grained analysis. This solves the technical problem of limited accuracy and reliability in semantic communication in existing technologies, alleviating bandwidth resource constraints, improving transmission efficiency, and achieving high-fidelity transmission of key image regions and semantically complete transmission of non-key image regions. Attached Figure Description

[0048] Figure 1 A flowchart of a semantic communication method provided in an embodiment of this application is shown;

[0049] Figure 2 An example flowchart of a semantic communication method based on a generative AI large model provided in an embodiment of this application is shown;

[0050] Figure 3 An example flowchart of a semantic communication method based on a generative AI large model provided in another embodiment of this application is shown;

[0051] Figure 4 A flowchart illustrating a method for generating a training dataset for a pre-trained text generation model according to an embodiment of this application is shown.

[0052] Figure 5 A flowchart of a text description small model training method provided in an embodiment of this application is shown;

[0053] Figure 6 An exemplary flowchart of a text description small model training method provided in an embodiment of this application is shown;

[0054] Figure 7 A flowchart illustrating the training process of a text generation model and an image reconstruction model provided in an embodiment of this application is shown.

[0055] Figure 8A structural diagram of a semantic communication device based on a generative AI large model provided in an embodiment of this application is shown.

[0056] Figure 9 A schematic diagram of an electronic device structure provided in an embodiment of this application is shown. Detailed Implementation

[0057] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, this application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0058] With the rapid development of internet and mobile communication technologies, and the widespread adoption of new application scenarios such as high-definition video calls, live streaming, and cloud gaming, the demand for high-quality image transmission is experiencing explosive growth. However, traditional communication systems typically transmit all pixels in real-time without discrimination during image transmission, striving for perfect fidelity reproduction of every pixel. Since image data often contains a large amount of redundant information, and this redundant information has the same data weight as critical information, it cannot be differentiated during transmission, leading to a significant consumption of network bandwidth by low-value data. In this context, adopting indiscriminate, full-pixel, lossless transmission inevitably results in the redundant transmission of a large amount of unnecessary data, causing a surge in data transmission volume, excessive consumption of network bandwidth resources, and a significant decrease in transmission efficiency. Ultimately, this further exacerbates the overall load pressure on multimedia communication systems, making it difficult to meet the actual needs of efficient and reliable transmission.

[0059] To alleviate the load pressure on multimedia communication systems, existing technologies propose reducing data volume through image compression. However, current image compression techniques only focus on image processing (such as texture complexity and color distribution) and lack a semantic importance ranking mechanism. This fails to effectively distinguish between semantically critical regions (such as the main object and core information) and semantically non-critical regions (such as background and semantic details). Consequently, semantically critical regions suffer distortion due to over-compression, while non-semantically critical regions become redundant due to the retention of excessive unnecessary details, failing to effectively optimize transmission efficiency. Ultimately, the transmission weights of semantically critical and non-semantically critical regions in the image are not rationally allocated, limiting the actual effectiveness of compression techniques in alleviating transmission load.

[0060] Another existing technology attempts to extract image features and convert them into semantic vectors for transmission. In the feature encoding stage, after extracting intermediate layer feature maps, a fixed-scale spatial pooling operation is used, and the pooled feature vectors are mapped to fixed-dimensional semantic representations through a fully connected layer. This coarse-grained spatial discretization process smooths out the fine-grained spatial relationships in the image. Furthermore, because the feature encoding process does not establish a visual attention mechanism, the encoder lacks the ability to distinguish the semantic importance of different regions of the image. During transmission, bandwidth resources are evenly distributed, resulting in inefficient use of bandwidth resources and limited improvement in transmission efficiency, making it difficult to meet the requirements of high-precision image transmission.

[0061] Some existing technologies also propose transmission schemes that convert images into text descriptions, transforming visual signals into natural language symbols. However, since text descriptions, as a symbolic representation, can only capture coarse-grained semantic concepts of images and cannot encode pixel spatial continuity and detail levels, they are difficult to fully cover all details and features in an image. This often leads to inconsistencies between the reconstructed image and the original content at the receiving end.

[0062] Furthermore, some existing technologies have attempted to introduce lightweight multimodal models to address the specific needs of multimedia communication systems. However, these solutions suffer from systemic flaws: while these models optimize computational resource consumption, their architecture design and training objectives are not specifically optimized for the low-latency, high-efficiency transmission requirements of communication systems. The decoupling of model functionality from semantic compression requirements makes it difficult to directly embed them into existing communication frameworks, resulting in significant system adaptation barriers.

[0063] Meanwhile, some studies have proposed lightweight multimodal models, but existing models typically require high computational resources, and their architecture and training objectives are not optimized for the specific needs of communication systems such as low latency and high-efficiency transmission. They are decoupled from the semantic compression requirements of semantic communication, making it difficult to apply them directly to existing communication frameworks and resulting in incompatibility with communication systems.

[0064] To address the aforementioned technical challenges, existing technologies have proposed semantic communication systems built upon large AI models (such as semantic encoders in the LAM-SC framework and MLM / LLM components in the multimodal LAM-MSC framework). These systems achieve semantic transmission through techniques like knowledge bases, attention integration, and multimodal alignment. However, limitations in data annotation and fine-grained analysis lead to low semantic extraction accuracy and undifferentiated encoding of key / non-key regions, resulting in low transmission efficiency, uneven image quality, unstable generation results, and high difficulty in model deployment. Furthermore, traditional algorithms fail to adequately adapt to image complexity and communication scenario requirements, and the randomness of noise addition further exacerbates fluctuations in generation quality, ultimately leading to bandwidth waste and decreased system reliability.

[0065] To address the technical problems existing in the prior art, this application provides a semantic communication method based on a generative AI large model. (Refer to...) Figure 1 It includes the following steps:

[0066] S1. The sending end performs semantic segmentation on the image to be transmitted to obtain masking information, the original pixel data of the key image region of the image to be transmitted, and the original pixel data of the non-key image region of the image to be transmitted. The masking information is used to characterize the region that needs to be reconstructed by the receiving end.

[0067] Specifically, semantic segmentation refers to dividing the image to be transmitted into regions with different semantic importance. This can be achieved using deep learning-based semantic segmentation algorithms. Convolutional neural networks extract multi-scale features from the image to be transmitted, generating pixel-by-pixel semantic classification results. These results distinguish between key image regions that require faithful transmission and non-key image regions that need reconstruction, providing a foundation for subsequent differentiated transmission. Key image regions are those that play a decisive role in the core semantics of the image or the user's focus; their original pixel data must be fully preserved. Non-key image regions are redundant regions with low visual importance or those that can be inferred from context. Masking information uses a binary mask matrix encoded with the same size as the image to be transmitted, clearly identifying the pixel positions of the regions to be reconstructed using 0 / 1 labels. This avoids over-fidelity in non-key image regions or distortion in key image regions due to region segmentation errors, improving the reconstruction efficiency of the reconstructed image.

[0068] For example, the key image region of the image to be transmitted refers to the core content such as the target object, face, or text area, while the non-key image region of the image to be transmitted refers to the background or other auxiliary visual elements.

[0069] S2. The sending end uses the original pixel data of the non-critical image region to generate descriptive text for the non-critical image region through a text generation model. The descriptive text for the non-critical image region is used to describe the semantic information of the non-critical image region.

[0070] Specifically, the text generation model refers to converting non-critical image regions into natural language descriptive text. It has image encoding and text decoding capabilities, can understand the semantic information in non-critical image regions, and generate descriptive text corresponding to the image content. The descriptive text retains the core information of non-critical image regions through semantic encoding, expresses the background information of the image through language, reduces the amount of data transmission, and provides semantic support.

[0071] S3. The sending end sends the masking information, the original pixel data of the key image area, and the descriptive text of the non-key image area to the receiving end.

[0072] Specifically, to improve transmission reliability and address channel noise and transmission errors, the transmitting end employs a phased coding strategy when transmitting masking information, raw pixel data of key image regions, and descriptive text of non-key image regions to the receiving end. First, a joint coding algorithm is used to jointly encode the raw pixel data of key image regions and the descriptive text of non-key image regions, generating a low-dimensional semantic vector. This low-dimensional semantic vector achieves efficient data compression and semantic preservation by associating pixel details of the raw pixel data of key image regions with the descriptive text of non-key image regions. Next, channel coding techniques such as Low-Density Parity-Check (LDPC) or Polar codes are used to channel-code the low-dimensional semantic vector, generating codewords for transmission. This coding process optimizes the combination of raw pixel data of key image regions and descriptive text of non-key image regions, thus optimizing the data representation. While reducing transmission redundancy, it retains the core information required for image reconstruction by the receiving end, thereby improving image reconstruction quality while ensuring transmission efficiency.

[0073] S4. The receiving end uses the masking information, the original pixel data of the key image region, and the descriptive text of the non-key image region to reconstruct the synthetic pixel data of the non-key image region through the image reconstruction model, and obtains the reconstructed image corresponding to the image to be transmitted based on the synthetic pixel data of the non-key image region and the original pixel data of the key image region.

[0074] Specifically, the image reconstruction model not only generates visual content based on masking information, original pixel data of key image regions, and descriptive text of non-key image regions, but also exhibits robustness against noise generated during transmission. Through its semantic understanding and contextual association capabilities, the model accurately understands the core semantic information of non-key image regions, effectively correcting transmission errors and ensuring consistency between the synthesized pixel data of reconstructed non-key image regions and their original pixel data. Upon receiving the masking information, original pixel data of key image regions, and descriptive text of non-key image regions, the receiving end first performs channel decoding to reconstruct low-dimensional semantic vectors. Then, through semantic decoding, it recovers the original pixel data of key image regions and the descriptive text of non-key image regions. Next, the image reconstruction model uses the reconstructed descriptive text of non-key image regions as semantic guidance and, combined with the masking information, determines the image regions to be reconstructed. Within the non-key regions identified by the mask, it synthesizes pixel data that matches the descriptive text of the non-key image regions. Finally, the receiving end fuses the original pixel data of the key image regions with the reconstructed pixel data that matches the descriptive text of the non-key image regions to generate the reconstructed image corresponding to the image to be transmitted.

[0075] This application embodiment performs semantic segmentation on the image to be transmitted at the sending end to obtain key image regions and non-key image regions. A lossless transmission strategy is employed for the key image regions, while a text generation model is used to generate concise descriptive text for the non-key image regions. This reduces data volume while preserving the semantic information of the non-key image regions, reducing transmission bandwidth and alleviating transmission pressure. At the receiving end, the non-key image regions are reconstructed based on the descriptive text, generating a reconstructed image that includes the original pixel data of the key image regions and the synthesized pixel data of the non-key image regions. This ensures the overall consistency and visual effect between the reconstructed image and the image to be transmitted. By distinguishing between key and non-key image regions through semantic segmentation, the semantic structure of the image is effectively utilized for image transmission, achieving fine-grained understanding of the content of the image to be transmitted. This allows for accurate identification and processing of semantic information from different components within the image to be transmitted. It is suitable for image transmission tasks involving complex scenes or multiple important objects, and avoids bandwidth waste and information loss due to a lack of fine-grained analysis. This solves the technical problem of limited accuracy and reliability in semantic communication in existing technologies, alleviating bandwidth resource constraints, improving transmission efficiency, and achieving high-fidelity transmission of key image regions and semantically complete transmission of non-key image regions.

[0076] For example, Figure 2 An example flowchart of a semantic communication method based on a generative AI large model provided in an embodiment of this application is shown. Figure 3 An example flowchart of a semantic communication method based on a generative AI large model, according to another embodiment of this application, is shown. (Refer to...) Figure 2 and Figure 3 The source end refers to the sending end, the destination end refers to the receiving end, the original image refers to the image to be transmitted, the restored image refers to the reconstructed image corresponding to the image to be transmitted, the description model refers to the text generation model, the description refers to the descriptive text of non-critical image regions, the semantic segmentation model refers to the model that performs semantic segmentation on the image to be transmitted, and the key part refers to the original pixel data of the key image regions.

[0077] After receiving the original image, the source end's semantic segmentation model first extracts the masking information of the original image. Then, it divides the original image into two regions: a key image region occupying a proportion α of the image to be transmitted, and a non-key image region occupying a proportion 1-α of the image to be transmitted. This yields the original pixel data for both the key image region (occupying a proportion α) and the non-key image region (occupying a proportion 1-α). Here, α is dynamically determined by the semantic segmentation model based on a preset foreground or task requirements, achieving adaptive region division by identifying content in the original image that needs to be transmitted without distortion.

[0078] Next, the description model uses the raw pixel data of non-critical image regions to generate descriptive text for those regions. This process encodes the raw pixel data of non-critical image regions into descriptive text, ensuring the preservation of key semantic features while representing less important details as structured text prompts.

[0079] Subsequently, the raw pixel data of the key image region and the raw pixel data of the non-key image region are respectively processed by the image semantic encoder and the text semantic encoder, and then jointly encoded by the channel encoder to generate a low-dimensional semantic vector. The encoding results of the image semantic encoder and the channel encoder are mapped to a shared latent space to generate codewords for transmission. These codewords are then transmitted to the receiving end via a wireless channel. The receiving end uses the text semantic decoder and the image semantic decoder to recover the raw pixel data of the key image region and the descriptive text of the non-key image region, respectively.

[0080] Subsequently, the image reconstruction model can specifically adopt the Stable Diffusion Inpainting model, using the descriptive text of the recovered non-critical image regions as prompts, and combining the non-critical image regions located with the masking information to perform image reconstruction on the non-critical regions. By fusing the original pixel data of the key image regions with the reconstructed non-critical region images, the final restored image is generated.

[0081] In some embodiments, considering that current generative AI-based semantic communication systems employ large multimodal models for processing, and that the capabilities of these large multimodal models extend far beyond generating descriptive text for non-critical image regions, computational and storage resources are inefficiently utilized. This not only increases hardware deployment costs but also limits the system's scalability in resource-constrained scenarios. The solution involves transferring the capabilities of the trained large text description model to a smaller text description model with fewer parameters, achieving model lightweighting while maintaining semantic description quality, and adapting to the system's requirement for low resource consumption. Specifically, this includes the following steps:

[0082] Before training the text generation model, a base model is first trained to obtain a trained text description model. Then, the trained text description model is used to generate the training dataset for the pre-trained text generation model.

[0083] Figure 4 A flowchart illustrating a method for generating a training dataset for a pre-trained text generation model according to an embodiment of this application is shown. (Refer to...) Figure 4 The training dataset for the pre-trained text generation model is generated according to the following steps:

[0084] S110. The original image is segmented into multiple local blocks, which are then input into the visual encoder of the basic large model to obtain the aggregated local features of the block encoding. The basic large model has image understanding capabilities.

[0085] Specifically, the basic large model has image understanding capabilities, and the visual encoder of the basic large model has image feature extraction capabilities. In the process of extracting the aggregated local features of block encoding, the original image is first divided into multiple local blocks, and then each local block is input into the visual encoder of the basic large model to obtain aggregated local features that represent the texture details of each local block.

[0086] By segmenting the original image into multiple local blocks, the visual encoder of the basic large model can focus on the local texture details of each local block, avoiding the loss of details caused by global scaling. At the same time, inputting multiple local blocks into the visual encoder of the basic large model can reduce the resolution requirement of a single processing, reduce the computational load of the visual encoder, and adapt to high-resolution image processing scenarios.

[0087] S120. Input the original image into the visual encoder of the basic large model to obtain the global features of the original image scaled and encoded.

[0088] Specifically, after the original image is input into the visual encoder of the basic large model, it is downsampled to obtain the global features of the scaled encoding of the original image. The global features extract macroscopic information such as the overall layout and scene category of the original image, providing contextual constraints for subsequent text descriptions. The global features and aggregated local features complement each other, providing data support for cross-scale feature fusion.

[0089] S130. The aggregated local features of the block coding and the global features of the original image scaling coding are fused to obtain the fused features.

[0090] Specifically, fusion features refer to a representation that integrates aggregated local features from block encoding and global features from scaling encoding of the original image through cross-dimensional splicing or attention mechanisms, resulting in a fusion feature vector containing multi-granularity information. By combining aggregated local features from block encoding and global features from scaling encoding of the original image, semantic ambiguity caused by single-scale features is avoided, and global features can correct noise interference in aggregated local features, thus improving feature robustness.

[0091] S140. Based on the fusion features, the text encoder of the basic large model generates descriptive text for the original image.

[0092] Specifically, the text encoder of the base model is used to map the fused features to descriptive text for the original image. By compressing the high-dimensional fused features into low-dimensional text, the core semantics of the original image are preserved, redundant pixel information is filtered out, providing structured training data for subsequent text generation models and establishing a visual-language alignment relationship.

[0093] S150. Based on the difference between the descriptive text of the original image and the correct descriptive text of the original image, keep the parameters of the visual encoder of the basic large model unchanged, update the parameters of the text encoder of the basic large model, and obtain the trained text description large model.

[0094] Specifically, by comparing the difference between the descriptive text of the original image and the correct descriptive text of the original image, the parameters of the text encoder of the base model are updated to minimize the difference between the generated descriptive text and the real description. By keeping the parameters of the visual encoder of the base model unchanged, the ability to extract existing image features is avoided, ensuring the stability of the fused feature-descriptive text mapping; at the same time, updating only the text encoder parameters can reduce the computational resource requirements and accelerate model convergence.

[0095] In an optional embodiment, obtaining the trained large text description model further includes the following steps:

[0096] S150-1. Input the original image into multiple reference small models to obtain multi-level descriptive text of the original image. The multiple reference small models are used for at least one of the following: extracting global semantic information of the original image and generating a preliminary description, refining the object attributes of local image regions of the original image, identifying the text embedded in the original image, and segmenting and associating the spatial location of key image regions of the original image.

[0097] Specifically, multiple reference mini-models refer to a set of specialized models with different semantic extraction capabilities. Specifically, the BLIP-2 model can be used to extract global semantic information from the original image and generate a preliminary description; the GRIT model can refine the object attributes of local image regions in the original image; the PPOCR model can identify embedded text in the original image; and the SAM model can segment key image regions in the original image and associate their spatial locations. These models complement each other to cover different dimensions of the semantics of the original image.

[0098] S150-2. Using the multi-level descriptive text of the original image as the contextual guidance signal, based on the difference between the descriptive text of the original image and the correct descriptive text of the original image, while keeping the parameters of the visual encoder of the basic large model unchanged, update the text encoder of the basic large model to obtain the trained text description large model.

[0099] Specifically, multi-level descriptive text refers to a set of semantic information generated by reference small models with different semantic extraction capabilities. This can be achieved by extracting global semantic information from the original image and generating a preliminary description, refining the object attributes of local image regions in the original image, identifying the text embedded in the original image, and segmenting key image regions of the original image and associating them with spatial combinations. Contextual guidance signals refer to using multi-level descriptive text as a conditional input for descriptive text generation. Specifically, this can be achieved by concatenating the outputs of multiple reference small models into a cue vector during the text encoder training phase, forcing the text encoder to learn a fusion expression of multi-source semantics.

[0100] This approach generates multi-level descriptive text for the original image using multiple reference mini-models, achieving multi-dimensional semantic coverage and improving the comprehensiveness and accuracy of semantic extraction. It avoids semantic gaps caused by the functional limitations of a single model. The multi-level descriptive text serves as a contextual guide signal, forcing the text encoder to learn the fusion expression of global features, local details, text content, and spatial relationships. This ensures that the generated semantic vectors reflect both the core content of the image and retain auxiliary semantic information, providing more accurate semantic guidance for subsequent image reconstruction. Through complementary training of multiple reference mini-models and semantic fusion, the richness and accuracy of the text description are enhanced, reducing the risk of image reconstruction distortion caused by semantic extraction bias and improving image transmission efficiency and image reconstruction quality.

[0101] S160. Using the trained text description model, add labeled text to the original image to obtain the training dataset of the pre-trained text generation model. The labeled text is used to describe the semantic information of non-critical regions of the original image.

[0102] Specifically, by utilizing a pre-trained text description model, annotated text is added to the original images to describe the semantic information of non-critical regions of the original images, generating a training dataset for a pre-trained text generation model with "image-text" pairings. The generalization ability of the pre-trained text description model is then leveraged to generate descriptive text covering multiple scenes and objects, thus generating large-scale training data and reducing the cost of constructing datasets generated by manual annotation.

[0103] By combining local feature extraction with global feature fusion, the original image details are preserved while overall features are taken into account, providing accurate multi-scale visual information for the text encoder to generate descriptive text. By fixing the visual encoder parameters and optimizing the text encoder, computational resource consumption is reduced while ensuring semantic alignment stability. The training dataset for the pre-trained text generation model with "image-text" pairing is automatically generated using the basic large model, breaking through the scale limitations of manual annotation and reducing the cost of training data construction.

[0104] Next, the selected text description mini-model is trained using the training dataset of the pre-trained text generation model to obtain the pre-trained text generation model.

[0105] Figure 5 A flowchart illustrating a text description small model training method according to an embodiment of this application is shown. (Refer to...) Figure 5 Specifically, it includes the following steps:

[0106] S210. Select a text description mini-model. The text description mini-model has the ability to describe the semantic information of the image.

[0107] Specifically, a text description small model refers to a lightweight neural network architecture with a smaller number of parameters than a text description large model. It can be implemented using basic models such as BLIP, and the computational complexity is reduced by simplifying the number of network layers and parameter scale.

[0108] S220. Using the training dataset of the pre-trained text generation model, with the goal of learning the ability to describe the semantic information of images from the trained large text description model, the small text description model is trained to improve the ability of the small text description model to describe the semantic information of images, thus obtaining the pre-trained text generation model.

[0109] In an optional embodiment, S220, the text description mini-model is trained to improve its ability to describe the semantic information of images, resulting in a pre-trained text generation model, including:

[0110] S220-1. Preserve the accuracy of the original model weights of the visual projection layer and the text encoder of the text description mini-model, and quantize the accuracy of the original model weights of the remaining layers of the text description mini-model to obtain the quantized text description mini-model.

[0111] S220-2. Train the quantized text description mini-model to improve its ability to describe the semantic information of images, thus obtaining a pre-trained text generation model.

[0112] Specifically, the visual projection layer refers to the cross-modal alignment module that maps image feature vectors to the text semantic space. This can be implemented using a linear transformation layer or an attention mechanism layer to maintain accurate mapping from image features to the text space. The text encoder refers to the neural network structure that transforms text sequences into semantic vectors. This can be implemented using a transformer-based encoder to ensure the coherence and logic of the semantic description. Quantization refers to the compression technique that converts model parameters from high-precision floating-point numbers to low-precision numbers. This can be implemented using dynamic bit-width quantization or mixed-precision quantization methods to reduce the computational complexity of the model. The remaining layers of the text description model refer to model components that are less sensitive to quantization errors, excluding the visual projection layer and the text encoder. These include the cross-modal fusion layer and the text decoding layer. During quantization, the text encoder and visual projection layer can retain their original FP32 precision to preserve key information, while the precision of the remaining layers is converted to INT4 format using quantization functions from the BitsAndBytes library.

[0113] By balancing performance and resource consumption through hierarchical quantization, the visual projection layer and text encoder retain their original precision, ensuring that the core semantic feature extraction capability is not affected by quantization errors. The remaining layers undergo low-precision quantization, reducing model size and computational overhead by decreasing parameter bit width. The quantized model undergoes fine-tuning of training step size and precision loss through parameter fine-tuning. Knowledge distillation is used to transfer the semantic description capabilities of the trained large text description model to the small text description model. During training, an image-text alignment loss function is used to maintain cross-modal feature matching, combined with a distillation loss function to constrain the consistency between the quantized model output and the original model. This collaborative mechanism of selectively preserving the precision of key layers and compressing non-key layers achieves model lightweighting while maintaining semantic description quality, adapting to the system's requirement for low resource consumption.

[0114] In another optional embodiment, the loss function used to train the text description mini-model includes an image-text alignment loss function and a distillation loss function. The image-text alignment loss function is used to measure the degree of matching between the text description output by the text description mini-model for an image and the semantic information of the image. The distillation loss function is used to measure the degree of consistency between the text description output by the text description mini-model for an image and the text description output by the trained text description large model for the image.

[0115] Specifically, the image-text alignment loss function calculates the matching degree between the generated text and the image feature space through a contrastive learning mechanism. This can be achieved using the cross-modal alignment mechanism of the CLIP model, with the cosine similarity between the text embedding vector and the image embedding vector used as the evaluation metric. The distillation loss function maintains the consistency of the output distribution between the small and large models through knowledge transfer methods. This can be achieved using the KL divergence calculation method, constructing the optimization objective by comparing the probability distribution differences of the output text from the two models.

[0116] During the training phase of the text description mini-model, the image-text alignment loss function, through a cross-modal contrastive learning mechanism, forces the model to establish a precise mapping relationship between local image features and linguistic descriptions. This mechanism matches image patch features with corresponding text fragments as positive samples, while simultaneously forming negative sample pairs with irrelevant text, thereby enhancing the model's ability to capture fine-grained semantics. The distillation loss function, through a temperature-regulated soft-labeling technique, transforms the descriptive text generated by the large model into a probability distribution-based supervisory signal, guiding the text description mini-model to learn the semantic description capabilities of the trained large model. A dynamic weight adjustment strategy is employed during training. For example, in the early stages of training, the distillation loss is emphasized to inherit the capabilities of the trained large model, while the alignment loss weight is gradually increased in the later stages to enhance semantic accuracy. By dynamically adjusting the image-text alignment loss function and the distillation loss function, key accuracy is maintained while compressing the size of the text description mini-model.

[0117] Figure 6 An exemplary flowchart of a text description small model training method according to an embodiment of this application is shown. (Refer to...) Figure 6The basic large model adopts the Tongyi Qianwen LargeVision Language Model (Qwen-VL). The visual encoder of the basic large model is represented by VIT (Vision Transformer). The ontology features refer to the aggregated local features of block encoding. The parameters of the text encoder of the basic large model are updated using Low-Rank Adaptation (LoRA). The integrated vision-text pipeline refers to multiple reference small models, including the Bootstrapping Language-Image Pre-training 2 model, the Grounded Reasoning with Images & Texts model, the Baidu PaddlePaddle Optical Character Recognition model (PPOCR model), the Segment Anything Model (SAM model), and the Generative Pre-trained Transformer (GPT model). The aforementioned reference mini-models are all publicly available through open-source technology websites or training model libraries. The specific parameters of each reference mini-model are set according to actual needs, and will not be elaborated further in this application embodiment.

[0118] In the training process of the text description small model, firstly, the original image is segmented into multiple local blocks (i.e. Figure 6 Inputting A, B, C, D, E, and F into the visual encoder of the base model, each local block is processed using shared weights to obtain the aggregated local features (i.e., ontological features) of the block encoding. Inputting the original image into the visual encoder of the base model yields the global features of the scaled-encoded original image.

[0119] Next, fused features are obtained by fusing the aggregated local features from block encoding and the global features from the scaled encoding of the original image. The text encoder of the base large model then generates descriptive text for the original image.

[0120] Then, the multi-level descriptive text of the original images output by multiple reference small models is used as contextual guidance signals. Based on the difference between the descriptive text of the original images and the correct descriptive text of the original images, the parameters of the visual encoder of the base large model are kept unchanged, and the parameters of the text encoder of the base large model are updated using low-rank adaptation, resulting in a trained text description large model. Using the trained text description large model, labeled text is added to the original images, resulting in a training dataset for a pre-trained text generation model with a sample size of 30K.

[0121] During the process of updating the parameters of the text encoder of the basic large model, features of the original image can be further extracted through multimodal task fusion training. These extracted features are then used as another contextual guidance signal to further ensure the accuracy of the text descriptions generated by the trained large text description model. Specifically, multimodal task fusion training includes at least an image scene text description model and a general document visual question answering model. Both the image scene text description model and the general document visual question answering model are publicly available through open-source technology websites or training model libraries. The specific parameters of the image scene text description model and the general document visual question answering model are set according to actual needs, and will not be elaborated further in this application embodiment.

[0122] Reference Figure 6 The text description mini-model includes at least a text encoder, a visual projection layer, and other layers. The remaining layers of the text description mini-model include a cross-attention layer, a query layer, a value layer, and other layers. Generative matching fine-tuning refers to training the text description mini-model using image-text alignment loss functions (including image-text contrast loss function and image-text matching loss function), while semantic matching fine-tuning refers to training the text description mini-model using distillation loss functions (including cross-entropy loss function and divergence loss function). The divergence loss function refers to the KL divergence loss function (Kullback-Leibler divergence loss function). Image-text contrast loss function, image-text matching loss function, cross-entropy loss function, divergence loss function, and KL divergence loss function are all prior art in this technical field and will not be described further in this application's embodiments.

[0123] Before training the text description mini-model, the FP32 precision of the original model weights of the visual projection layer and text encoder of the text description mini-model is preserved. The precision of the original model weights of other layers of the text description mini-model is quantized so that the model weights of other layers of the text description mini-model are in INT4 format. Here, FP32 and INT4 are both numerical precision formats of the model. FP32 precision is the original high-precision format before quantization, which is used for high-precision numerical calculation to ensure the accuracy of key feature extraction. INT4 is the low-precision format after quantization. By reducing the precision, the model storage and computational overhead is reduced, and the model is lightweighted. Finally, a text description mini-model that balances accuracy and efficiency is obtained after quantization.

[0124] Using a training dataset of 30K pre-trained text generation models generated from a pre-trained large text description model, the image-text alignment loss function (i.e., generative matching fine-tuning) forces the model to establish a precise mapping relationship between local image features and linguistic descriptions through a cross-modal contrastive learning mechanism. This mechanism performs positive sample matching between image patch features and corresponding text fragments, while simultaneously forming negative sample pairs with irrelevant text, thereby enhancing the model's ability to capture fine-grained semantics. The distillation loss function (i.e., semantic matching fine-tuning) uses a temperature-adjusted soft-labeling technique to transform the descriptive text generated by the large model into a probability distribution-based supervision signal, guiding the small text description model to learn the semantic description capabilities of the pre-trained large model. A dynamic weight adjustment strategy is employed during training. For example, in the early stages of training, the distillation loss is emphasized to inherit the capabilities of the pre-trained large model, while the alignment loss weight is gradually increased in the later stages to enhance semantic accuracy. By dynamically adjusting the image-text alignment loss function and the distillation loss function, key accuracy is maintained while compressing the size of the small text description model.

[0125] Finally, the pre-trained text generation model and the pre-trained image reconstruction model are jointly trained. The pre-trained image reconstruction model has basic image reconstruction capabilities. Through joint training, the pre-trained text generation model and the pre-trained image reconstruction model are further fine-tuned to obtain the text generation model and the image reconstruction model. Figure 7 A flowchart illustrating the training process of a text generation model and an image reconstruction model provided in an embodiment of this application is shown. (Refer to...) Figure 7 The training process for text generation models and image reconstruction models includes:

[0126] S10. The sending end performs semantic segmentation on the sample image to obtain sample masking information, original pixel data of the key image region of the sample that accounts for α proportion of the sample image, and original pixel data of the non-key image region of the sample image. The sample masking information is used to characterize the region that needs to be reconstructed by the receiving end.

[0127] Specifically, in training the text generation model and the image reconstruction model, semantic segmentation of the sample images is first required to generate pixel-by-pixel semantic classification results. By identifying regions in the sample images with significantly different visual importance, the sample images are segmented into raw pixel data of key image regions (accounting for α proportion) and raw pixel data of non-key image regions (accounting for 1-α proportion). Sample masking information clearly identifies the locations of non-key image regions that need to be reconstructed.

[0128] By using fine-grained semantic segmentation, the pre-trained text generation model is forced to focus on the extraction of semantic features from non-critical image regions of the samples. This enables the pre-trained text generation model to learn the accurate mapping relationship from visual signals to structured text, while ensuring the lossless transmission of the original pixel data of the key image regions of the samples. This provides high-fidelity benchmark information for the pre-trained image reconstruction model, reduces the risk of reconstruction distortion caused by region segmentation errors from the source, and improves the model's ability to understand the semantics of complex scenes.

[0129] S20. The sending end uses the original pixel data of the non-critical image region of the sample to generate descriptive text for the non-critical image region of the sample through a pre-trained text generation model. The descriptive text for the non-critical image region of the sample is used to describe the semantic information of the non-critical image region of the sample.

[0130] Specifically, a pre-trained text generation model is used to encode the original pixel data of non-critical image regions of the samples into text describing them in natural language, thereby covering the core semantics of the non-critical image regions of the samples.

[0131] By transforming the visual information of non-critical image regions into semantic text, a clear semantic guide is provided for the pre-trained image reconstruction model. This enables the pre-trained image reconstruction model to generate pixel data that is semantically consistent with the sample image based on the descriptive text of the sample non-critical image regions, thus preserving key semantic information for image reconstruction.

[0132] S30. The sending end sends the sample masking information, the original pixel data of the key image region of the sample, and the descriptive text of the non-key image region of the sample to the receiving end.

[0133] Specifically, in the process of sending sample masking information, original pixel data of key image regions of the sample, and descriptive text of non-key image regions of the sample to the receiving end, the original pixel data of key image regions of the sample and the descriptive text of non-key image regions of the sample are first mapped to a shared latent space through a joint coding algorithm to generate a low-dimensional semantic vector; then, the low-dimensional semantic vector is error-corrected and encoded through channel coding to generate transmission codewords adapted to the characteristics of the wireless channel; finally, it is transmitted through the wireless channel.

[0134] By jointly encoding the original pixel data of the key image regions of the samples and the descriptive text of the non-key image regions of the samples, the data representation is optimized, which reduces transmission redundancy while retaining core information and improving transmission efficiency. Channel coding enhances robustness to channel noise, ensures reliable data transmission, and guarantees efficient transmission even in resource-constrained scenarios (low bandwidth, high bit error rate channels). It provides complete and low-distortion input information for the pre-trained image reconstruction model, ensuring the accuracy of the final reconstructed image.

[0135] S40. The receiving end uses the sample masking information, the original pixel data of the key image region of the sample, and the descriptive text of the non-key image region of the sample to reconstruct the synthetic pixel data of the non-key image region of the sample through a pre-trained image reconstruction model, and obtains the sample reconstruction image corresponding to the sample image based on the synthetic pixel data of the non-key image region of the sample and the original pixel data of the key image region of the sample.

[0136] Specifically, after receiving the transmitted data, the receiving end first restores the low-dimensional semantic vector through channel decoding, and then recovers the original pixel data of the key image regions and the descriptive text of the non-key regions through semantic decoding. Subsequently, the pre-trained image reconstruction model uses the descriptive text of the sample's non-key image regions as semantic guidance, combines sample masking information to locate the sample's non-key image regions to be reconstructed, synthesizes pixel data that conforms to the semantic description, and finally fuses the original pixels of the sample's key image regions with the synthesized pixel data of the reconstructed sample's non-key image regions to generate a complete sample reconstructed image.

[0137] By collaboratively reconstructing the original pixel data of key image regions and the descriptive text of non-key image regions, the semantic and visual coherence between the reconstructed image and the original image is ensured. This guarantees reconstruction accuracy for complex backgrounds or low-interest areas.

[0138] S50, The difference between the reconstructed image corresponding to the sample image and the sample image is less than... To achieve this, the model parameters of the pre-trained text generation model and the pre-trained image reconstruction model are updated, resulting in the text generation model and the image reconstruction model, respectively. This is a preset constant.

[0139] Specifically, This refers to the proportion of key image regions in a sample image. This is a preset constant, set according to actual needs; this application embodiment does not impose specific limitations. The difference between the reconstructed image and the original sample image is less than a preset threshold. To achieve this, the parameters of the pre-trained text generation model and the pre-trained image reconstruction model are jointly optimized. During the optimization process, the model simultaneously adjusts the semantic encoding strategy (such as attention weight allocation) of the pre-trained text generation model and various model parameters of the pre-trained image reconstruction model until the difference between the reconstructed image and the original sample image meets the transmission quality requirements.

[0140] Through joint optimization, a close semantic alignment relationship is established between the pre-trained text generation model and the pre-trained image reconstruction model in the latent space. This enables text descriptions to more accurately guide image reconstruction, reduces the perceptual difference between the reconstructed sample image and the sample image, and reduces the amount of data transmission while ensuring high fidelity of the core content, thereby improving transmission efficiency.

[0141] Referring to the same inventive concept, embodiments of this application also provide a semantic communication device based on a generative AI large model, referring to... Figure 8 Semantic communication devices based on generative AI large models include:

[0142] The segmentation module 100 is used by the transmitting end to perform semantic segmentation on the image to be transmitted, and to obtain masking information, the original pixel data of the key image region of the image to be transmitted and the original pixel data of the non-key image region of the image to be transmitted. The masking information is used to characterize the region that needs to be reconstructed by the receiving end.

[0143] The generation module 200 is used by the sending end to generate descriptive text for the non-critical image regions using the original pixel data of the non-critical image regions through a text generation model. The descriptive text for the non-critical image regions is used to describe the semantic information of the non-critical image regions.

[0144] The sending module 300 is used to send the masking information, the original pixel data of the key image area, and the descriptive text of the non-key image area to the receiving end.

[0145] The reconstruction module 400 is used by the receiving end to reconstruct the synthetic pixel data of the non-critical image regions using masking information, the original pixel data of the key image regions, and the descriptive text of the non-critical image regions through an image reconstruction model, and to obtain the reconstructed image corresponding to the image to be transmitted based on the synthetic pixel data of the non-critical image regions and the original pixel data of the key image regions.

[0146] In some embodiments, the semantic communication device based on a generative AI large model further includes:

[0147] The sample segmentation module is used by the sending end to perform semantic segmentation on the sample image, and obtain sample masking information, original pixel data of the key image region of the sample that accounts for α proportion of the sample image, and original pixel data of the non-key image region of the sample image. The sample masking information is used to characterize the region that needs to be reconstructed by the receiving end.

[0148] The sample generation module is used by the sending end to generate descriptive text for the non-critical image regions of the sample using the original pixel data of the non-critical image regions and a pre-trained text generation model. The descriptive text for the non-critical image regions of the sample is used to describe the semantic information of the non-critical image regions of the sample.

[0149] The sample sending module is used by the sending end to send sample masking information, raw pixel data of key image regions of the sample, and descriptive text of non-key image regions of the sample to the receiving end.

[0150] The sample reconstruction module is used by the receiving end to reconstruct the synthetic pixel data of the non-critical image regions of the sample using sample masking information, the original pixel data of the key image regions of the sample, and the descriptive text of the non-critical image regions of the sample through a pre-trained image reconstruction model. Based on the synthetic pixel data of the non-critical image regions of the sample and the original pixel data of the key image regions of the sample, the sample reconstruction image corresponding to the sample image is obtained.

[0151] The parameter update module is used to reconstruct the image from the sample image, where the difference between the reconstructed image and the sample image is less than [a certain value]. To achieve this, the model parameters of the pre-trained text generation model and the pre-trained image reconstruction model are updated, resulting in the text generation model and the image reconstruction model, respectively. This is a preset constant.

[0152] As the system implementation is basically similar to the method implementation, it is described in a relatively simple way. For relevant details, please refer to the description of the method implementation.

[0153] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0154] Figure 9 A schematic diagram of an electronic device structure according to an embodiment of this application is shown. (Refer to...) Figure 9 This application also provides another electronic device, including:

[0155] processor;

[0156] Memory is used to store processor-executable instructions;

[0157] The processor is configured to execute instructions to implement any semantic communication method based on a generative AI large model.

[0158] In this embodiment, the computer device includes a processor, memory, and network interface connected via a system bus.

[0159] The computer device's processor provides computational and control capabilities. Its memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The computer device's database stores data samples. Its network interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements any semantic communication method based on a generative AI large model.

[0160] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0161] This application also provides a computer-readable storage medium that, when the instructions in the computer-readable storage medium are executed by the processor of a terminal, enables the terminal to execute any semantic communication method based on a generative AI large model.

[0162] The aforementioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0163] Optionally, a readable storage medium can be coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Alternatively, the readable storage medium can be an integral part of the processor. Both the processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components within the device.

[0164] This application also provides a computer program product, which includes a computer program that, when executed by a processor, implements any semantic communication method based on a generative AI large model.

[0165] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0166] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0167] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0168] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0169] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0170] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

[0171] The semantic communication method, apparatus, and device based on generative AI large models provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A semantic communication method based on a generative AI large model, characterized in that, include: The sending end performs semantic segmentation on the image to be transmitted to obtain masking information, the original pixel data of the key image region of the image to be transmitted, and the original pixel data of the non-key image region of the image to be transmitted. The masking information is used to characterize the region that needs to be reconstructed by the receiving end. The sending end uses the original pixel data of the non-critical image region to generate descriptive text for the non-critical image region through a text generation model. The descriptive text for the non-critical image region is used to describe the semantic information of the non-critical image region. The sending end sends the masking information, the original pixel data of the key image region, and the descriptive text of the non-key image region to the receiving end; The receiving end uses the masking information, the original pixel data of the key image region, and the descriptive text of the non-key image region to reconstruct the synthetic pixel data of the non-key image region through an image reconstruction model, and obtains the reconstructed image corresponding to the image to be transmitted based on the synthetic pixel data of the non-key image region and the original pixel data of the key image region.

2. The method as described in claim 1, characterized in that, The training process for the text generation model and the image reconstruction model includes: The transmitting end performs semantic segmentation on the sample image to obtain sample masking information, original pixel data of the key image region of the sample that accounts for α proportion of the sample image, and original pixel data of the non-key image region of the sample image. The sample masking information is used to characterize the region that the receiving end needs to reconstruct the image. The sending end uses the original pixel data of the non-critical image region of the sample to generate descriptive text of the non-critical image region of the sample through a pre-trained text generation model. The descriptive text of the non-critical image region of the sample is used to describe the semantic information of the non-critical image region of the sample. The sending end sends the sample masking information, the original pixel data of the key image region of the sample, and the descriptive text of the non-key image region of the sample to the receiving end. The receiving end uses the sample masking information, the original pixel data of the key image region of the sample, and the descriptive text of the non-key image region of the sample to reconstruct the synthetic pixel data of the non-key image region of the sample through a pre-trained image reconstruction model, and obtains the sample reconstruction image corresponding to the sample image based on the synthetic pixel data of the non-key image region of the sample and the original pixel data of the key image region of the sample. The difference between the reconstructed image corresponding to the sample image and the sample image is less than [a certain value]. To achieve this, the model parameters of the pre-trained text generation model and the pre-trained image reconstruction model are updated to obtain the text generation model and the image reconstruction model, wherein, This is a preset constant.

3. The method as described in claim 2, characterized in that, The training dataset for the pre-trained text generation model was generated according to the following steps: The original image is segmented into multiple local blocks, which are then input into the visual encoder of the base model to obtain the aggregated local features encoded in blocks. The base model has image understanding capabilities. The original image is input into the visual encoder of the basic large model to obtain the global features of the original image scaled and encoded. The aggregated local features of the block encoding and the global features of the original image scaling encoding are fused to obtain the fused features; Based on the fusion features, the text encoder of the basic large model generates the descriptive text of the original image; Based on the difference between the descriptive text of the original image and the correct descriptive text of the original image, keep the parameters of the visual encoder of the basic large model unchanged, update the parameters of the text encoder of the basic large model, and obtain the trained text description large model. Using the trained text description model, labeled text is added to the original image to obtain the training dataset of the pre-trained text generation model. The labeled text is used to describe the semantic information of non-critical regions of the original image.

4. The method as described in claim 3, characterized in that, Also includes: The original image is input into multiple reference small models to obtain multi-level descriptive text of the original image. The multiple reference small models are used for at least one of the following: extracting global semantic information of the original image and generating a preliminary description, refining the object attributes of local image regions of the original image, identifying the text embedded in the original image, and segmenting and associating the spatial location of key image regions of the original image. Based on the difference between the descriptive text of the original image and the correct descriptive text of the original image, while keeping the parameters of the visual encoder of the basic large model unchanged, the parameters of the text encoder of the basic large model are updated to obtain the trained text description large model, including: Using the multi-level descriptive text of the original image as a contextual guidance signal, and based on the difference between the descriptive text of the original image and the correct descriptive text of the original image, while keeping the parameters of the visual encoder of the basic large model unchanged, the text encoder of the basic large model is updated to obtain the trained text description large model.

5. The method as described in claim 3, characterized in that, Also includes: A text description mini-model is selected, which has the ability to describe the semantic information of the image; Using the training dataset of the pre-trained text generation model, with the goal of learning the ability to describe the semantic information of images from the trained large text description model, the small text description model is trained to improve the ability of the small text description model to describe the semantic information of images, thus obtaining the pre-trained text generation model.

6. The method as described in claim 5, characterized in that, Also includes: The accuracy of the original model weights of the visual projection layer and the text encoder of the text description mini-model is preserved, and the accuracy of the original model weights of the remaining layers of the text description mini-model is quantized to obtain the quantized text description mini-model. The text description mini-model is trained to improve its ability to describe the semantic information of images, resulting in the pre-trained text generation model, comprising: The quantized text description mini-model is trained to improve its ability to describe the semantic information of images, thus obtaining the pre-trained text generation model.

7. The method as described in claim 6, characterized in that, The loss functions used to train the text description mini-model include the image-text alignment loss function and the distillation loss function. The image-text alignment loss function is used to measure the degree of matching between the text description output by the text description mini-model for an image and the semantic information of the image. The distillation loss function is used to measure the degree of consistency between the text description output by the small text description model for the image and the text description output by the trained large text description model for the image.

8. A semantic communication device based on a generative AI large model, characterized in that, include: The segmentation module is used by the sending end to perform semantic segmentation on the image to be transmitted, and obtain masking information, original pixel data of key image regions of the image to be transmitted, and original pixel data of non-key image regions of the image to be transmitted. The masking information is used to characterize the regions that need to be reconstructed by the receiving end. The generation module is used by the sending end to generate descriptive text for the non-critical image region using the original pixel data of the non-critical image region and a text generation model. The descriptive text for the non-critical image region is used to describe the semantic information of the non-critical image region. The sending module is used by the sending end to send the masking information, the original pixel data of the key image region, and the descriptive text of the non-key image region to the receiving end; The reconstruction module is used by the receiving end to reconstruct the synthetic pixel data of the non-critical image region using the masking information, the original pixel data of the key image region, and the descriptive text of the non-critical image region through an image reconstruction model, and to obtain the reconstructed image corresponding to the image to be transmitted based on the synthetic pixel data of the non-critical image region and the original pixel data of the key image region.

9. An electronic device, characterized in that, include: processor; Memory for storing processor-executable instructions; The processor is configured to execute the instructions to implement the semantic communication method based on a generative AI large model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the terminal, the terminal is able to perform the semantic communication method based on a generative AI large model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image processing method and device, computer, storage medium and program product

    CN117252947A

  • Medical visual question and answer method and system based on multi-task modeling

    CN119202334A