Semantic communication method, device and equipment based on generative AI large model and storage medium
By using a generative AI large model for semantic segmentation and text generation, combined with channel coding technology, the accuracy and reliability problems of semantic communication in existing technologies are solved, and efficient and reliable image transmission is achieved, which is suitable for image transmission tasks in complex scenarios.
Patent Information
- Application Number
- CN202511141469.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-15
AI Technical Summary
The accuracy and reliability of semantic communication in existing technologies are limited. Traditional communication systems cannot effectively distinguish the semantic importance of different areas of an image, resulting in tight bandwidth resources and reduced transmission efficiency.
Semantic segmentation is performed through a generative AI large model to distinguish between key and non-key areas of the image, transmit the key areas without distortion, generate concise description text for non-key areas, combine channel coding technology for data compression and transmission, and restore the image through an image reconstruction model.
It achieves high efficiency, reliability and accuracy in image transmission, reduces bandwidth usage, ensures the overall consistency and visual effect of the reconstructed image, and is suitable for image transmission tasks in complex scenes.
Smart Images

Figure CN120725928A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a semantic communication method, apparatus, device and storage medium based on a generative AI big model. Background Art
[0002] With the rapid development of the internet and mobile communications, the demand for high-quality image transmission has increased dramatically. Traditional communication systems generally use an indiscriminate transmission mode, striving for distortion-free reproduction of every pixel. However, due to the natural differences in visual importance and information carrying capacity of different image regions, a large amount of redundant data is inevitably generated, leading to bandwidth constraints and reduced transmission efficiency.
[0003] To alleviate bandwidth constraints and improve transmission efficiency, existing technologies have proposed semantic communication methods based on large AI models. These methods extract and transmit semantic information from the image to be transmitted. The receiver, based on a knowledge base, fuses the received semantic information with prior knowledge to reconstruct the image to be transmitted. However, because the knowledge base relies on limited training data, it struggles to fully cover all semantic scenarios. This leads to discrepancies between the reconstructed image and the transmitted image, reducing the accuracy and reliability of semantic communication. Summary of the Invention
[0004] The present application provides a semantic communication method, apparatus, device and storage medium based on a generative AI big model to solve the problem of limited accuracy and reliability of semantic communication in the prior art.
[0005] To solve the above problems, this application discloses a semantic communication method based on a generative AI large model, comprising the following steps: The transmitting end performs semantic segmentation on the image to be transmitted to obtain mask information, original pixel data of the key image area of the image to be transmitted, and original pixel data of the non-key image area of the image to be transmitted, wherein the mask information is used to represent the area where the receiving end needs to perform image reconstruction; The sending end generates a description text of the non-key image area using the original pixel data of the non-key image area through a text generation model, wherein the description text of the non-key image area is used to describe the semantic information of the non-key image area; The sending end sends the mask information, the original pixel data of the key image area, and the description text of the non-key image area to the receiving end; The receiving end uses the mask information, the original pixel data of the key image area, and the description text of the non-key image area to reconstruct the synthetic pixel data of the non-key image area through an image reconstruction model, and obtains the reconstructed image corresponding to the image to be transmitted based on the synthetic pixel data of the non-key image area and the original pixel data of the key image area.
[0006] Furthermore, the training process of the text generation model and the image reconstruction model includes: The sending end performs semantic segmentation on the sample image to obtain sample mask information, original pixel data of a sample key image area occupying an α ratio of the sample image, and original pixel data of a sample non-key image area of the sample image, wherein the sample mask information is used to represent: an area where the receiving end needs to perform image reconstruction; The sending end generates a description text of the sample non-key image area using the original pixel data of the sample non-key image area through a pre-trained text generation model, wherein the description text of the sample non-key image area is used to describe the semantic information of the sample non-key image area; The sending end sends the sample mask information, the original pixel data of the sample key image area, and the description text of the sample non-key image area to the receiving end; The receiving end uses the sample mask information, the original pixel data of the sample key image area, and the description text of the sample non-key image area to reconstruct the synthesized pixel data of the sample non-key image area through a pre-trained image reconstruction model, and obtains a sample reconstructed image corresponding to the sample image based on the synthesized pixel data of the sample non-key image area and the original pixel data of the sample key image area; The difference between the sample reconstructed image corresponding to the sample image and the sample image is less than As the target, the model parameters of the pre-trained text generation model and the pre-trained image reconstruction model are updated to obtain the text generation model and the image reconstruction model, wherein, is a preset constant.
[0007] Furthermore, the training dataset of the pre-trained text generation model is generated according to the following steps: The original image is divided into multiple local blocks, which are respectively input into the visual encoder of the basic large model to obtain aggregated local features of the block encoding. The basic large model has image understanding capabilities; Inputting the original image into the visual encoder of the basic large model to obtain the global features of the scaled encoding of the original image; Fusing the aggregated local features of block coding and the global features of original image scaling coding to obtain fused features; Based on the fused features, a description text of the original image is generated through a text encoder of the basic large model; Based on the difference between the description text of the original image and the correct description text of the original image, keep the parameters of the visual encoder of the basic large model unchanged, update the parameters of the text encoder of the basic large model, and obtain a trained text description large model; The trained text description model is used to add annotated text to the original image to obtain a training data set for the pre-trained text generation model, where the annotated text is used to describe semantic information of non-critical areas of the original image.
[0008] Furthermore, it also includes: Inputting the original image into a plurality of reference miniature models to obtain multi-level description text of the original image, wherein the plurality of reference miniature models are used for at least one of the following: extracting global semantic information of the original image and generating a preliminary description, refining object attributes of local image regions of the original image, identifying text embedded in the original image, and segmenting key image regions of the original image and associating spatial positions; Based on the difference between the description text of the original image and the correct description text of the original image, keeping the parameters of the visual encoder of the basic large model unchanged, updating the parameters of the text encoder of the basic large model, and obtaining a trained text description large model, including: Taking the multi-level description text of the original image as the context guidance signal, based on the difference between the description text of the original image and the correct description text of the original image, the parameters of the visual encoder of the basic large model are kept unchanged, the text encoder of the basic large model is updated, and a trained text description large model is obtained.
[0009] Furthermore, it also includes: Selecting a text description model, wherein the text description model has the ability to describe semantic information of the image; Using the training data set of the pre-trained text generation model, the text description small model is trained with the goal of learning the ability to describe the semantic information of the image from the trained text description large model, thereby improving the ability of the text description small model to describe the semantic information of the image, and obtaining the pre-trained text generation model.
[0010] Furthermore, it also includes: retaining the accuracy of the original model weights of the visual projection layer and the text encoder of the text description small model, and quantizing the accuracy of the original model weights of the remaining layers of the text description small model to obtain a quantized text description small model; Training the text description model to improve its ability to describe semantic information of images, thereby obtaining the pre-trained text generation model, includes: The quantized text description small model is trained to improve the ability of the quantized text description small model to describe the semantic information of the image, thereby obtaining the pre-trained text generation model.
[0011] Furthermore, the loss function used to train the text description model includes an image-text alignment loss function and a distillation loss function. The image-text alignment loss function is used to measure the degree of match between the text description output by the text description model for an image and the semantic information of the image. The distillation loss function is used to measure the degree of consistency between the text description output by the small text description model for an image and the text description output by the trained large text description model for the image.
[0012] In order to solve the above technical problems, the present application also discloses a semantic communication device based on a generative AI large model, comprising: A segmentation module is configured to perform semantic segmentation of the image to be transmitted at the transmitting end to obtain mask information, raw pixel data of a key image area of the image to be transmitted, and raw pixel data of a non-key image area of the image to be transmitted, wherein the mask information is used to represent an area requiring image reconstruction at the receiving end; A generating module, configured for the sending end to generate a description text of the non-key image area using the original pixel data of the non-key image area through a text generation model, wherein the description text of the non-key image area is used to describe the semantic information of the non-key image area; A sending module, configured for the sending end to send the mask information, the original pixel data of the key image area, and the description text of the non-key image area to the receiving end; A reconstruction module is used for the receiving end to use the mask information, the original pixel data of the key image area, and the description text of the non-key image area to reconstruct the synthetic pixel data of the non-key image area through an image reconstruction model, and obtain the reconstructed image corresponding to the image to be transmitted based on the synthetic pixel data of the non-key image area and the original pixel data of the key image area.
[0013] In order to solve the above technical problems, the present application also discloses an electronic device, comprising: processor; a memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement any one of the semantic communication methods based on the generative AI big model.
[0014] In order to solve the above technical problems, the present application also discloses a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by the processor of the terminal, the terminal is able to execute the semantic communication method based on the generative AI big model described in any one of the claims.
[0015] Compared with the prior art, this application has the following advantages: The embodiment of the present application performs semantic segmentation of the image to be transmitted on the transmitting end to obtain key image areas and non-key image areas, adopts a lossless transmission strategy for the key image areas, and uses a text generation model to generate concise description text for the non-key image areas, thereby reducing the amount of data while retaining the semantic information of the non-key image areas, reducing the transmission bandwidth, and alleviating the transmission pressure; the receiving end reconstructs the non-key image areas based on the description text of the non-key image areas, and generates a reconstructed image including the original pixel data of the key image areas and the synthesized pixel data of the non-key image areas, thereby ensuring the overall consistency and visual effect of the reconstructed image with the image to be transmitted; the key image areas and non-key image areas are distinguished by semantic segmentation, and the semantic structure of the image is effectively utilized for image transmission, thereby achieving a fine-grained understanding of the content of the image to be transmitted, accurately identifying and processing the semantic information of different components in the image to be transmitted, and being suitable for image transmission tasks containing complex scenes or multiple important objects, and avoiding bandwidth waste and information loss caused by lack of fine-grainedness. The invention solves the technical problem of limited accuracy and reliability of semantic communication in the prior art, which not only relieves bandwidth resources and improves transmission efficiency, but also achieves efficient transmission of high fidelity in key image areas and complete semantics in non-key image areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A flowchart of a semantic communication method provided by an embodiment of the present application is shown; Figure 2 An example flow chart of a semantic communication method based on a generative AI big model provided by an embodiment of the present application is shown; Figure 3 An example flow chart of a semantic communication method based on a generative AI big model provided by another embodiment of the present application is shown; Figure 4 A flow chart of a method for generating a training data set for a pre-trained text generation model provided in one embodiment of the present application is shown; Figure 5A flowchart of a text description small model training method provided in one embodiment of the present application is shown; Figure 6 An exemplary flow chart of a text description small model training method provided by an embodiment of the present application is shown; Figure 7 A flowchart of a method for training a text generation model and an image reconstruction model provided in an embodiment of the present application is shown; Figure 8 The figure shows a structure diagram of a semantic communication device based on a generative AI large model provided by an embodiment of the present application; Figure 9 A schematic structural diagram of an electronic device provided in one embodiment of the present application is shown. DETAILED DESCRIPTION
[0017] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0018] With the rapid development of the internet and mobile communication technologies, and the increasing popularity of new application scenarios such as high-definition video calls, live streaming, and cloud gaming, the demand for high-quality image transmission is exploding. However, traditional communication systems typically transmit all pixels in real time, without distinction, and strive for complete fidelity reproduction of all image pixels. Because image data often contains a large amount of redundant information, and this redundant information has the same data weight as key information, it cannot be treated differently during transmission, resulting in a large amount of network bandwidth being consumed by low-value data. In this context, adopting indiscriminate, full-pixel, lossless transmission inevitably leads to the redundant transmission of a large amount of unnecessary data, which in turn causes a surge in data transmission volume, excessive network bandwidth utilization, and a significant decrease in transmission efficiency. Ultimately, this further increases the overall load on multimedia communication systems, making it difficult to meet the actual requirements for efficient and reliable transmission.
[0019] To alleviate the load pressure on multimedia communication systems, existing technologies have proposed reducing data volume through image compression technology. However, existing image compression technology only stays at the image processing level (such as texture complexity and color distribution), and has not established a semantic importance grading mechanism. It is unable to effectively distinguish between semantically critical areas (such as the main object, core information) and semantically non-critical areas (such as background, semantic details) in the image. As a result, semantically critical areas are distorted due to excessive compression, while non-semantically critical areas are redundant due to retaining too many unnecessary details, failing to achieve effective optimization of transmission efficiency. Ultimately, the transmission weights of semantically critical areas and non-semantically critical areas in the image are not reasonably distributed, limiting the actual effect of compression technology on alleviating transmission load.
[0020] Another existing technology attempts to extract image features and convert them into semantic vectors for transmission. During the feature encoding stage, after extracting the intermediate layer feature maps, a fixed-scale spatial pooling operation is performed, and the pooled feature vectors are mapped into fixed-dimensional semantic representations through a fully connected layer. This coarse-grained spatial discretization process smooths out fine-grained spatial relationships in the image. Furthermore, because the feature encoding process lacks a visual attention mechanism, the encoder lacks the ability to discern the semantic importance of different image regions. Consequently, bandwidth resources are evenly distributed during transmission, resulting in inefficient bandwidth utilization and limited transmission efficiency improvements, making it difficult to meet the requirements of high-precision image transmission.
[0021] Some existing technologies also propose transmission schemes that convert images into text descriptions, translating visual signals into natural language symbols. However, as a symbolic representation, text descriptions can only capture coarse-grained semantic concepts of an image and cannot encode pixel spatial continuity and levels of detail. This makes it difficult to fully capture all details and features in an image. Consequently, when reconstructing the image at the receiving end, it often results in inconsistencies between the image and the original content.
[0022] Furthermore, some existing technologies attempt to introduce lightweight multimodal models to address the specific needs of multimedia communication systems. However, these approaches suffer from systemic flaws: while these models optimize computing resource usage, their architecture and training objectives are not specifically optimized for the low-latency, high-efficiency transmission requirements of communication systems. The decoupling of model functionality from semantic compression requirements makes them difficult to directly embed into existing communication frameworks, presenting significant system adaptation barriers.
[0023] At the same time, some studies have also proposed lightweight multimodal models, but the existing ones usually have high computing resource requirements, and their architecture and training objectives are not optimized for specific requirements of communication systems such as low latency and high-efficiency transmission. They are also decoupled from the semantic compression requirements of semantic communication, making them difficult to directly apply to existing communication frameworks and causing incompatibility with communication systems.
[0024] To address the aforementioned technical issues, existing technologies have proposed semantic communication systems based on large AI models (such as the semantic encoder of the LAM-SC framework and the MLM / LLM components of the multimodal LAM-MSC framework). These systems achieve semantic transmission through technologies such as knowledge bases, attention integration, and multimodal alignment. However, these systems are limited by insufficient data annotation and the lack of fine-grained analysis, resulting in low semantic extraction accuracy and non-differentiated encoding of key / non-key areas. This in turn leads to problems such as low transmission efficiency, uneven image quality, unstable generation results, and difficulty in model deployment. Furthermore, traditional algorithms are not fully adapted to the complexity of images and the requirements of communication scenarios. The randomness of noise further exacerbates fluctuations in generation quality, ultimately wasting bandwidth and reducing system reliability.
[0025] In order to solve the technical problems existing in the above-mentioned prior art, the embodiment of the present application provides a semantic communication method based on a generative AI large model. Figure 1 , including the following steps: S1. The sending end performs semantic segmentation on the image to be transmitted to obtain mask information, original pixel data of the key image area of the image to be transmitted, and original pixel data of the non-key image area of the image to be transmitted. The mask information is used to represent: the area where the receiving end needs to perform image reconstruction.
[0026] Specifically, semantic segmentation refers to dividing the image to be transmitted into regions with different semantic importance. This can be achieved by using a semantic segmentation algorithm based on deep learning. The multi-scale features of the image to be transmitted are extracted through a convolutional neural network, and pixel-by-pixel semantic classification results are generated to distinguish between key image areas that need to be transmitted with fidelity and non-key image areas that need to be reconstructed, providing a basis for subsequent differentiated transmission. Key image areas refer to areas that play a decisive role in the core semantics of the image or the user's focus, and their original pixel data needs to be fully preserved; non-key image areas refer to redundant areas with low visual importance or that can be inferred from the context; the mask information is encoded using a binary mask matrix that is consistent with the size of the image to be transmitted, and the pixel positions of the areas that need to be reconstructed are clearly identified through 0 / 1 labels to avoid excessive fidelity in non-key image areas or distortion in key image areas due to errors in area division, thereby improving the reconstruction efficiency of the reconstructed image.
[0027] Exemplarily, the key image area of the image to be transmitted refers to core content such as a target object, a face or a text area, and the non-key image area of the image to be transmitted refers to the background part or other auxiliary visual elements.
[0028] S2. The sending end uses the original pixel data of the non-key image area to generate a description text of the non-key image area through a text generation model. The description text of the non-key image area is used to describe the semantic information of the non-key image area.
[0029] Specifically, the text generation model refers to the process of converting non-critical image areas into text described in natural language. It has the ability to encode images and decode text, understand the semantic information in non-critical image areas, and generate descriptive text corresponding to the image content. The descriptive text retains the core information of non-critical image areas through semantic encoding, expresses image background information through language, reduces the amount of data transmission, and provides semantic support.
[0030] S3. The sending end sends the mask information, the original pixel data of the key image area, and the description text of the non-key image area to the receiving end.
[0031] Specifically, to improve transmission reliability and mitigate channel noise and transmission errors, the transmitter adopts a phased encoding strategy when transmitting mask information, raw pixel data of key image areas, and descriptive text of non-key image areas to the receiver. First, a joint encoding algorithm is used to jointly encode the raw pixel data of key image areas and the descriptive text of non-key image areas to generate a low-dimensional semantic vector. This low-dimensional semantic vector achieves efficient data compression and semantic preservation by associating pixel details of the raw pixel data of key image areas with the descriptive text of non-key image areas. Next, channel coding techniques such as low-density parity-check codes (LDPC) or polar codes are used to channel encode the low-dimensional semantic vector to generate codewords for transmission. This encoding process optimizes the combination of raw pixel data of key image areas and descriptive text of non-key image areas, optimizing data representation. This reduces transmission redundancy while retaining the core information required for image reconstruction at the receiver, thereby ensuring transmission efficiency and improving image reconstruction quality.
[0032] S4. The receiving end uses the mask information, the original pixel data of the key image area, and the descriptive text of the non-key image area to reconstruct the synthetic pixel data of the non-key image area through the image reconstruction model, and obtains the reconstructed image corresponding to the image to be transmitted based on the synthetic pixel data of the non-key image area and the original pixel data of the key image area.
[0033] Specifically, the image reconstruction model not only generates visual content based on the mask information, the original pixel data of the key image area, and the descriptive text of the non-key image area, but is also robust to noise generated during transmission. Through the image reconstruction model's semantic understanding and contextual association capabilities, it accurately understands the core semantic information of the non-key image area, effectively corrects transmission errors, and ensures that the synthesized pixel data of the reconstructed non-key image area is consistent with the original pixel data of the non-key image area. After receiving the mask information, the original pixel data of the key image area, and the descriptive text of the non-key image area, the receiver first performs channel decoding to restore the low-dimensional semantic vector. Then, through semantic decoding, the original pixel data of the key image area and the descriptive text of the non-key image area are restored. Next, the image reconstruction model uses the restored descriptive text of the non-key image area as a semantic guide and combines it with the mask information to determine the image area to be reconstructed. Pixel data consistent with the descriptive text of the non-key image area is synthesized within the non-key area identified by the mask. Finally, the receiver fuses the original pixel data of the key image area with the reconstructed pixel data consistent with the descriptive text of the non-key image area to generate a reconstructed image corresponding to the image to be transmitted.
[0034] The embodiment of the present application performs semantic segmentation of the image to be transmitted on the transmitting end to obtain key image areas and non-key image areas, adopts a lossless transmission strategy for the key image areas, and uses a text generation model to generate concise description text for the non-key image areas, thereby reducing the amount of data while retaining the semantic information of the non-key image areas, reducing the transmission bandwidth, and alleviating the transmission pressure; the receiving end reconstructs the non-key image areas based on the description text of the non-key image areas, and generates a reconstructed image including the original pixel data of the key image areas and the synthesized pixel data of the non-key image areas, thereby ensuring the overall consistency and visual effect of the reconstructed image with the image to be transmitted; the key image areas and non-key image areas are distinguished by semantic segmentation, and the semantic structure of the image is effectively utilized for image transmission, thereby achieving a fine-grained understanding of the content of the image to be transmitted, accurately identifying and processing the semantic information of different components in the image to be transmitted, and being suitable for image transmission tasks containing complex scenes or multiple important objects, and avoiding bandwidth waste and information loss caused by lack of fine-grainedness. The invention solves the technical problem of limited accuracy and reliability of semantic communication in the prior art, which not only relieves bandwidth resources and improves transmission efficiency, but also achieves efficient transmission of high fidelity in key image areas and complete semantics in non-key image areas.
[0035] For example, Figure 2 An example flowchart of a semantic communication method based on a generative AI big model provided in one embodiment of the present application is shown. Figure 3 An example flow chart of a semantic communication method based on a generative AI big model provided by another embodiment of the present application is shown. Figure 2 and Figure 3 , the source end refers to the sending end, the destination end refers to the receiving end, the original image refers to the image to be transmitted, the restored image refers to the reconstructed image corresponding to the image to be transmitted, the description model refers to the text generation model, the description refers to the description text of the non-critical image area, the semantic segmentation model refers to the model for semantic segmentation of the image to be transmitted, and the key part refers to the original pixel data of the key image area.
[0036] After receiving the original image at the source, the semantic segmentation model first extracts the mask information from the original image and then divides the original image into two regions: a key image region (a proportion of the image to be transmitted) and a non-key image region (a proportion of the image to be transmitted - 1 - α). The original pixel data for the key image region (a proportion of the image to be transmitted) and the non-key image region (a proportion of the image to be transmitted - 1 - α) are obtained. The α value is dynamically determined by the semantic segmentation model based on the preset foreground or task requirements. Adaptive region division is achieved by identifying the content in the original image that needs to be transmitted without distortion.
[0037] Next, the description model uses the raw pixel data of non-critical image regions to generate description text for non-critical image regions. This process encodes the raw pixel data of non-critical image regions into description text, ensuring that key semantic features are preserved while representing unimportant details as structured text cues.
[0038] The original pixel data of the key image area and the original pixel data of the non-key image area are then passed through the image semantic encoder and text semantic encoder, respectively, and then jointly encoded by the channel encoder to generate a low-dimensional semantic vector. The encoding results of the image semantic encoder and channel encoder are mapped into a shared latent space. After generating a codeword for transmission, the codeword is sent to the receiver via a wireless channel. The receiver recovers the original pixel data of the key image area and the descriptive text of the non-key image area through the text semantic decoder and image semantic decoder, respectively.
[0039] Subsequently, the image reconstruction model can specifically adopt a local repainting (Stable Diffusion Inpainting) model, using the descriptive text of the restored non-critical image area as a prompt word, and combining the non-critical image area located by mask information to reconstruct the image of the non-critical area. By fusing the original pixel data retained in the critical image area with the reconstructed non-critical area image, the final restored image is generated.
[0040] In some embodiments, considering that the current semantic communication system based on generative AI uses a large multimodal model for processing, and the current large multimodal model is capable of far more than just generating descriptive text for non-critical image areas, resulting in inefficient use of computing resources and storage resources, which not only increases the cost of hardware deployment, but also limits the system's ability to be applied on a large scale in resource-constrained scenarios. The capabilities of the trained large text description model are migrated to a small text description model with a low parameter level, achieving model lightweighting while maintaining the quality of semantic description, and adapting to the system's demand for low resource usage. Specifically, the following steps are included: Before training the text generation model, the basic large model is first trained to obtain a trained text description large model, and then the trained text description large model is used to generate a training data set for the pre-trained text generation model.
[0041] Figure 4 The flowchart of the method for generating a training data set for a pre-trained text generation model provided by an embodiment of the present application is shown. Figure 4 ,The training dataset of the pre-trained text generation model is generated by following the steps below: S110, dividing the original image into multiple local blocks, inputting each of the blocks into the visual encoder of the basic large model, and obtaining aggregated local features of the block encoding, so that the basic large model has the ability to understand the image.
[0042] Specifically, the basic large model has the ability to understand images, and the visual encoder of the basic large model has the ability to extract image features. In the process of extracting the aggregated local features of block coding, the original image is first divided into multiple local blocks, and each local block is input into the visual encoder of the basic large model respectively to obtain the aggregated local features that characterize the texture details of each local block.
[0043] By dividing the original image into multiple local blocks, the visual encoder of the basic large model can focus on the local texture details of each local block, avoiding the loss of details caused by global scaling; at the same time, multiple local blocks are input into the visual encoder of the basic large model respectively, which can reduce the resolution requirement of a single processing, reduce the computational load of the visual encoder, and adapt to high-resolution image processing scenarios.
[0044] S120: Input the original image into the visual encoder of the basic large model to obtain the global features of the scaled encoding of the original image.
[0045] Specifically, after the original image is input into the visual encoder of the underlying large model, it is downsampled to obtain the global features of the original image scaled encoding. Global features extract macro information such as the overall layout of the original image and the scene category, providing contextual constraints for subsequent text descriptions. Global features complement aggregated local features and provide data support for cross-scale feature fusion.
[0046] S130 , fusing the aggregated local features of the block coding and the global features of the original image scaling coding to obtain a fused feature.
[0047] Specifically, fused features are representations that combine the aggregated local features of block coding and the global features of original image scale coding through cross-dimensional splicing or attention mechanisms. This creates a fused feature vector containing multi-granularity information. By combining the aggregated local features of block coding and the global features of original image scale coding, semantic ambiguity caused by single-scale features is avoided. Furthermore, global features can correct for noise interference in aggregated local features, improving feature robustness.
[0048] S140: Based on the fused features, a description text of the original image is generated through the text encoder of the basic large model.
[0049] Specifically, the text encoder in the base large model is used to map the fused features into text describing the original image. By compressing the high-dimensional fused features into low-dimensional text, the core semantics of the original image are preserved, redundant pixel information is filtered out, and structured training data is provided for the subsequent text generation model, establishing a visual-language alignment relationship.
[0050] S150. Based on the difference between the description text of the original image and the correct description text of the original image, the parameters of the visual encoder of the basic large model are kept unchanged, and the parameters of the text encoder of the basic large model are updated to obtain a trained text description large model.
[0051] Specifically, by comparing the original image description with the correct description, the parameters of the text encoder of the base model are updated to minimize the difference between the generated description and the true description. By keeping the parameters of the visual encoder of the base model unchanged, the existing image feature extraction capabilities are not damaged, ensuring the stability of the fused feature-to-description mapping. Furthermore, updating only the text encoder parameters reduces computing resource requirements and accelerates model convergence.
[0052] In an optional embodiment, obtaining the trained text description model further includes the following steps: S150-1. Input the original image into multiple reference small models to obtain multi-level description text of the original image. The multiple reference small models are used for at least one of the following: extracting the global semantic information of the original image and generating a preliminary description, refining the object attributes of the local image area of the original image, identifying the text embedded in the original image, and segmenting the key image areas of the original image and associating them with spatial positions.
[0053] Specifically, multiple reference mini-models refer to a collection of specialized models with different semantic extraction capabilities. Specifically, the BLIP-2 model extracts global semantic information from the original image and generates a preliminary description; the GRIT model refines object attributes in local image regions; the PPOCR model recognizes embedded text in the original image; and the SAM model segments key image regions in the original image and associates their spatial locations. These models complement each other in covering different dimensions of the original image semantics.
[0054] S150-2. Using the multi-level description text of the original image as the context guidance signal, based on the difference between the description text of the original image and the correct description text of the original image, the parameters of the visual encoder of the basic large model are kept unchanged, and the text encoder of the basic large model is updated to obtain a trained text description large model.
[0055] Specifically, multi-level description text refers to a set of semantic information generated by reference small models with different semantic extraction capabilities. It can be achieved by extracting the global semantic information of the original image and generating a preliminary description, refining the object attributes of the local image area of the original image, identifying the text embedded in the original image, and segmenting the key image areas of the original image and associating the spatial positions. The context guidance signal refers to the use of multi-level description text as a conditional input for description text generation. Specifically, it can be achieved by splicing the outputs of multiple reference small models into a prompt vector during the text encoder training stage to force the text encoder to learn the fusion expression of multi-source semantics.
[0056] Multiple reference small models are used to generate multi-level description text of the original image to achieve multi-dimensional semantic coverage, improve the comprehensiveness and accuracy of semantic extraction, and avoid semantic loss caused by functional limitations of a single model. The multi-level description text is used as a context guidance signal to force the text encoder to learn the fusion expression of global features and local details, text content and spatial relationships, so that the generated semantic vector can not only reflect the core content of the image, but also retain auxiliary semantic information, providing more accurate semantic guidance for subsequent image reconstruction. Through the complementarity and semantic fusion training of multiple reference small models, the richness and accuracy of text description are enhanced, the risk of image reconstruction distortion caused by semantic extraction deviation is reduced, and the image transmission efficiency and image reconstruction quality are improved.
[0057] S160: Using the trained text description model, add annotation text to the original image to obtain a training data set for the pre-trained text generation model. The annotation text is used to describe the semantic information of non-critical areas of the original image.
[0058] Specifically, the trained large-scale text description model is used to annotate original images with text describing the semantic information of non-critical areas of the original image, generating a training dataset for a pre-trained text generation model for "image-text" pairing. Leveraging the generalization capabilities of the trained large-scale text description model, descriptive text covering multiple scenes and objects is generated to generate large-scale training data, reducing the dataset construction costs associated with manual annotation.
[0059] By extracting local features in blocks and fusing them with global features, the detailed features of the original image are retained while taking into account the overall features, providing accurate multi-scale visual information for the text encoder to generate descriptive text; by fixing the visual encoder parameters and optimizing the text encoder, the consumption of computing resources is reduced while ensuring the stability of semantic alignment; the basic large model is used to automatically generate a training data set for the pre-trained text generation model of "image-text" pairing, breaking through the scale limitations of manual annotation and reducing the cost of building training data.
[0060] Next, the selected text description small model is trained using the training data set of the pre-trained text generation model to obtain a pre-trained text generation model.
[0061] Figure 5 The flowchart of the text description small model training method provided by an embodiment of the present application is shown. Figure 5 , specifically including the following steps: S210: Select a small text description model, where the small text description model has the ability to describe the semantic information of the image.
[0062] Specifically, the small text description model refers to a lightweight neural network architecture with a lower parameter level than the large text description model. It can be implemented using basic models such as BLIP, reducing computational complexity by streamlining the number of network layers and parameter scale.
[0063] S220. Using the training data set of the pre-trained text generation model, with the goal of learning the ability to describe the semantic information of the image from the trained text description large model, the text description small model is trained to improve the ability of the text description small model to describe the semantic information of the image, and a pre-trained text generation model is obtained.
[0064] In an optional embodiment, S220, training the text description model to improve the ability of the text description model to describe the semantic information of the image, and obtaining a pre-trained text generation model, includes: S220-1. Retain the accuracy of the original model weights of the visual projection layer and the text encoder of the text description small model, and quantize the accuracy of the original model weights of the remaining layers of the text description small model to obtain a quantized text description small model.
[0065] S220-2. Train the quantized text description model to improve the ability of the quantized text description model to describe the semantic information of the image, and obtain a pre-trained text generation model.
[0066] Specifically, the visual projection layer refers to a cross-modal alignment module that maps image feature vectors to text semantic space. It can be implemented using a linear transformation layer or an attention mechanism layer to maintain accurate mapping of image features to text space. The text encoder refers to a neural network structure that converts text sequences into semantic vectors. It can be implemented using a transformer-based encoder to ensure the consistency and logic of the semantic description. Quantization refers to a compression technology that converts model parameters from high-precision floating-point numbers to low-precision numbers. It can be implemented using dynamic bit width quantization or mixed precision quantization methods to reduce the computational complexity of the model. The remaining layers of the text description model refer to model components other than the visual projection layer and text encoder that are less sensitive to quantization errors, specifically including the cross-modal fusion layer and the text decoding layer. During the quantization process, the text encoder and visual projection layers can retain the original FP32 precision to retain key information, and the precision of the remaining layers can be converted to INT4 format by applying the quantization function in the BitsAndBytes library.
[0067] Through layered quantization, performance and resource usage are balanced. The visual projection layer and text encoder retain their original accuracy, ensuring that the core semantic feature extraction capability is not affected by quantization errors. The remaining layers are processed with low-precision quantization, reducing the model size and computational overhead by reducing the parameter bit width. The quantized model fine-tunes the training step precision loss through parameter fine-tuning, and uses knowledge distillation technology to transfer the semantic description capabilities of the trained large text description model to the small text description model. During training, the image-text alignment loss function is used to maintain cross-modal feature matching, combined with the distillation loss function to constrain the consistency of the quantized model output with the original model. This synergistic mechanism of selectively retaining the accuracy of key layers and compressing non-key layers achieves model lightweightness while maintaining the quality of semantic descriptions, adapting to the system's demand for low resource usage.
[0068] In another optional embodiment, the loss function used to train the text description small model includes an image-text alignment loss function and a distillation loss function. The image-text alignment loss function is used to measure the degree of matching between the text description output by the text description small model for an image and the semantic information of the image; the distillation loss function is used to measure the degree of consistency between the text description output by the text description small model for an image and the text description output by the trained text description large model for the image.
[0069] Specifically, the image-text alignment loss function calculates the degree of match between the generated text and the image feature space through a contrastive learning mechanism. This can be achieved using the cross-modal alignment mechanism of the CLIP model, using the cosine similarity between the text embedding vector and the image embedding vector as the evaluation metric. The distillation loss function maintains the consistency of the output distribution of the small model and the large model through knowledge transfer methods. This can be achieved using the KL divergence calculation method, which constructs the optimization target by comparing the probability distribution differences of the text output by the two models.
[0070] During the training phase of the small text description model, the image-text alignment loss function, through a cross-modal contrastive learning mechanism, forces the model to establish a precise mapping between local image features and linguistic descriptions. This mechanism matches image patch features with corresponding text fragments as positive samples, while also forming negative pairs with unrelated text, thereby enhancing the model's ability to capture fine-grained semantics. The distillation loss function, through temperature-adjusted soft labeling, converts the description text generated by the large model into a supervisory signal in the form of a probability distribution, guiding the small text description model to learn the semantic description capabilities of the trained large text description model. A dynamic weight adjustment strategy is employed during training. For example, in the early stages of training, the distillation loss is emphasized to inherit the capabilities of the trained large text description model. Later, the alignment loss is gradually weighted to enhance semantic accuracy. By dynamically adjusting the image-text alignment loss and distillation loss functions, the size of the small text description model is compressed while maintaining critical accuracy.
[0071] Figure 6 An exemplary flow chart of a text description small model training method provided by an embodiment of the present application is shown. Figure 6The basic large model adopts the Tongyi Qianwen Large Vision Language Model (Qwen-VL), the visual encoder of the basic large model is represented by VIT (Vision Transformer), the ontology feature refers to the aggregated local features of block encoding, the parameters of the text encoder of the basic large model are updated by low-rank adaptation (LoRA), and the integrated visual-text pipeline refers to multiple reference small models, including the Bootstrapping Language-Image Pre-training Model (BLIP-2 Model), the Grounded Reasoning with Images&Texts Model (GRIT Model), the Baidu PaddlePaddle Optical Character Recognition Model (PPOCR Model), the Segment Anything Model (SAM Model) and the Generative Pre-trained Transformer (GPT). Among them, the above-mentioned multiple reference small models are all disclosed through open source technology websites or training model libraries. The specific parameters of each reference small model are set according to actual needs and will not be repeated in the embodiments of this application.
[0072] In the text description small model training process, first, the original image is divided into multiple local blocks (i.e. Figure 6 In the example, A, B, C, D, E, and F are fed into the visual encoder of the large base model. Each local block is processed using shared weights to obtain aggregated local features (i.e., ontological features) encoded in each block. Feeding the original image into the visual encoder of the large base model yields global features encoded in the scaled form of the original image.
[0073] Next, the fused features are obtained by fusing the aggregated local features of the block coding with the global features of the original image scaling coding. The text description of the original image is generated through the text encoder of the basic large model.
[0074] Then, the multi-level description text of the original image output by multiple reference small models is used as a contextual guidance signal. Based on the difference between the description text of the original image and the correct description text of the original image, the parameters of the visual encoder of the basic large model are kept unchanged, and the parameters of the text encoder of the basic large model are updated using low-rank adaptation to obtain the trained text description large model. Using the trained text description large model, annotated text is added to the original image, resulting in a training dataset with a sample size of 30K for the pre-trained text generation model.
[0075] In the process of updating the parameters of the text encoder of the basic large model, the features of the original image can be further extracted through multimodal task fusion training, and the extracted features can be used as another context guidance signal to further ensure the accuracy of the text description generated by the trained text description large model. Specifically, the multimodal task fusion training includes at least an image scene text description model and a general document visual question answering model. The above-mentioned image scene text description model and general document visual question answering model are disclosed through open source technology web pages or training model libraries. The specific parameters of the image scene text description model and the general document visual question answering model are set according to actual needs and will not be repeated in the embodiments of this application.
[0076] Reference Figure 6 , the text description small model includes at least a text encoder, a visual projection layer and other layers. The remaining layers of the text description small model include a cross-attention layer, a query layer, a value layer and other layers. Generative matching fine-tuning refers to training the text description small model using the image-text alignment loss function (including the image-text contrast loss function and the image-text matching loss function), and semantic matching fine-tuning refers to training the text description small model using the distillation loss function (including the cross entropy loss function and the divergence loss function), wherein the divergence loss function refers to the KL divergence loss function (Kullback-Leibler divergence loss function). The image-text contrast loss function, the image-text matching loss function, the cross entropy loss function, the divergence loss function and the KL divergence loss function are all existing technologies in this technical field and will not be repeated in the embodiments of this application.
[0077] Before training the text description small model, the FP32 precision of the original model weights of the visual projection layer and text encoder of the text description small model is first retained, and the precision of the original model weights of other layers of the text description small model is quantized, so that the model weights of other layers of the text description small model are in INT4 format. Among them, FP32 and INT4 are both numerical precision formats of the model. FP32 precision is the original high-precision format before quantization, which is used for high-precision numerical calculations to ensure the accuracy of key feature extraction; INT4 is the low-precision format after quantization. By reducing the precision, the model storage and computing overhead are reduced, and the model is lightweight. Finally, a text description small model that takes into account both accuracy and efficiency after quantization is obtained.
[0078] When training the small text description model using a training dataset of 30K pre-trained text generation models generated from a trained large text description model, the image-text alignment loss (i.e., generative matching fine-tuning) uses a cross-modal contrastive learning mechanism to force the model to establish a precise mapping between local image features and linguistic descriptions. This mechanism matches image patch features with corresponding text snippets for positive examples, while also creating negative pairs with unrelated text, thereby enhancing the model's ability to capture fine-grained semantics. The distillation loss (i.e., semantic matching fine-tuning) uses a temperature-adjusted soft labeling technique to convert the descriptions generated by the large model into a supervisory signal in the form of a probability distribution, guiding the small text description model to learn the semantic description capabilities of the trained large text description model. A dynamic weight adjustment strategy is employed during training. For example, the distillation loss is initially emphasized to inherit the capabilities of the trained large text description model, while the alignment loss is gradually weighted to enhance semantic accuracy. By dynamically adjusting the image-text alignment loss and distillation loss, the small text description model is compressed while maintaining critical accuracy.
[0079] Finally, the pre-trained text generation model and the pre-trained image reconstruction model are jointly trained. The pre-trained image reconstruction model has basic image reconstruction capabilities. Through joint training, the pre-trained text generation model and the pre-trained image reconstruction model are further fine-tuned to obtain the text generation model and image reconstruction model. Figure 7 The flowchart of the method for training the text generation model and image reconstruction model provided by an embodiment of the present application is shown. Figure 7 ,The training process of the text generation model and image reconstruction model includes: S10. The sending end performs semantic segmentation on the sample image to obtain sample mask information, original pixel data of the sample key image area that occupies the α ratio of the sample image, and original pixel data of the sample non-key image area of the sample image. The sample mask information is used to represent: the area where the receiving end needs to perform image reconstruction.
[0080] Specifically, during the training of the text generation model and image reconstruction model, semantic segmentation of the sample image is first required to generate pixel-by-pixel semantic classification results. By identifying regions of significant visual importance within the sample image, the sample image is segmented into raw pixel data for the sample key image region, which accounts for a proportion α of the sample image, and raw pixel data for the sample non-key image region, which accounts for a proportion 1-α of the sample image. The sample mask information clearly identifies the locations of the sample non-key image regions that require image reconstruction.
[0081] Through fine-grained semantic segmentation, the pre-trained text generation model is forced to focus on the semantic feature extraction of non-critical image areas of the sample, so that the pre-trained text generation model can learn the precise mapping relationship from visual signals to structured text. At the same time, it ensures the distortion-free transmission of the original pixel data of the key image areas of the sample, providing high-fidelity benchmark information for the pre-trained image reconstruction model, reducing the risk of reconstruction distortion caused by area division errors from the source, and improving the model's semantic understanding ability of complex scenes.
[0082] S20. The sending end uses the original pixel data of the sample non-key image area to generate a description text of the sample non-key image area through a pre-trained text generation model. The description text of the sample non-key image area is used to describe the semantic information of the sample non-key image area.
[0083] Specifically, through a pre-trained text generation model, the original pixel data of the non-key image area of the sample is encoded into text described in natural language to cover the core semantics of the non-key image area of the sample.
[0084] By converting the visual information of non-critical image areas into semantic text, clear semantic guidance is provided for the pre-trained image reconstruction model, so that the pre-trained image reconstruction model generates pixel data consistent with the semantics of the sample image based on the descriptive text of the sample non-critical image areas, retaining key semantic information for image reconstruction.
[0085] S30: The sending end sends the sample mask information, the original pixel data of the sample key image area, and the description text of the sample non-key image area to the receiving end.
[0086] Specifically, in the process of sending the sample mask information, the original pixel data of the sample key image area, and the descriptive text of the sample non-key image area to the receiving end, the original pixel data of the sample key image area and the descriptive text of the sample non-key area are first mapped to a shared latent space through a joint coding algorithm to generate a low-dimensional semantic vector; then the low-dimensional semantic vector is error-corrected and encoded through channel coding to generate a transmission codeword that adapts to the characteristics of the wireless channel; finally, it is sent through the wireless channel.
[0087] By jointly encoding the original pixel data of the key image areas of the sample and the descriptive text of the non-key image areas of the sample, the data representation is optimized, the core information is retained while reducing transmission redundancy, and the transmission efficiency is improved; channel coding enhances the robustness to channel noise, ensures reliable data transmission, and ensures efficient transmission even in resource-constrained scenarios (low bandwidth, high bit error rate channels), providing complete and low-distortion input information for the pre-trained image reconstruction model, and ensuring the accuracy of the final reconstructed image.
[0088] S40. The receiving end uses the sample mask information, the original pixel data of the sample key image area, and the description text of the sample non-key image area to reconstruct the synthetic pixel data of the sample non-key image area through a pre-trained image reconstruction model, and obtains the sample reconstructed image corresponding to the sample image based on the synthetic pixel data of the sample non-key image area and the original pixel data of the sample key image area.
[0089] Specifically, after receiving the transmitted data, the receiver first restores the low-dimensional semantic vector through channel decoding, and then recovers the original pixel data of the key image area and the descriptive text of the non-key area through semantic decoding. Subsequently, the pre-trained image reconstruction model uses the descriptive text of the sample non-key image area as a semantic guide, combines the sample mask information to locate the sample non-key image area to be reconstructed, and synthesizes pixel data that conforms to the semantic description. Finally, the original pixels of the sample key image area are fused with the synthesized pixel data of the reconstructed sample non-key image area to generate a complete sample reconstructed image.
[0090] By collaboratively reconstructing the original pixel data of the sample's key image area and the descriptive text of the sample's non-key image area, the reconstructed image of the sample is ensured to be both semantically and visually coherent with the sample image, ensuring the reconstruction accuracy of complex backgrounds or low-attention areas.
[0091] S50, the difference between the sample reconstructed image corresponding to the sample image and the sample image is less than As the goal, the model parameters of the pre-trained text generation model and the pre-trained image reconstruction model are updated to obtain the text generation model and the image reconstruction model, where is a preset constant.
[0092] Specifically, Refers to the proportion of the sample key image area in the sample image, The difference between the sample reconstructed image and the original sample image is less than the preset threshold. To achieve this goal, the parameters of a pre-trained text generation model and a pre-trained image reconstruction model are jointly optimized. During the optimization process, the model simultaneously adjusts the semantic encoding strategy (such as attention weight distribution) of the pre-trained text generation model and various model parameters of the pre-trained image reconstruction model until the difference between the sample reconstructed image and the sample image corresponding to the sample image meets the transmission quality requirements.
[0093] Through joint optimization, the pre-trained text generation model and the pre-trained image reconstruction model establish a close semantic alignment relationship in the latent space, so that the text description can more accurately guide image reconstruction, reduce the perceptual difference between the sample reconstructed image and the sample image, and reduce the data transmission volume and improve the transmission efficiency while ensuring the high fidelity of the core content.
[0094] Referring to the same inventive concept, the embodiment of the present application also provides a semantic communication device based on a generative AI large model, referring to Figure 8 , semantic communication devices based on generative AI big models include: The segmentation module 100 is used to perform semantic segmentation of the image to be transmitted at the transmitting end to obtain mask information, raw pixel data of the key image area of the image to be transmitted, and raw pixel data of the non-key image area of the image to be transmitted. The mask information is used to represent the area where the receiving end needs to perform image reconstruction; A generating module 200 is configured to generate, at the sending end, a description text of the non-key image region using the original pixel data of the non-key image region through a text generation model, wherein the description text of the non-key image region is used to describe the semantic information of the non-key image region; The sending module 300 is used for the sending end to send the mask information, the original pixel data of the key image area, and the description text of the non-key image area to the receiving end; The reconstruction module 400 is used for the receiving end to use the mask information, the original pixel data of the key image area, and the description text of the non-key image area to reconstruct the synthetic pixel data of the non-key image area through the image reconstruction model, and obtain the reconstructed image corresponding to the image to be transmitted based on the synthetic pixel data of the non-key image area and the original pixel data of the key image area.
[0095] In some embodiments, the semantic communication device based on the generative AI big model further includes: The sample segmentation module is used to perform semantic segmentation on the sample image at the sending end to obtain sample mask information, raw pixel data of the sample key image area that accounts for the α ratio of the sample image, and raw pixel data of the sample non-key image area of the sample image. The sample mask information is used to represent: the area where the receiving end needs to perform image reconstruction.
[0096] The sample generation module is used to generate description text of the sample non-key image area by using the original pixel data of the sample non-key image area at the sending end through a pre-trained text generation model. The description text of the sample non-key image area is used to describe the semantic information of the sample non-key image area.
[0097] The sample sending module is used for sending the sample mask information, the original pixel data of the sample key image area, and the description text of the sample non-key image area to the receiving end.
[0098] The sample reconstruction module is used at the receiving end to use the sample mask information, the original pixel data of the sample key image area, and the description text of the sample non-key image area to reconstruct the synthetic pixel data of the sample non-key image area through a pre-trained image reconstruction model, and obtain the sample reconstructed image corresponding to the sample image based on the synthetic pixel data of the sample non-key image area and the original pixel data of the sample key image area.
[0099] The parameter updating module is used to reconstruct the sample image corresponding to the sample image and the difference between the sample image and the sample image is less than As the goal, the model parameters of the pre-trained text generation model and the pre-trained image reconstruction model are updated to obtain the text generation model and the image reconstruction model, where is a preset constant.
[0100] As for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0101] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0102] Figure 9 FIG1 shows a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 9 , an embodiment of the present application also provides another electronic device, including: processor; a memory for storing processor-executable instructions; The processor is configured to execute instructions to implement any semantic communication method based on a generative AI big model.
[0103] In this embodiment, the computer device includes a processor, a memory, and a network interface connected via a system bus.
[0104] The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data samples. The network interface of the computer device is used to communicate with an external terminal via a network connection. When executed by the processor, the computer program implements any one of the semantic communication methods based on the generative AI large model.
[0105] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0106] An embodiment of the present application also provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by the processor of the terminal, the terminal is enabled to execute any semantic communication method based on the generative AI big model.
[0107] The computer-readable storage medium mentioned above may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic storage, flash memory, magnetic disk, or optical disk. The computer-readable storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0108] Optionally, a readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.
[0109] An embodiment of the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements any semantic communication method based on a generative AI big model.
[0110] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0111] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0112] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0113] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0114] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0115] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
[0116] The above is a detailed introduction to the semantic communication method, device and equipment based on the generative AI big model provided by this application. Specific examples are used in this article to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this application.
Claims
1. A semantic communication method based on a generative AI large model, characterized in that: include: The transmitting end performs semantic segmentation on the image to be transmitted to obtain mask information, original pixel data of the key image area of the image to be transmitted, and original pixel data of the non-key image area of the image to be transmitted, wherein the mask information is used to represent the area where the receiving end needs to perform image reconstruction; The sending end generates a description text of the non-key image area using the original pixel data of the non-key image area through a text generation model, wherein the description text of the non-key image area is used to describe the semantic information of the non-key image area; The sending end sends the mask information, the original pixel data of the key image area, and the description text of the non-key image area to the receiving end; The receiving end uses the mask information, the original pixel data of the key image area, and the description text of the non-key image area to reconstruct the synthetic pixel data of the non-key image area through an image reconstruction model, and obtains the reconstructed image corresponding to the image to be transmitted based on the synthetic pixel data of the non-key image area and the original pixel data of the key image area.
2. The method according to claim 1, wherein The training process of the text generation model and the image reconstruction model includes: The sending end performs semantic segmentation on the sample image to obtain sample mask information, original pixel data of a sample key image area occupying an α ratio of the sample image, and original pixel data of a sample non-key image area of the sample image, wherein the sample mask information is used to represent: an area where the receiving end needs to perform image reconstruction; The sending end generates a description text of the sample non-key image area using the original pixel data of the sample non-key image area through a pre-trained text generation model, wherein the description text of the sample non-key image area is used to describe the semantic information of the sample non-key image area; The sending end sends the sample mask information, the original pixel data of the sample key image area, and the description text of the sample non-key image area to the receiving end; The receiving end uses the sample mask information, the original pixel data of the sample key image area, and the description text of the sample non-key image area to reconstruct the synthesized pixel data of the sample non-key image area through a pre-trained image reconstruction model, and obtains a sample reconstructed image corresponding to the sample image based on the synthesized pixel data of the sample non-key image area and the original pixel data of the sample key image area; The difference between the sample reconstructed image corresponding to the sample image and the sample image is less than As the target, the model parameters of the pre-trained text generation model and the pre-trained image reconstruction model are updated to obtain the text generation model and the image reconstruction model, wherein, is a preset constant.
3. The method according to claim 2, wherein The training dataset of the pre-trained text generation model is generated according to the following steps: The original image is divided into multiple local blocks, which are respectively input into the visual encoder of the basic large model to obtain aggregated local features of the block encoding. The basic large model has image understanding capabilities; Inputting the original image into the visual encoder of the basic large model to obtain global features of the scaled encoding of the original image; Fusing the aggregated local features of the block coding and the global features of the original image scaling coding to obtain a fused feature; Based on the fused features, a description text of the original image is generated through a text encoder of the basic large model; Based on the difference between the description text of the original image and the correct description text of the original image, keep the parameters of the visual encoder of the basic large model unchanged, update the parameters of the text encoder of the basic large model, and obtain a trained text description large model; The trained text description model is used to add annotated text to the original image to obtain a training data set for the pre-trained text generation model, where the annotated text is used to describe semantic information of non-critical areas of the original image.
4. The method according to claim 3, wherein Also includes: Inputting the original image into a plurality of reference miniature models to obtain multi-level description text of the original image, wherein the plurality of reference miniature models are used for at least one of the following: extracting global semantic information of the original image and generating a preliminary description, refining object attributes of local image regions of the original image, identifying text embedded in the original image, and segmenting key image regions of the original image and associating them with spatial positions; Based on the difference between the description text of the original image and the correct description text of the original image, keeping the parameters of the visual encoder of the basic large model unchanged, updating the parameters of the text encoder of the basic large model, and obtaining a trained text description large model, including: Taking the multi-level description text of the original image as the context guidance signal, based on the difference between the description text of the original image and the correct description text of the original image, the parameters of the visual encoder of the basic large model are kept unchanged, the text encoder of the basic large model is updated, and a trained text description large model is obtained.
5. The method according to claim 3, wherein Also includes: Selecting a text description model, wherein the text description model has the ability to describe semantic information of the image; Using the training data set of the pre-trained text generation model, the text description small model is trained with the goal of learning the ability to describe the semantic information of the image from the trained text description large model, thereby improving the ability of the text description small model to describe the semantic information of the image, and obtaining the pre-trained text generation model.
6. The method according to claim 5, wherein Also includes: retaining the accuracy of the original model weights of the visual projection layer and the text encoder of the text description small model, and quantizing the accuracy of the original model weights of the remaining layers of the text description small model to obtain a quantized text description small model; Training the text description model to improve its ability to describe semantic information of images, thereby obtaining the pre-trained text generation model, includes: The quantized text description small model is trained to improve the ability of the quantized text description small model to describe the semantic information of the image, thereby obtaining the pre-trained text generation model.
7. The method according to claim 6, wherein The loss function used to train the text description model includes an image-text alignment loss function and a distillation loss function. The image-text alignment loss function is used to measure the degree of match between the text description output by the text description model for an image and the semantic information of the image. The distillation loss function is used to measure the degree of consistency between the text description output by the small text description model for an image and the text description output by the trained large text description model for the image.
8. A semantic communication device based on a generative AI large model, characterized in that: include: A segmentation module is configured to perform semantic segmentation of the image to be transmitted at the transmitting end to obtain mask information, raw pixel data of a key image area of the image to be transmitted, and raw pixel data of a non-key image area of the image to be transmitted, wherein the mask information is used to represent an area requiring image reconstruction at the receiving end; A generating module, configured for the sending end to generate a description text of the non-key image area using the original pixel data of the non-key image area through a text generation model, wherein the description text of the non-key image area is used to describe the semantic information of the non-key image area; A sending module, configured for the sending end to send the mask information, the original pixel data of the key image area, and the description text of the non-key image area to the receiving end; A reconstruction module is used for the receiving end to use the mask information, the original pixel data of the key image area, and the description text of the non-key image area to reconstruct the synthetic pixel data of the non-key image area through an image reconstruction model, and obtain the reconstructed image corresponding to the image to be transmitted based on the synthetic pixel data of the non-key image area and the original pixel data of the key image area.
9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the semantic communication method based on the generative AI big model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by the processor of the terminal, the terminal is enabled to execute the semantic communication method based on the generative AI big model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Image registration method combining target detection and semantic segmentation
CN110097584A
Image processing method and device, computer, storage medium and program product
CN117252947A
Real-time AI image background dynamic generation method, medium and system
CN118097076A
Medical visual question and answer method and system based on multi-task modeling
CN119202334A
Semantic segmentation model training method and device
CN119919669A
Cited By
Multi-modal feature alignment method and device for heterogeneous data and medium
CN121527790A