Image transmission method and system based on meta learning

By dynamically updating the pretrained semantic feature extraction model and image reconstruction model on the small sample dataset, the problem of poor overfitting and generalization capabilities of image semantic communication technology in a small sample environment is solved, and higher quality image reconstruction is achieved.

CN120223907APending Publication Date: 2025-06-27BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510305787.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing image semantic communication technologies are prone to poor overfitting and generalization capabilities in small sample environments, which affects the semantic fidelity and detail accuracy of image reconstruction.

Method used

Using a meta-learning-based image transmission method, the pre-trained semantic feature extraction model and image reconstruction model are dynamically updated on small sample data sets to improve the model's adaptability and generalization ability.

Benefits of technology

Improve the quality of image reconstruction in image communication, enhance the processing capability of small sample data sets, and improve semantic fidelity and detail accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223907A_ABST
    Figure CN120223907A_ABST
Patent Text Reader

Abstract

The invention discloses an image transmission method based on meta-learning and related equipment, and aims to solve the problem of low image reconstruction quality in an image transmission process. The method comprises the steps that a transmitting end extracts semantic features of images in a to-be-transmitted data set through a pre-trained semantic feature extraction model in a semantic encoder, and the to-be-transmitted data set is a small sample data set; the transmitting end converts the semantic features into semantic symbols through a channel encoder and transmits the semantic symbols; the receiving end recovers the semantic symbols into semantic features through a channel decoder; the receiving end reconstructs the semantic features through a pre-trained image reconstruction model in a semantic decoder to obtain a reconstructed image; wherein the pre-trained semantic feature extraction model and the pre-trained image reconstruction model are dynamically updated based on the images in the small sample data set and the central element learning module. According to the invention, the image reconstruction quality in the image transmission process can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present application relate to the field of communication technologies, and in particular, to an image transmission method and system based on meta-learning. Background Art

[0002] In traditional image processing and communication technologies, image communication usually relies on pixel-level representation methods. This method encodes the input image pixel by pixel and then transmits the encoded bitstream to achieve image communication. However, traditional syntax communication schemes generate a large amount of redundancy during information transmission and have a low error tolerance, making it difficult to achieve ideal effects in terms of image reconstruction quality and transmission efficiency.

[0003] On this basis, related technologies have proposed image semantic communication technologies. Image semantic communication technology is a communication method centered on transmitting the key semantic features of images, aiming to improve information transmission efficiency and reduce bandwidth requirements by extracting and transmitting the semantic information of images rather than the specific values of each pixel. The semantic information extracted by image semantic communication technology based on deep learning includes shallow texture, shape information, and high-level conceptual semantic information of images. However, this technology faces challenges such as overfitting and poor generalization ability of deep learning models in small-sample environments, which affect the semantic fidelity and detail accuracy of image reconstruction. Summary of the Invention

[0004] In view of this, the purpose of one or more embodiments of the present application is to propose an image transmission method and system based on meta-learning to solve the problems raised in the background art.

[0005] Based on the above purpose, one or more embodiments of the present application provide an image transmission method based on meta-learning, including:

[0006] The transmitting end extracts the semantic features of the images in the dataset to be transmitted through a pre-trained semantic feature extraction model in the semantic encoder, and the dataset to be transmitted is a small-sample dataset;

[0007] The transmitting end converts the semantic features into semantic symbols through a channel encoder and transmits them;

[0008] The receiving end restores the semantic symbols to semantic features through a channel decoder;

[0009] The receiving end reconstructs the semantic features through a pre-trained image reconstruction model in the semantic decoder to obtain a reconstructed image;

[0010] Wherein, the pre-trained semantic feature extraction model and the pre-trained image reconstruction model are dynamically updated based on the images in the small-sample dataset and the central meta-learning module.

[0011] Optionally, the number of layers and hyperparameters of the pre-trained semantic feature extraction model and the pre-trained image reconstruction model are dynamically updated based on the images in the few-shot dataset and the central meta-learning module.

[0012] Optionally, the dynamic update steps of the pre-trained semantic feature extraction model and / or the pre-trained image reconstruction model include:

[0013] In response to receiving the few-shot dataset, constructing a support set and a query set according to the few-shot dataset;

[0014] According to the support set and the first initialization parameters of the pre-trained semantic feature extraction model and / or the pre-trained image reconstruction model, the second initialization parameters of the pre-trained semantic feature extraction model and / or the pre-trained image reconstruction model are obtained through iterative calculation; the first initialization parameters represent the model parameters determined in the pre-training stage or the model parameters at the end of the previous task, and the second initialization parameters represent the model parameters adapted to the dataset to be transmitted;

[0015] Based on the first initialization parameters, optimize the first initialization parameters according to the query set.

[0016] Optionally, the second initialization parameters are calculated through the following steps:

[0017] Obtain the first initialization parameters;

[0018] According to the first initialization parameters, use the learning rate and the first gradient loss corresponding to the first initialization parameters to iteratively calculate the second initialization parameters until the second initialization parameters converge.

[0019] Optionally, in each iteration of the second initialization parameters, the total parameter update amount is determined according to the learning rate, the loss gradient, and the regularization term.

[0020] Optionally, the dynamic update steps of the pre-trained semantic feature extraction model and / or the pre-trained image reconstruction model further include:

[0021] Determine the complexity of the images in the dataset to be transmitted through a preset method;

[0022] In response to determining that the complexity exceeds a preset threshold, increase the number of network layers of the model.

[0023] Optionally, the dynamic update steps of the pre-trained semantic feature extraction model and / or the pre-trained image reconstruction model further include:

[0024] According to the second initialization parameter of the pre-trained semantic feature extraction model, the third initialization parameter of the pre-trained semantic feature extraction model is iteratively calculated by using the number of newly added network layers, the second gradient loss, and the regularization term; the second gradient loss represents the feature loss of the deep layer in the model, and the third initialization parameter represents the model parameters of the pre-trained semantic feature extraction model after adding network layers, and the third initialization parameter adapts to the model parameters of the dataset to be transmitted.

[0025] Optionally, it further includes:

[0026] The transmitter optimizes the images in the dataset to be transmitted through the pre-trained feature enhancement model in the semantic encoder to obtain feature-enhanced images;

[0027] Among them, the pre-trained feature enhancement model is dynamically updated based on the images in the dataset to be transmitted and the central meta-learning module.

[0028] Optionally, the dynamic update step of the pre-trained feature enhancement model includes:

[0029] According to the fourth initialization parameter of the pre-trained feature enhancement model and the dataset to be transmitted, the fifth initialization parameter is iteratively calculated by using the learning rate and the third gradient loss until the fifth initialization parameter of the pre-trained feature enhancement model converges, and the third gradient loss represents the semantic difference between the original image and the reconstructed image;

[0030] Among them, the fourth initialization parameter is the model parameter determined in the pre-training stage or the model parameter when the previous task is completed, and the fifth initialization parameter represents the model parameter adapted to the dataset to be transmitted.

[0031] Based on the same inventive concept, one or more embodiments of the present application further provide a meta-learning-based image transmission system, including a transmitter and a system terminal:

[0032] The transmitter extracts the semantic features of the images in the dataset to be transmitted through the pre-trained semantic feature extraction model in the semantic encoder, and the dataset to be transmitted is a small-sample dataset; the semantic features are converted into semantic symbols through the channel encoder and transmitted;

[0033] The receiver restores the semantic symbols to semantic features through the channel decoder; the semantic features are reconstructed through the pre-trained image reconstruction model in the semantic decoder to obtain the reconstructed images;

[0034] Among them, the pre-trained semantic feature extraction model and the pre-trained image reconstruction model are dynamically updated based on the images in the small-sample dataset and the central meta-learning module.

[0035] As can be seen from the above, in the image transmission method based on meta-learning provided by one or more embodiments of the present application, the transmitter extracts the semantic features of the images in the dataset to be transmitted through the pre-trained semantic feature extraction model in the semantic encoder, and the dataset to be transmitted is a small-sample dataset; the transmitter converts the semantic features into semantic symbols through the channel encoder and transmits them; the receiver restores the semantic symbols to semantic features through the channel decoder; the receiver reconstructs the semantic features through the pre-trained image reconstruction model in the semantic decoder to obtain the reconstructed images; wherein, the pre-trained semantic feature extraction model and the pre-trained image reconstruction model are dynamically updated based on the images in the small-sample dataset and the central meta-learning module.

[0036] In the image transmission method based on meta-learning of the present application, a central meta-learning unit is adopted to dynamically update the pre-trained semantic feature extraction model and the pre-trained image reconstruction model in the semantic encoder and the semantic decoder based on the small-sample dataset to be transmitted, so as to improve the image reconstruction quality of the small-sample dataset.

[0037] A meta-learning-based image transmission system provided by the present application can implement the steps of the above image transmission method, and thus also has the beneficial effects of the above meta-learning-based image transmission method. Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in one or more embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only one or more embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0039] Figure 1 It is a schematic flow chart of the meta-learning-based image transmission method for one or more embodiments of the present application;

[0040] Figure 2 It is a schematic structural diagram of the meta-learning-based image transmission system for one or more embodiments of the present application;

[0041] Figure 3 It is a schematic flow chart of adjusting the number of model layers of the meta-learning-based image transmission method for one or more embodiments of the present application;

[0042] Figure 4 It is a schematic diagram of the experimental results of the meta-learning-based image transmission method for one or more embodiments of the present application;

[0043] Figure 5Schematic diagram of experimental results of the meta-learning-based image transmission method according to one or more embodiments of the present application;

[0044] Figure 6 Schematic diagram of the structure of the meta-learning-based image transmission system according to one or more embodiments of the present application;

[0045] Figure 7 Schematic diagram of the hardware structure of an electronic device according to one or more embodiments of the present application. Detailed implementation manners

[0046] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the following further describes the present disclosure in detail with reference to specific embodiments and the accompanying drawings.

[0047] It should be noted that unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present application should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The "first", "second", and similar terms used in one or more embodiments of the present application do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "include" or "comprise" mean that the elements or objects appearing before the term cover the elements or objects listed after the term and their equivalents, without excluding other elements or objects. The terms such as "connect" or "couple" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0048] As described in the background art, the image semantic communication technology adopted in the related art is prone to problems such as overfitting in training and poor generalization ability in a small-sample environment, making the semantic fidelity and detail accuracy of image reconstruction unable to achieve satisfactory effects.

[0049] Meta-Learning technology is a technology that enables a model to quickly adapt to multiple tasks through training. Meta-Learning can learn how to effectively transfer knowledge from multiple different tasks to enable the model to quickly learn and optimize when facing a new task with small samples.

[0050] In the process of implementing the present application, the applicant found that Meta-Learning technology can improve the efficient adaptive ability of the model when facing new tasks. Specifically, Meta-Learning technology helps the model accumulate experience through multiple learning tasks, extract general learning strategies and knowledge representations applicable to various tasks, and provide a solution with high generalization ability for image semantic information reconstruction under small samples and changing requirements.

[0051] Therefore, the present application uses meta-learning technology to propose an image transmission method based on meta-learning, so as to improve the reconstruction quality of images in image communication.

[0052] Referring to Figure 1 , the meta-learning-based image transmission method according to one or more embodiments of the present application includes the following steps:

[0053] Step S101: The transmitter extracts the semantic features of the images in the dataset to be transmitted through a pre-trained semantic feature extraction model in the semantic encoder, and the dataset to be transmitted is a small-sample dataset;

[0054] Step S102: The transmitter converts the semantic features into semantic symbols through a channel encoder and transmits them;

[0055] Step S103: The receiver restores the semantic symbols to semantic features through a channel decoder;

[0056] Step S104: The receiver reconstructs the semantic features through a pre-trained image reconstruction model in the semantic decoder to obtain a reconstructed image;

[0057] Wherein, the pre-trained semantic feature extraction model and the pre-trained image reconstruction model are dynamically updated based on the images in the small-sample dataset and the central meta-learning module.

[0058] In an embodiment of the present application, the transmitter may include a semantic encoder, and a pre-trained semantic feature extraction model is deployed on the semantic encoder; the receiver may include a semantic decoder, and a pre-trained image reconstruction model is deployed on the semantic decoder.

[0059] The transmitter should also be deployed with a channel encoder for converting the semantic features obtained by the semantic encoder into semantic symbols and transmitting them. The receiver should also be deployed with a channel decoder for restoring the semantic symbols to semantic features.

[0060] In an embodiment of the present application, the semantic feature extraction model and the image reconstruction model may be trained using meta-learning technology on a preset small-sample dataset, or may be trained using machine learning technology on a preset sample dataset.

[0061] In the technical solution of the present application, the fast learning and adaptation capabilities of meta-learning technology are used to dynamically adjust the model each time a small-sample dataset of a new task is received, so as to improve the image reconstruction ability when facing different tasks.

[0062] In an embodiment of the present application, the hyperparameters of the model and the number of layers of the model can be dynamically adjusted in combination with the central meta-learning module.

[0063] Specifically, in the embodiments of the present application, a meta-learning algorithm based on gradient update (Model-Agnostic Meta-Learning, MAML) can be used, and a task-driven learning method is adopted to update model parameters. The MAML algorithm can be applied to a variety of different types of tasks and can be quickly optimized through a small number of gradient updates on new tasks. In the case of scarce task data, it can effectively utilize a small number of samples for rapid learning, which is suitable for the technical solution of the present application.

[0064] In the embodiments of the present application, every time a new task is received, at least part of the images are extracted from the image dataset of the new task as the support set and at least part of the images are extracted as the query set; based on the first initialization parameters of the model, the second initialization parameters of the model are iteratively calculated using the above support set; based on the second initialization parameters of the model, the model parameters are optimized using the above query set. Among them, the first initialization parameters are the model parameters obtained by pre-training in the above text or the parameters of the model at the end of the previous task. The second initialization parameters are the initialization parameters adapted to the current new task and are adapted to the dataset to be transmitted. Optimizing the model based on the query set is to further optimize the model parameters during the image transmission process.

[0065] Specifically, in the embodiments of the present application, the parameters of the pre-trained semantic feature extraction model can be dynamically adjusted according to the learning rate and the gradient loss corresponding to the dataset to be transmitted. In the embodiments of the present application, the parameter update formula in the pre-trained semantic feature extraction model can be expressed as:

[0066]

[0067] Among them, Θ new represents the updated model parameters, Θ old represents the model parameters before update, α represents the learning rate, represents the gradient loss, T represents the total number of images in the task set, D i represents the i-th image in the task combination. The task set represents the above support set or query set.

[0068] In the embodiments of the present application, the total update amount of the parameters updated each time can be determined according to the learning rate, the loss gradient, and the regularization term. Specifically, in the embodiments of the present application, the calculation formula for the total update amount can be expressed as:

[0069]

[0070] Among them, g j (Θold,C j ) represents the regularization term, C j represents the constraint, β and λ respectively represent the learning rates of the gradient loss and the regularization term, w iDenotes the weight of the i-th image in the task set. The values of β and λ affect the learning rate α in the above text. It should be noted that the above learning rates α, β, and λ are also updated in each calculation.

[0071] The pre-trained image reconstruction model of this application uses enhanced decoded features to reconstruct the image, converting the decoded semantic features back into image data, specifically including multi-layer convolution and upsampling operations:

[0072]

[0073] Among them, I rec Denotes the reconstructed image, σ denotes ReLU or LeakyReLU, W k and b k Denote the convolutional kernel weight and bias term of the k-th layer respectively. Up() denotes the upsampling operation, which is used to increase the spatial dimension of the feature map and match higher-resolution image layers. F k-1 Denotes the feature output of the previous layer. Res(F k-1 , Θ meta,k ) is a residual connection, and the residual connection combines the parameters Θ optimized by the meta-learning algorithm meta,k , which is used to adjust the feature map of each layer. The residual learning part is adjusted in the following way:

[0074]

[0075] Among them, Θ old,k Denotes the meta-learning parameters before the update of the k-th layer network. α k Denotes the learning rate for the k-th layer network. Denotes the gradient of the meta-learning loss function with respect to Θ old,k , and the loss function evaluates the difference between the features G k generated by the current layer and the ideal features .

[0076] In the embodiments of this application, the number of network layers of the model can also be adjusted.

[0077] In the process of implementing this application, the applicant found that for images with higher complexity, deeper semantic feature extraction can be performed by increasing the number of network layers, and for images with lower complexity, the number of network layers can be reduced to lower the model training difficulty.

[0078] Thus, as Figure 3As shown, the embodiments of the present application can adjust the number of model network layers through the following steps: First, calculate the complexity of the images in the dataset to be transmitted; then determine whether the complexity of the images in the dataset to be transmitted is greater than a preset threshold; if the complexity is greater than the preset threshold, increase the number of convolutional layers of the model; if the complexity is less than or equal to the preset threshold, maintain the number of convolutional layers of the model; use the adjusted model to extract semantic information or reconstruct images from the images in the dataset to be transmitted.

[0079] In the embodiments of the present application, the information entropy of the image can be used to evaluate the image complexity. The specific calculation formula can be expressed as:

[0080]

[0081] where H(I) represents the comprehensive information entropy of image I, p(x) represents the probability distribution of pixel value x in image I, and G i,j represents the gradient intensity of pixel point (i,j), and the gradient intensity can be calculated by a Sobel filter.

[0082] In the embodiments of the present application, the increased number of network layers can be achieved by introducing additional deep convolutional layers and residual blocks to deeply extract high-level semantic features. The present application does not limit the specific method for increasing the number of network layers.

[0083] To achieve the fusion of the texture features at the bottom layer and the high-level semantic features of the image, a shortcut path with a 1×1 convolution operation can be added between the shallow convolutional layer and the deep convolutional layer of the pre-trained semantic feature extraction model. This 1×1 convolution operation processes the feature map output by the encoder, enabling the cross-fusion of information between channels. This operation can be expressed as:

[0084]

[0085] where F merged represents the fused feature map, F textural represents the texture feature map transferred from the bottom network layer, F semantic represents the high-level semantic feature map, W represents the weight matrix of the 1×1 convolution, and the operator represents the weighted combination of features.

[0086] In the embodiments of the present application, when the number of network layers increases, the parameters of the model can be further optimized using the newly added number of network layers, gradient loss, and regularization terms. Specifically, the calculation formula for updating the parameters can be:

[0087]

[0088] where K addDenotes the number of additional network layers according to the image complexity, γ k Denotes the adjustment coefficient of the number of network layers in the k-th layer to balance the influence of the gradient loss function on each network layer, F k Denotes the output feature of the k-th layer network. μ denotes the regularization parameter to promote the similarity of features between adjacent layers and enhance the consistency of the model, N k Denotes other network layers adjacent to the k-th layer network, F l Denotes the output feature of the l-th layer network, which is the shallow feature of the previous part of the network. It should be noted that the gradient loss in the above formula is the loss of feature extraction by the deep network in the model.

[0089] In the embodiments of the present application, before the transmitting end extracts semantic features through the pre-trained semantic feature extraction model, the image can also be enhanced by the pre-trained feature enhancement model. The pre-trained semantic feature extraction model can also be dynamically adjusted by the central meta-learning model.

[0090] Specifically, in the embodiments of the present application, the parameter update formula in the pre-trained semantic feature extraction model can be expressed as:

[0091]

[0092] Among them, η is the learning rate, is the loss function For the gradient of the network parameters, the loss function evaluates the semantic difference between the reconstructed image I' and the original image I.

[0093] In the embodiments of the present application, the semantic encoder can also include a pre-trained image feature extraction model. Specifically, in the embodiments of the present application, the Canny edge detection algorithm can be used to process the images in the dataset to be transmitted, strengthen the edge information of the images, and determine the edge positions by calculating the image gradients:

[0094]

[0095] G x = I * [-1,0,1;-2,0,2;-1,0,1], G y = I * [-1,-2,-1;0,0,0;1,2,1];

[0096] Among them, * represents the convolution operation.

[0097] In the embodiments of the present application, the edges detected by the Canny algorithm can also be weighted. Specifically, the edge weight enhancement factor λ can be introduced and dynamically adjusted according to the clarity and importance of the edges to increase the pixel weights of the key edge points:

[0098]

[0099] Among them, Edge_Strength(x, y) is the edge strength, which can be calculated from Gradient(I), and τ is a threshold parameter used to adjust the influence strength of the edge weight.

[0100] In addition, in the embodiments of the present application, the semantic encoder may further include an image preprocessing module for preprocessing the input image data, applying a low-pass filter to weaken the high-frequency information in the image to reduce the image sharpness, and cropping the image to make the input image size consistent.

[0101] The preprocessing process may specifically include:

[0102] First, apply an adaptive low-pass filter to perform frequency filtering on the input image data, automatically adjust the cut-off frequency according to the local variance of the image data, and calculate the cut-off frequency of the filter through the following formula:

[0103]

[0104] Among them, α and β are preset adjustment coefficients that control the baseline and flexibility of the cut-off frequency, is the average variance of the local area of the image, representing the detail complexity of the area, is the maximum value of the variances of all local areas of the image;

[0105] Then, perform dynamic size adjustment on the image data processed by the low-pass filter, and automatically select the optimal cropping frame according to the density of the image data content: analyze the information entropy of the image data, set a dynamic cropping window, and automatically adjust the cropping boundary according to the high-information-entropy area, so that the cropped image retains the most critical visual information, rather than cropping according to a fixed ratio or position, optimizing the size and information retention of the input image:

[0106]

[0107] Among them, w and h are the width and height of the cropping window respectively, and E ij is the information entropy value of the image at coordinates (i, j), indicating the information richness of the area, and w opt and h opt are the cropping window sizes that maximize the average information entropy within the window, and the cropped image retains the most critical visual information.

[0108] In the embodiments of the present application, the working principle of the channel encoder may include:

[0109] First, a deep learning network is used to perform channel coding on the image features received from the semantic encoder, which are optimized by edge enhancement and multi-task learning. The auto-encoder network structure is utilized to reduce the coding error and enhance the transmission robustness of semantic features. The activation function and network weights of the coding layer are dynamically adjusted to adapt to different image features:

[0110]

[0111] Among them, Θ old and Θ new represent the old parameters and new parameters of the coding network respectively, η represents the learning rate, the loss function measures the difference between the coded image X and the reconstructed image , γ is the adjustment factor of meta-learning, M t is the meta-knowledge adjustment matrix obtained from multi-task learning, and Θ t represents the optimal parameters of the t-th task.

[0112] Then, the coded features are transmitted through an additive Gaussian channel, and the channel model expression is:

[0113] Y = X + N, N ~ N(0, σ 2 I);

[0114] Among them, Y is the received signal, X is the signal after channel coding, N is the noise vector that follows a Gaussian distribution with a mean of 0 and a variance of σ 2 , and I is the identity matrix.

[0115] According to the above content, the overall semantic reconstruction system of this application is integrated:

[0116] S meta = σ(W1 * Θ enc + W2 * Θ dec + W3 * f extract (D small ));

[0117] Among them, σ represents the non-linear activation function, and the weight matrices W1, W2, and W3 are used to weight the encoder parameters, decoder parameters, and the features of the small sample dataset respectively. f extract is the feature extraction, which can be expressed as average pooling:

[0118]

[0119] Among them, n is the number of samples in the small sample dataset, is the i-th small sample.

[0120] The image transmission method provided by this application utilizes a dynamic network layer number adjustment mechanism to intelligently adjust the network depth according to the image complexity, so as to be able to flexibly adapt to different image reconstruction tasks. Specifically, when the image complexity is high, this mechanism increases the number of network layers to deeply mine the semantic information of the image, and when the image is relatively simple, it reduces the number of layers to improve the training speed. During the adjustment process, the model uses the strategy of meta-learning adaptive learning rate to ensure that it can be trained quickly and stably after each change in the network structure, effectively improving the problem of sensitive model initialization parameters in Comparative Document 3, enabling the entire network to flexibly adapt to the complexity of different images, and optimizing the computational efficiency and processing performance of the model. In addition, a shortcut path is introduced into the semantic feature extraction network, which effectively improves the training speed, reduces the loss during the information transmission process, retains more image details, and thus improves the quality of average image reconstruction.

[0121] To verify the technical effects of the technical solution of this application, the applicant conducted a verification experiment. In this verification experiment, the number of training images was set to 2000, and the amount of test image data was 500.

[0122] As Figure 4 shown, when the snr is 10dB, compared with the small-sample metaCT model using the teacher-student model and the semantic communication system based on traditional deep learning (Auto_EC), the proposed framework improves the image reconstruction quality by 8.23% and 12.32% respectively.

[0123] As Figure 5 shown, the average reconstruction quality of different solutions for multiple heterogeneous image dataset reconstruction tasks, where the image complexity indicators such as image semantic details and pixel matrix size are different in each dataset, and the data sample size is small. Compared with the semantic communication system based on traditional deep learning (Auto_EC), the average reconstruction quality of the proposed algorithm is improved by about 13.39%.

[0124] It can be understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities.

[0125] It should be noted that the method of one or more embodiments of this application can be executed by a single device, such as a computer or a server, etc. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of these multiple devices can only execute one or more steps of the method of one or more embodiments of this application, and these multiple devices will interact with each other to complete the described method.

[0126] It should be noted that the above description is for specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0127] Based on the same inventive concept, corresponding to any of the above method embodiments, the present application also provides an image transmission system based on meta-learning. As Figure 2 shown, the system includes a transmitting end 11 and a system end 12:

[0128] The above-mentioned transmitting end 11 extracts the semantic features of the images in the dataset to be transmitted through a pre-trained semantic feature extraction model in the semantic encoder, and the dataset to be transmitted is a small-sample dataset; the semantic features are converted into semantic symbols through a channel encoder and then transmitted;

[0129] The above-mentioned receiving end 12 restores the semantic symbols to semantic features through a channel decoder; the semantic features are reconstructed through a pre-trained image reconstruction model in the semantic decoder to obtain the reconstructed images;

[0130] Among them, the above-mentioned pre-trained semantic feature extraction model and the above-mentioned pre-trained image reconstruction model are dynamically updated based on the images in the above-mentioned small-sample dataset and the central meta-learning module.

[0131] In other words, as Figure 6 shown, the image transmission system of the present application includes an image preprocessing module, a semantic encoder, a channel encoder, a channel decoder, a semantic decoder, and a central meta-learning unit. The central meta-learning unit is used to dynamically adjust the parameters of the semantic encoder and the semantic decoder.

[0132] For the convenience of description, when describing the above devices, various modules are described separately according to their functions. Of course, when implementing one or more embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0133] The devices in the above embodiments are used to implement the corresponding methods in the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated herein.

[0134] Figure 7FIG. 0 shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0135] The processor 1010 may be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0136] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of the present application through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0137] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0138] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module may implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).

[0139] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0140] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of the present application, and does not necessarily include all the components shown in the figure.

[0141] The electronic device of the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here.

[0142] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0143] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary, and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments of the present application as described above, and they are not provided in detail for the sake of brevity.

[0144] Additionally, for simplicity of explanation and discussion, and so as not to render one or more embodiments of the present application difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid rendering one or more embodiments of the present application difficult to understand, and this also takes into account the fact that details regarding the implementation of these block diagram devices are highly dependent on the platform on which one or more embodiments of the present application are to be implemented (i.e., these details should be entirely within the understanding of those of ordinary skill in the art). In cases where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those of ordinary skill in the art that one or more embodiments of the present application may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be regarded as illustrative rather than restrictive.

[0145] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0146] One or more embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of the present application shall be included within the scope of protection of the present disclosure.

Claims

1. A meta-learning-based image transmission method, characterized in that: include: The transmitter extracts semantic features of images in the data set to be transmitted by using a pre-trained semantic feature extraction model in the semantic encoder, where the data set to be transmitted is a small sample data set; The transmitting end converts the semantic features into semantic symbols through a channel encoder and transmits the semantic symbols; The receiving end restores the semantic symbol into a semantic feature through a channel decoder; The receiving end reconstructs the semantic features through a pre-trained image reconstruction model in a semantic decoder to obtain a reconstructed image; Wherein, the pre-trained semantic feature extraction model and the pre-trained image reconstruction model are dynamically updated based on the images in the small sample data set and the central meta-learning module.

2. The method according to claim 1, characterized in that The number of layers and hyperparameters of the pre-trained semantic feature extraction model and the pre-trained image reconstruction model are dynamically updated based on the images in the small sample dataset and the central meta-learning module.

3. The method according to claim 2, characterized in that The step of dynamically updating the pre-trained semantic feature extraction model and / or the pre-trained image reconstruction model comprises: In response to receiving the small sample data set, constructing a support set and a query set according to the small sample data set; According to the support set and the first initialization parameters of the pre-trained semantic feature extraction model and / or the pre-trained image reconstruction model, second initialization parameters of the pre-trained semantic feature extraction model and / or the pre-trained image reconstruction model are obtained by iterative calculation; the first initialization parameters represent the model parameters determined in the pre-training stage or the model parameters when the previous task is completed, and the second initialization parameters represent the model parameters adapted to the data set to be transmitted; Based on the first initialization parameter, the first initialization parameter is optimized according to the query set.

4. The method according to claim 3, characterized in that The second initialization parameter is calculated by the following steps: Obtaining the first initialization parameter; According to the first initialization parameter, a second initialization parameter is iteratively calculated using a learning rate and a first gradient loss corresponding to the data set to be transmitted until the second initialization parameter converges.

5. The method according to claim 4, characterized in that In each round of iteration of the second initialization parameters, the total parameter update amount is determined according to the learning rate, the loss gradient and the regularization term.

6. The method according to claim 3, characterized in that The step of dynamically updating the pre-trained semantic feature extraction model and / or the pre-trained image reconstruction model further includes: Determining the complexity of the image in the data set to be transmitted in a preset manner; In response to determining that the complexity exceeds a preset threshold, the number of network layers of the model is increased.

7. The method according to claim 6, characterized in that The step of dynamically updating the pre-trained semantic feature extraction model and / or the pre-trained image reconstruction model further includes: According to the second initialization parameter of the pre-trained semantic feature extraction model, the third initialization parameter of the pre-trained semantic feature extraction model is obtained by iterative calculation using the newly added network layers, the second gradient loss and the regularization term; the second gradient loss represents the feature loss of the deep layer in the model, the third initialization parameter represents the model parameter of the pre-trained semantic feature extraction model after adding the network layer, and the third initialization parameter adapts to the model parameter of the data set to be transmitted.

8. The method according to claim 1, characterized in that Also includes: The transmitting end optimizes the image in the data set to be transmitted by using a pre-trained feature enhancement model in the semantic encoder to obtain a feature-enhanced image; The pre-trained feature enhancement model is dynamically updated based on the images in the data set to be transmitted and the central meta-learning module.

9. The method according to claim 8, characterized in that The step of dynamically updating the pre-trained feature enhancement model includes: Iteratively calculating a fifth initialization parameter of the pre-trained feature enhancement model according to the fourth initialization parameter of the pre-trained feature enhancement model and the data set to be transmitted using a learning rate and a third gradient loss until the fifth initialization parameter converges, wherein the third gradient loss represents a semantic difference between an original image and a reconstructed image; The fourth initialization parameter is a model parameter determined in a pre-training phase or a model parameter when a previous task is completed, and the fifth initialization parameter represents a model parameter adapted to the data set to be transmitted.

10. A meta-learning-based image transmission system, characterized in that: Including the transmitter and system: The transmitting end extracts semantic features of images in a data set to be transmitted by using a semantic feature extraction model pre-trained in a semantic encoder, wherein the data set to be transmitted is a small sample data set; converts the semantic features into semantic symbols by using a channel encoder, and transmits the semantic symbols; The receiving end restores the semantic symbol into a semantic feature through a channel decoder; Reconstructing the semantic features through a pre-trained image reconstruction model in a semantic decoder to obtain a reconstructed image; Wherein, the pre-trained semantic feature extraction model and the pre-trained image reconstruction model are dynamically updated based on the images in the small sample data set and the central meta-learning module.

Citation Information

Cited By

  • Visual optimization processing method and system based on real-time image processing

    CN120634893A

  • A visual optimization processing method and system based on real-time image processing

    CN120634893B