Image transmission method and device, electronic equipment and storage medium
By optimizing image transmission through spatial synonymous clustering and adaptive power allocation, the problems of insufficient image transmission efficiency and anti-interference capability in traditional architecture are solved, and high-reliability and low-latency image transmission effects are achieved.
Patent Information
- Application Number
- CN202510683680.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-10-03
AI Technical Summary
The traditional Shannon separation architecture has difficulty adapting to complex channel environments and limited bandwidth constraints in high-resolution image transmission, resulting in a decrease in reconstructed image quality. In addition, the existing DJSCC scheme ignores semantic redundancy and fixed power allocation, resulting in insufficient transmission efficiency and anti-interference ability.
Key semantic features are extracted through spatial synonymous clustering, and transmission power is adaptively allocated according to feature importance. Combined with channel resource scheduling, the image transmission process is optimized.
It improves the efficiency and quality of image transmission, enhances the noise resistance of key semantic features, and meets the needs of high-reliability, low-latency wireless communication.
Smart Images

Figure CN120751066A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present disclosure relate to the field of wireless communication technologies, and in particular, to an image transmission method, apparatus, electronic device, and storage medium. Background Art
[0002] It should be noted that the above technical background is merely provided to provide a clear and complete description of the technical solutions of the present invention and to facilitate understanding by those skilled in the art. Simply because these solutions are described in the technical background section of the present invention, it should not be assumed that the above technical solutions are well known to those skilled in the art.
[0003] With the rapid development of technologies such as autonomous driving, virtual reality (VR), augmented reality (AR), and the metaverse, wireless communication systems are facing a sharp increase in demand for real-time and reliable high-resolution image transmission. Traditional Shannon separation architectures, which utilize independently designed source and channel coding modules, face significant challenges in latency-sensitive scenarios. First, the strictly separated coding framework struggles to adapt to complex channel environments and limited bandwidth constraints, resulting in a significant decrease in reconstructed image quality. Second, the theoretical performance of separation coding is limited under short code lengths, making it ineffective in extracting and protecting semantic features, resulting in insufficient transmission efficiency and anti-interference capabilities. Summary of the Invention
[0004] In view of this, an object of one or more embodiments of the present disclosure is to provide an image transmission method, apparatus, electronic device, and storage medium to solve the problems raised in the background technology.
[0005] Based on the above objectives, one or more embodiments of the present disclosure provide an image transmission method, including:
[0006] Obtaining the image to be transmitted;
[0007] Extracting latent features of the image to be transmitted;
[0008] According to the latent features, obtaining features to be transmitted through spatial synonymous clustering, and determining allocation factors of the features to be transmitted;
[0009] Allocating transmission power to the feature to be transmitted according to the allocation factor of the feature to be transmitted;
[0010] A sequence to be transmitted is obtained according to the characteristics to be transmitted and the transmission power corresponding to the characteristics to be transmitted.
[0011] Optionally, obtaining the features to be transmitted by spatial synonymous clustering based on the latent features of the image to be transmitted includes:
[0012] Determining similarity between the latent feature and other latent features;
[0013] Determine the local density of the k nearest neighbors of the latent feature according to the similarity, where the local density reflects the density of the latent feature in the feature space;
[0014] Constructing a latent feature cluster according to the similarity between the latent feature and other latent features and the density of the latent features;
[0015] According to the preset transmission rules, the features to be transmitted are filtered from all the latent features.
[0016] Optionally, according to a preset transmission rule, filtering features to be transmitted from the latent features includes:
[0017] Determining a cluster center in the latent feature cluster according to the density and similarity of all latent features in the latent feature cluster;
[0018] Generate masks for the latent features other than the cluster center in the latent feature cluster and discard them;
[0019] The remaining latent features of the image to be transmitted are used as the features to be transmitted.
[0020] Optionally, the following steps are performed to determine the cluster center in the latent feature cluster:
[0021] Obtaining a distance index of the latent feature according to the local density of the latent feature in the latent feature cluster;
[0022] Obtaining a center score of the latent feature according to the distance index and the local density of the latent feature;
[0023] The latent feature with the highest center score of the latent feature cluster is determined as the cluster center.
[0024] Optionally, determining the allocation factor of the feature to be transmitted includes:
[0025] A power allocation factor of the feature to be transmitted is determined according to the initial power and the total power of the feature to be transmitted, where the power allocation factor is inversely proportional to the initial power of the feature to be transmitted.
[0026] Optionally, obtaining the sequence to be transmitted according to the feature to be transmitted and the transmission power corresponding to the feature to be transmitted includes:
[0027] The to-be-transmitted sequence is obtained by multiplying the to-be-transmitted characteristic by a power allocation factor of the to-be-transmitted characteristic.
[0028] Optionally, receiving a transmission sequence;
[0029] Obtaining a reconstructed semantic feature sequence according to the transmission sequence and the power allocation factor;
[0030] Obtaining a reconstructed latent feature sequence based on the reconstructed semantic features and a preset latent feature prediction model;
[0031] The reconstructed latent feature sequence is subjected to source decoding to obtain a reconstructed image.
[0032] Based on the same inventive concept, one or more embodiments of the present disclosure further provide an image transmission device, including:
[0033] An acquisition module, configured to acquire an image to be transmitted;
[0034] A first computing module is configured to extract latent features of the image to be transmitted;
[0035] a second calculation module configured to obtain features to be transmitted by spatial synonymous clustering based on the latent features, and determine a distribution factor of the features to be transmitted, where the distribution factor is positively correlated with the importance of the features to be transmitted;
[0036] a third calculation module, configured to allocate transmission power to the feature to be transmitted according to the allocation factor of the feature to be transmitted;
[0037] The fourth calculation module is configured to obtain a sequence to be transmitted according to the feature to be transmitted and the transmission power corresponding to the feature to be transmitted.
[0038] Based on the same inventive concept, one or more embodiments of the present disclosure also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the image transmission method as described in any one of the above items is implemented.
[0039] Based on the same inventive concept, one or more embodiments of the present disclosure further provide a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute any of the above-mentioned image transmission methods.
[0040] From the above description, it can be seen that the image transmission method provided by one or more embodiments of the present disclosure is collaboratively optimized from the two perspectives of semantic redundancy compression and channel resource optimization. By jointly optimizing semantic feature extraction, redundancy elimination and channel resource scheduling, it breaks through the performance bottleneck of the existing DJSCC framework and meets the core requirements of the next generation of wireless communications for highly reliable and low-latency image transmission.
[0041] The image transmission device, electronic device, and computer-readable storage medium provided by the present disclosure are all capable of implementing the steps of the above-mentioned image transmission method, and therefore also have the beneficial effects of the above-mentioned image transmission method. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate one or more embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only one or more embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0043] Figure 1 A flowchart of an image transmission method according to one or more embodiments of the present disclosure is provided;
[0044] Figure 2 A schematic structural diagram of an image transmission device according to one or more embodiments of the present disclosure;
[0045] Figure 3 A schematic diagram of a process for generating features to be transmitted according to one or more embodiments of the present disclosure;
[0046] Figure 4 A schematic diagram of the structure of a latent feature prediction model according to one or more embodiments of the present disclosure;
[0047] Figure 5 A schematic diagram of performance experimental results of one or more embodiments of the present disclosure;
[0048] Figure 6 A schematic diagram of experimental results of the effects of one or more embodiments of the present disclosure;
[0049] Figure 7 A schematic diagram of the experimental results of semantic compression effect according to one or more embodiments of the present disclosure;
[0050] Figure 8 A schematic diagram of the hardware structure of an electronic device according to one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0051] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0052] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present disclosure should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in one or more embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0053] As described in the background technology section, with the development of science and technology, the requirements for real-time and reliability of high-resolution image transmission in wireless communication systems have increased dramatically.
[0054] In order to improve the effect of image transmission, the related technology has proposed Deep Joint Source-Channel Coding (DJSCC) based on deep learning. However, although Deep Joint Source-Channel Coding (DJSCC) based on deep learning has achieved a breakthrough in bit rate-distortion performance through end-to-end optimization, it still has inherent defects: on the one hand, the existing DJSCC scheme regards the spatial domain features of the image as equally important, ignoring the compressibility of semantic redundancy in the spatial domain (such as repeated textures and similar structural areas); on the other hand, the fixed power allocation strategy leads to low channel resource utilization, and key semantic features are easily interfered under low signal-to-noise ratio conditions, further restricting the improvement of system performance.
[0055] In addition, the research on related technologies for semantic redundancy compression and channel resource optimization is still fragmented.
[0056] Spatial synonymy theory suggests that feature fusion can be used to significantly compress data in adjacent semantically similar regions of an image, while dynamic power allocation technology can enhance the noise resistance of key features based on the differentiation of semantic importance. However, existing technologies have not yet achieved the coordinated optimization of these two types of technologies, making it difficult to balance transmission efficiency and robustness. Therefore, a new transmission scheme that integrates spatial synonymy compression and adaptive power allocation is urgently needed. By jointly optimizing semantic feature extraction, redundancy elimination, and channel resource scheduling, it can break through the performance bottleneck of the existing DJSCC framework and meet the core requirements of next-generation wireless communications for highly reliable and low-latency image transmission.
[0057] According to some implementations of the present disclosure, a scheme for image transmission is provided. In this scheme, the latent features of the image are first extracted; then, a spatial synonymous clustering algorithm is used to adaptively process the latent features of the image, extracting key semantic features for wireless transmission to reduce the amount of data while retaining important information; finally, an adaptive power allocation method is used to allocate corresponding power according to the importance of different semantic features to reduce the impact of channel noise on transmission, improve the noise resistance of key semantic features, and enhance the reliability of transmission. Through the scheme of the present disclosure, the image transmission scheme can be jointly optimized based on spatial synonymous feature compression and dynamic power allocation, so that during the image transmission process, key semantic information is retained while key power is increased, thereby improving the efficiency and quality of image transmission.
[0058] refer to Figure 1 The image transmission method of one or more embodiments of the present disclosure includes the following steps:
[0059] Step S101: Acquire the image to be transmitted;
[0060] Step S102: extracting latent features of the image to be transmitted;
[0061] Step S103: Based on the latent features, spatial synonymous clustering is performed to obtain features to be transmitted, and a distribution factor of the features to be transmitted is determined. The distribution factor is positively correlated with the importance of the features to be transmitted.
[0062] Step S104: allocating transmission power to the characteristic to be transmitted according to the allocation factor of the characteristic to be transmitted;
[0063] Step S105: obtaining a sequence to be transmitted according to the characteristics to be transmitted and the transmission power corresponding to the characteristics to be transmitted.
[0064] In the implementation of the present disclosure, after obtaining the grammatical symbol sequence of the image to be transmitted through the source coding method, the corresponding latent features can be obtained according to the grammatical symbol sequence.
[0065] In this disclosed example, latent features can be obtained using a Swin Transformer neural network model. Specifically, the Swin Transformer neural network model takes the image to be transmitted as input and outputs serialized tokens. Swin Transformer can capture local and global information of the image, thereby generating rich high-dimensional latent features. These feature representations are further processed to adapt to different channel conditions and transmission rates, providing an efficient foundation for subsequent image transmission.
[0066] In order to further improve the efficiency and quality of image transmission, the technical solution disclosed in the present invention further screens latent features to obtain a small number of representative semantic features to improve transmission efficiency.
[0067] In the implementation of the present disclosure, spatial synonymous clustering is first performed on the latent features to obtain the similarity between each latent feature and other latent features; then, all latent features are clustered according to the similarity between each latent feature and other latent features and the distance index of the latent features to obtain latent feature clusters; finally, the latent feature clusters are screened to eliminate other latent features outside the cluster center in the latent feature cluster.
[0068] In the example disclosed herein, the specific process is as follows:
[0069] First, calculate the similarity between each latent feature and other latent features. This similarity can be measured by the distance between the latent features. In other words, it can be calculated using the following formula:
[0070] d(y i ,y j )=||y i -y j ||;
[0071] Among them, y i represents the i-th latent feature, y j represents the jth latent feature.
[0072] Afterwards, the distance index of each latent feature is calculated.
[0073] In one or more examples of the present disclosure, the local density of the latent features may be calculated first, and then the distance index of each latent feature may be determined based on the local density of the latent features.
[0074] The local density of the k nearest neighbors of each latent feature reflects its density in the feature space.
[0075] The local density of the i-th latent feature can be calculated by the following formula:
[0076]
[0077] Among them, k represents the number of latent features involved in the calculation.
[0078] The distance index can be calculated by the following formula:
[0079]
[0080] Then, all latent features are clustered according to the similarity between each latent feature and other latent features and the distance index of the latent features to obtain latent feature clusters. In the implementation of the present disclosure, the cluster center in each latent feature cluster represents the latent feature of the corresponding latent feature cluster.
[0081] In the example disclosed herein, the product of the local density and distance index of all latent features in a latent feature cluster is calculated to obtain a score for each latent feature, and the latent feature with the highest score is used as the cluster center. The cluster center represents the key semantics of the latent feature cluster.
[0082] In the implementation of the present disclosure, the other latent features in each latent feature cluster except the cluster center can be masked and discarded, wherein the mask is a 0-1 matrix of the same size as the latent feature.
[0083] Through the above-mentioned approach, the present disclosure greatly reduces the amount of transmitted data while retaining key semantic information, optimizes bandwidth, and thereby improves transmission efficiency while ensuring image transmission quality.
[0084] In one example of the present disclosure, the process of generating features to be transmitted is as follows: Figure 3 As shown in Figure 1. Each latent feature cluster, outlined in red, contains latent features that are considered synonymous and belong to the same cluster. The cluster center of a latent feature cluster is marked with a red cross, indicating the representative feature. Masks, shown in black, represent latent features that are masked and discarded during transmission. The channel-to-bandwidth ratio can be adjusted by changing the retention rate of latent features.
[0085] In the implementation of the present disclosure, while reducing the amount of transmitted data, the transmission power is adaptively allocated to the features to be transmitted to improve the transmission quality, thereby collaboratively optimizing the image transmission efficiency and quality.
[0086] In the implementation of the present disclosure, the power allocation factors are allocated with the goal of obtaining the optimal noise resistance capability within the total power.
[0087] In the implementation of the present disclosure, the initial power of each feature to be transmitted is first determined by the following formula.
[0088]
[0089] in, Indicates the characteristics to be transferred.
[0090] Then, according to the power allocation calculation formula, the power allocation factor of each feature to be transmitted is obtained.
[0091]
[0092] Where P represents the total power.
[0093] It can be understood that a feature to be transmitted with a large initial power can be regarded as having a stronger ability to resist noise, and a smaller power allocation factor can be allocated to it; conversely, a feature to be transmitted with a small initial power can be allocated a larger power allocation factor.
[0094] By multiplying the above characteristics to be transmitted by the corresponding power allocation factor, the corresponding sequence to be transmitted can be obtained.
[0095] The above steps can all be implemented by the transmitter. For the receiver, the reconstructed image can be obtained by performing the following steps:
[0096] Step S201: receiving a transmission sequence;
[0097] Step S202: obtaining a reconstructed semantic feature sequence according to the transmission sequence and the power allocation factor;
[0098] Step S203: obtaining a reconstructed latent feature sequence based on the above-mentioned reconstructed semantic features and the preset latent feature prediction model;
[0099] Step S204: performing source decoding on the reconstructed latent feature sequence to obtain a reconstructed image.
[0100] In the implementation of the present disclosure, when performing channel transmission, the sequence to be transmitted is converted into an analog sequence and can be directly transmitted to the receiving end through the wireless channel.
[0101] In the example of the present disclosure, for the channel model W, its transfer function is:
[0102]
[0103] Where ⊙ represents element-by-element multiplication, h is the channel response vector, and n is the noise vector. The components of the noise vector are independently sampled from a Gaussian distribution, i.e. Represents the noise power. The received sequence is the reconstructed transmission sequence after the influence of channel noise. It contains the information of the original transmission sequence s as well as the noise and fading introduced during the transmission process. This process simulates the signal transmission in actual wireless communication and provides a basis for subsequent signal detection and image reconstruction.
[0104] After receiving the transmission sequence, the receiver can use the known power allocation factor g i The received signal Perform signal detection and reconstruct semantic features through reverse power adjustment
[0105] That is, for each received signal value The formula can be Calculate and reconstruct semantic features to eliminate the power scaling effect at the transmitter. This process is crucial for accurately restoring the original semantic features because it minimizes the impact of channel noise on transmission and ensures the reconstructed semantic feature sequence. The semantic information of the original image can be restored as accurately as possible even in the presence of noise interference.
[0106] Power allocation and signal detection adaptively optimize the distribution of transmission power, ensuring that key semantic features are more resistant to noise during transmission. This strategy not only improves transmission reliability but also allows for optimal image quality within given constraints. Because power allocation is a linear transformation, it can be easily integrated into neural networks, allowing scaling factors to be adjusted via gradient backpropagation during training. This end-to-end joint training enables the entire framework to learn all parameters simultaneously, improving overall performance and image reconstruction quality.
[0107] In the implementation of the present disclosure, a reconstructed latent feature sequence is obtained through a preset latent feature prediction model.
[0108] In the example disclosed herein, the latent feature prediction model can be constructed using a CNN network. Figure 4 , which is a structural diagram of a latent feature prediction model of one or more examples of the present disclosure. The latent feature prediction model consists of a convolutional layer and a LeakyReLU activation function, and can learn local details and global semantic information of an image.
[0109] The reconstructed semantic feature sequence can be calculated by the following formula:
[0110]
[0111] The architecture uses a prediction model to accurately restore the masked areas, thereby improving the quality and completeness of the decoded features and providing a high-quality foundation for the final image reconstruction.
[0112] In the implementation of the present disclosure, a Swin Transformer-based decoder can be used to implement source decoding. The Swin Transformer decoder has a powerful feature representation capability and can gradually restore the semantic information and details of the image.
[0113] The decoder gradually reconstructs the grammatical symbol sequence of the original image based on the received feature representation This process not only restores the semantic information of the image, but also retains the details of the image as much as possible, ensuring the integrity and accuracy of the reconstructed image.
[0114] In order to verify the technical effect of the technical solution disclosed in the present invention, the applicant conducted experiments.
[0115] The experimental conditions include using the public OpenImages dataset for training and Kodak for testing, covering images of diverse scenes and content. The input images are preprocessed, resized to a uniform size (256×256), and normalized to the range [0, 1].
[0116] Figure 5 The PSNR and LPIPS performance comparisons between this embodiment and other existing methods under different signal-to-noise ratio (SNR) and channel bandwidth ratio (CBR) conditions are demonstrated, demonstrating the performance advantages of the semantic communication method based on spatial synonymy feature compression and power allocation.
[0117] Figure 6 The figure shows a comparison between the original image and the reconstructed image. The first and second columns show the original image and its local image block, respectively, while the third through fifth columns show the reconstruction results using different transmission methods. The corresponding resolution, CBR, PSNR, and LPIPS values are annotated below each image, visually demonstrating the advantages of this method in reconstruction quality under different signal-to-noise ratio and bandwidth conditions.
[0118] Figure 7 The ability of this method to compress features and preserve semantic information is further demonstrated by visualizing the clustering and masking results at different CBRs. The first row of the figure shows clustered regions, identified by blocks of the same color, indicating semantically similar regions within the image, or synonyms. The second row shows masked regions in gray, with the remaining regions retaining key semantic information. This result demonstrates that this method can effectively preserve key semantic features of the image while reducing data volume, achieving a good balance between transmission efficiency and image quality.
[0119] It should be noted that the methods of one or more embodiments of the present disclosure can be performed by a single device, such as a computer or server. The methods of these embodiments can also be applied in a distributed scenario, performed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the methods of one or more embodiments of the present disclosure, and the multiple devices will interact with each other to complete the described methods.
[0120] It should be noted that the above description is of specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0121] Based on the same inventive concept, corresponding to any of the above embodiments and methods, the present disclosure also provides an image transmission device. Figure 2 As shown, the device includes:
[0122] An acquisition module 11 is configured to acquire an image to be transmitted;
[0123] A first computing module 12 is configured to extract latent features of the image to be transmitted;
[0124] The second calculation module 13 is configured to obtain the features to be transmitted by spatial synonymous clustering based on the latent features, and determine the allocation factors of the features to be transmitted;
[0125] The third calculation module 14 is configured to allocate transmission power to the above-mentioned feature to be transmitted according to the allocation factor of the above-mentioned feature to be transmitted;
[0126] The fourth calculation module 15 is configured to obtain a sequence to be transmitted according to the characteristics to be transmitted and the transmission power corresponding to the characteristics to be transmitted.
[0127] For the convenience of description, the above devices are described as being functionally divided into various modules. Of course, when implementing one or more embodiments of the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0128] The apparatus of the above embodiment is used to implement the corresponding method in the above embodiment and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0129] Figure 8 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0130] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure.
[0131] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present disclosure are implemented through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0132] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0133] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0134] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0135] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, those skilled in the art will understand that the above device may only include the components necessary to implement the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.
[0136] The electronic devices of the above embodiments are used to implement the corresponding methods in the above embodiments and have the beneficial effects of the corresponding method embodiments, which will not be described in detail here.
[0137] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0138] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Within the scope of the present disclosure, the technical features of the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.
[0139] In addition, to simplify the description and discussion, and so as not to obscure one or more embodiments of the present disclosure, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, devices may be shown in the form of block diagrams to avoid obscuring one or more embodiments of the present disclosure, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which one or more embodiments of the present disclosure will be implemented (i.e., these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that one or more embodiments of the present disclosure may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0140] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0141] The one or more embodiments of the present disclosure are intended to encompass all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the one or more embodiments of the present disclosure should be included within the scope of protection of the present disclosure.
Claims
1. An image transmission method, characterized in that: include: Obtaining the image to be transmitted; Extracting latent features of the image to be transmitted; According to the latent features, obtaining features to be transmitted through spatial synonymous clustering, and determining allocation factors of the features to be transmitted; Allocating transmission power to the feature to be transmitted according to the allocation factor of the feature to be transmitted; A sequence to be transmitted is obtained according to the characteristics to be transmitted and the transmission power corresponding to the characteristics to be transmitted.
2. The method according to claim 1, characterized in that The step of obtaining the features to be transmitted by spatial synonymous clustering based on the latent features of the image to be transmitted includes: Determining similarity between the latent feature and other latent features; Determine the local density of the k nearest neighbors of the latent feature according to the similarity, where the local density reflects the density of the latent feature in the feature space; Constructing a latent feature cluster according to the similarity between the latent feature and other latent features and the density of the latent features; According to the preset transmission rules, the features to be transmitted are filtered from all the latent features.
3. The method according to claim 2, characterized in that According to the preset transmission rules, the features to be transmitted are selected from the latent features, including: Determining a cluster center in the latent feature cluster according to the density and similarity of all latent features in the latent feature cluster; Generate masks for the latent features other than the cluster center in the latent feature cluster and discard them; The remaining latent features of the image to be transmitted are used as the features to be transmitted.
4. The method according to claim 3, characterized in that Perform the following steps to determine the cluster center in the latent feature cluster: Obtaining a distance index of the latent feature according to the local density of the latent feature in the latent feature cluster; Obtaining a center score of the latent feature according to the distance index and the local density of the latent feature; The latent feature with the highest center score of the latent feature cluster is determined as the cluster center.
5. The method according to claim 1, wherein The determining of the allocation factor of the feature to be transmitted includes: A power allocation factor of the feature to be transmitted is determined according to the initial power and the total power of the feature to be transmitted, where the power allocation factor is inversely proportional to the initial power of the feature to be transmitted.
6. The method according to claim 5, characterized in that The obtaining a sequence to be transmitted according to the characteristics to be transmitted and the transmission power corresponding to the characteristics to be transmitted includes: The to-be-transmitted sequence is obtained by multiplying the to-be-transmitted characteristic by a power allocation factor of the to-be-transmitted characteristic.
7. The method according to claim 1, characterized in that Also includes: Receive transmission sequence; Obtaining a reconstructed semantic feature sequence according to the transmission sequence and the power allocation factor; Obtaining a reconstructed latent feature sequence based on the reconstructed semantic features and a preset latent feature prediction model; The reconstructed latent feature sequence is subjected to source decoding to obtain a reconstructed image.
8. An image transmission device, characterized in that: include: An acquisition module, configured to acquire an image to be transmitted; A first computing module is configured to extract latent features of the image to be transmitted; a second calculation module configured to obtain features to be transmitted by spatial synonymous clustering based on the latent features, and determine a distribution factor of the features to be transmitted; a third calculation module, configured to allocate transmission power to the feature to be transmitted according to the allocation factor of the feature to be transmitted; The fourth calculation module is configured to obtain a sequence to be transmitted according to the feature to be transmitted and the transmission power corresponding to the feature to be transmitted.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executed by the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.