Image transmission method and device, image compression component and computer equipment
By downsampling the image to be transmitted and extracting feature information using multiple sub-models, the problem of inaccurate feature extraction in image transmission is solved, the image compression quality and speed are improved, and power consumption and delay are reduced.
Patent Information
- Application Number
- CN202510344105.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-10
AI Technical Summary
During the transmission process of images, the prior art is inaccurate in the extraction of the feature of the image, which affects the image compression quality, and it is difficult to ensure the quality of image compression while obtaining a higher compression ratio.
An image transmission method is adopted to downsample the image to be transmitted, and the downsampled results are input into the image compression model. Using the image compression model including the first sub-model, the second sub-model and the third sub-model, the local spatial feature information, global spatial feature information and channel feature information of the downsampled results are extracted, and stitched to generate a compressed image.
By accurately extracting and retaining important features of the image, we improve image compression quality and compression rate, while reducing the power consumption required for compressed images, improving image compression speed, and reducing transmission delay.
Smart Images

Figure CN120128720A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to an image transmission method, device, image compression component, and computer device. Background Art
[0002] In the aspect of remote monitoring and management of a server, the management end of the server obtains the server desktop image, and realizes real-time monitoring of the running state of the server through the information in the server desktop image, and analyzes the monitoring data, so as to realize remote control of the server.
[0003] The resolution of the server desktop image is high. If the original image is directly transmitted during remote transmission, it will occupy a large amount of memory and network bandwidth, and it needs to be compressed. However, the image compression technology currently used for transmitting images extracts inaccurate features of the image and cannot accurately retain various feature information of the image, so the image compression quality is affected. It is difficult to ensure the quality of image compression on the premise of obtaining a high compression ratio.
[0004] Therefore, in the process of transmitting images in related technologies, the feature extraction of the images is inaccurate, which affects the image compression quality. Summary of the Invention
[0005] In view of this, the present invention provides an image transmission method, device, image compression component, and computer device to solve the problem that in the process of transmitting images, the feature extraction of the images is inaccurate, which affects the image compression quality.
[0006] In a first aspect, the present invention provides an image transmission method, which is applied to an image compression component, and the method includes:
[0007] Obtain an image to be transmitted;
[0008] Perform downsampling on the image to be transmitted, input the downsampling result into an image compression model, and obtain a compressed image. The image compression model includes a first sub-model, a second sub-model, and a third sub-model. The first sub-model is used to extract the first local spatial feature information of the downsampling result, the second sub-model is used to extract the first global spatial feature information of the downsampling result, and the third sub-model is used to extract the first channel feature information of the downsampling result. The compressed image is obtained by splicing the first local spatial feature information, the first global spatial feature information, and the first channel feature information;
[0009] Send the compressed image to the control end, so that the control end obtains the image to be transmitted according to the compressed image and an image decompression model. The image decompression model includes a first sub-model, a second sub-model, and a third sub-model.
[0010] Second aspect, the present invention provides an image transmission method, which is applied to a control end, and the method includes:
[0011] Obtain the compressed image from an image compression component, wherein the compressed image is obtained by splicing first local spatial feature information, first global spatial feature information, and first channel feature information. The first local spatial feature information is extracted by a first sub-model from the downsampling result, the first global spatial feature information is extracted by a second sub-model from the downsampling result, the first channel feature information is extracted by a third sub-model from the downsampling result. The first sub-model, the second sub-model, and the third sub-model are included in an image compression model, and the downsampling result is obtained by downsampling the image to be transmitted;
[0012] Upsample the compressed image, and input the upsampling result into an image decompression model to obtain the image to be transmitted. The image decompression model is used to generate the image to be transmitted according to the local spatial feature information, global spatial feature information, and channel feature information of the upsampling result. The image decompression model includes a first sub-model, a second sub-model, and a third sub-model.
[0013] Third aspect, the present invention provides an image compression component, which includes: a compression module, a storage module, and a controller; the compression module includes an image compression model;
[0014] The compression module is used to obtain the image to be transmitted, downsample the image to be transmitted, and input the downsampling result into the image compression model to obtain the compressed image. The image compression model includes a first sub-model, a second sub-model, and a third sub-model. The first sub-model is used to extract the first local spatial feature information of the downsampling result, the second sub-model is used to extract the first global spatial feature information of the downsampling result, the third sub-model is used to extract the first channel feature information of the downsampling result, and the compressed image is obtained by splicing the first local spatial feature information, the first global spatial feature information, and the first channel feature information;
[0015] The compression module is used to store the compressed image into the storage module;
[0016] The controller is used to read the compressed image from the storage module and send the compressed image to the control end, so that the control end obtains the image to be transmitted according to the compressed image and the image decompression model. The image decompression model includes a first sub-model, a second sub-model, and a third sub-model.
[0017] Fourth aspect, the present invention provides an image transmission device, which is deployed in the image compression component, and the device includes:
[0018] A first image acquisition module, which is used to acquire the image to be transmitted;
[0019] An image compression module, configured to downsample an image to be transmitted, input the downsampling result into an image compression model, and obtain a compressed image. The image compression model includes a first sub-model, a second sub-model, and a third sub-model. The first sub-model is used to extract first local spatial feature information of the downsampling result, the second sub-model is used to extract first global spatial feature information of the downsampling result, and the third sub-model is used to extract first channel feature information of the downsampling result. The compressed image is obtained by splicing the first local spatial feature information, the first global spatial feature information, and the first channel feature information;
[0020] An image sending module, configured to send the compressed image to a control end, so that the control end obtains the image to be transmitted according to the compressed image and an image decompression model. The image decompression model includes a first sub-model, a second sub-model, and a third sub-model.
[0021] In a fifth aspect, the present invention provides an image transmission device, which is deployed at the control end. The device includes:
[0022] A second image acquisition module, configured to acquire the compressed image from an image compression component. The compressed image is obtained by splicing the first local spatial feature information, the first global spatial feature information, and the first channel feature information. The first local spatial feature information is extracted by the first sub-model from the downsampling result, the first global spatial feature information is extracted by the second sub-model from the downsampling result, the first channel feature information is extracted by the third sub-model from the downsampling result. The first sub-model, the second sub-model, and the third sub-model are included in the image compression model, and the downsampling result is obtained by downsampling the image to be transmitted;
[0023] An image decompression module, configured to upsample the compressed image, input the upsampling result into an image decompression model, and obtain the image to be transmitted. The image decompression model is used to generate the image to be transmitted according to the local spatial feature information, global spatial feature information, and channel feature information of the upsampling result. The image decompression model includes a first sub-model, a second sub-model, and a third sub-model.
[0024] In a sixth aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the image transmission method according to the first aspect or any corresponding embodiment thereof.
[0025] In a seventh aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored. The computer instructions are used to cause a computer to execute the image transmission method according to the first aspect or any corresponding embodiment thereof.
[0026] In an eighth aspect, the present invention provides a computer program product, including computer instructions for causing a computer to execute the image transmission method according to the first aspect or any corresponding embodiment thereof above.
[0027] Through this application, the image to be transmitted is downsampled, the downsampling result is input into an image compression model, the first local spatial feature information of the downsampling result is extracted by a first sub-model in the image compression model, the first global spatial feature information of the downsampling result is extracted by a second sub-model, and the first channel feature information of the downsampling result is extracted by a third sub-model. This solves the problem that the feature extraction of the image is inaccurate during the image transmission process, which affects the image compression quality. It has the effects of accurately extracting and retaining important features of the image, improving the image compression quality and compression ratio, reducing the power consumption required for compressing the image, increasing the image compression speed, and reducing the transmission delay. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the related art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the related art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0029] Figure 1 is a schematic flowchart of an image transmission method applied to an image compression component according to an embodiment of the present invention;
[0030] Figure 2 is a flowchart of image compression and decompression for server remote desktop monitoring and management according to an embodiment of the present invention;
[0031] Figure 3 is a schematic structural diagram of an image compression model according to an embodiment of the present invention;
[0032] Figure 4 is a schematic diagram of an image compression and decompression deep neural network architecture according to an embodiment of the present invention;
[0033] Figure 5 is a flowchart of extracting global spatial feature information and channel feature information according to an embodiment of the present invention;
[0034] Figure 6 is a schematic flowchart of an image transmission method applied to a control end according to an embodiment of the present invention;
[0035] Figure 7 is a structural diagram of an image compression component according to an embodiment of the present invention;
[0036] Figure 8 It is a structural diagram of a compression module in an image compression component according to an embodiment of the present invention;
[0037] Figure 9 It is a structural block diagram of an image transmission device deployed in an image compression component according to an embodiment of the present invention;
[0038] Figure 10 It is a structural block diagram of an image transmission device deployed at a control end according to an embodiment of the present invention;
[0039] Figure 11 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed implementation manners
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] In terms of remote monitoring and management of a server, remote desktop management is an important application for monitoring and managing a server system. By using the server desktop image, the running state of the server can be monitored in real time remotely, and the monitoring data can be analyzed, and then functions such as remote control of the server, hardware configuration management, hardware diagnosis, and troubleshooting can be realized. Since the resolution of the server desktop image is high, directly transmitting the original image during the image transmission process will occupy a large amount of memory and network bandwidth, with a large transmission delay and it is difficult to achieve real-time performance. Therefore, it is necessary to compress the image during the image transmission process.
[0042] Currently, traditional image compression technologies are mainly divided into video compression technologies for multi-frame time series and image compression technologies for single-frame space such as JPEG (Joint Photographic Experts Group, an image compression standard), etc. However, as server desktop images become increasingly complex, it is difficult for these traditional image compression technologies to ensure the quality of image compression while achieving a high compression ratio. In addition, currently, some image compression methods based on CNN (Convolutional neural networks) or ViT (Vision Transformer) are adopted. Although these methods have good compression effects, the CNN-based image compression method mainly focuses on the local information of the image, and the ViT-based image compression method mainly focuses on the global information of the image. Both methods ignore the channel information of the image, and neither method can simultaneously focus on and process the local information, global information, and channel information of the image, resulting in inaccurate feature extraction and the inability to accurately retain the important information of the image, thus affecting the image compression quality. In addition, the computational complexity of the ViT-based image compression method is huge, and it is difficult for the Baseboard Management Controller (BMC) to provide sufficient computing resources.
[0043] Based on the above, an embodiment of the present invention provides an image transmission method. During the entire compression and decompression process, first, obtain the server host image, and use the image compression deep neural network based on long and short spatial attention and channel attention to compress the image. This network can capture the local information, global information, and channel information of the image, thereby extracting the most important information. Then, quantize and encode the compressed image to further compress the amount of data to be transmitted. The quantized and encoded image data is transmitted to the control end via Ethernet. After receiving the data, the control end first decodes and inverse-quantizes it, and then sends the inverse-quantized image data into the image decompression deep neural network based on long and short spatial attention and channel attention to decompress the image data, accurately restore the image, and display the restored server desktop image on the control end display. It has the effects of accurately extracting and retaining the important and effective information of the image, improving the image compression quality and compression ratio, reducing power consumption at the same time, increasing the image compression speed, and reducing the transmission delay.
[0044] According to an embodiment of the present invention, an image transmission embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in components with data processing capabilities, such as: the baseboard management controller, the central processing unit, the mobile terminal, etc. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0045] In this embodiment, an image transmission method is provided, which is applied to an image compression component. Figure 1 It is a flowchart of the image transmission method according to the embodiment of the present invention, as Figure 1 shown, the process includes the following steps:
[0046] Step S101, obtain the image to be transmitted.
[0047] Specifically, the image compression component is, for example, a baseboard management controller having a KVM (Keyboard, Video, Mouse, a server remote management technology) module and a compression module. The baseboard management controller obtains the server desktop image from the server through the KVM module through the PCIe (Peripheral Component Interconnect express, a high-speed serial computer expansion bus standard) interface, and uses the server desktop image as the image to be transmitted. It is necessary to transmit the image to be transmitted to the control end.
[0048] Step S102, perform downsampling on the image to be transmitted, input the downsampling result into an image compression model, and obtain a compressed image. Among them, the image compression model includes a first sub-model, a second sub-model, and a third sub-model. The first sub-model is used to extract the first local spatial feature information of the downsampling result, the second sub-model is used to extract the first global spatial feature information of the downsampling result, and the third sub-model is used to extract the first channel feature information of the downsampling result. The compressed image is obtained by splicing the first local spatial feature information, the first global spatial feature information, and the first channel feature information.
[0049] Specifically, perform downsampling on the image to be transmitted to obtain a downsampling result. For example, 2-fold downsampling can be completed by a pooling layer or a convolutional layer with a stride of 2. The downsampling result is a feature map extracted from the image to be transmitted. When a deep neural network processes an image, the image feature map has a spatial dimension and a channel dimension. In this regard, this embodiment provides an image compression model based on long spatial attention, short spatial attention, and channel attention. Among them, long spatial attention refers to the global feature of the image feature map in the spatial dimension, short spatial attention refers to the local feature of the image feature map in the spatial dimension, and channel attention refers to adding an attention in the channel dimension of the image feature map, such as SENet (Squeeze-and-Excitation Network).
[0050] The image compression model includes a first sub-model, a second sub-model, and a third sub-model. The first sub-model is a convolutional layer created based on a convolutional neural network, and the first sub-model can be used to extract the first local spatial feature information of the downsampling result. The second sub-model is a spatial self-attention model created based on a visual self-attention mechanism model, and the second sub-model is used to extract the first global spatial feature information of the downsampling result. The third sub-model is a channel self-attention model created based on a visual self-attention mechanism model, and the third sub-model is used to extract the first channel feature information of the downsampling result. After the extraction of the feature information is completed, the first local spatial feature information, the first global spatial feature information, and the first channel feature information are concatenated to obtain the compressed image. Local spatial features focus on the detailed information of small regions in an image or data, such as edges, textures, corners, etc. Global spatial features focus on the overall structure of an image or data, such as the shape of an object, the scene layout, the context relationship, etc. Channel features focus on the importance and interaction relationships of different feature channels, and each channel may represent a specific pattern (such as color, texture, or object parts).
[0051] The above process is as Figure 2 shown, obtaining the server desktop image from the server, and image compression based on the long-short spatial attention and channel attention deep neural network. The image compression model and the image decompression model of this embodiment are constructed based on a deep neural network. It is necessary to pre-train the compression neural network of the image compression model and the decompression neural network of the image decompression model. It is necessary to collect a sufficient number and variety of server desktop images, construct a dataset, and use the constructed dataset to train the entire compression neural network and decompression neural network, and then the trained network can be deployed to implement image compression and decompression.
[0052] It should be noted that the image compression model in this embodiment and the method of compressing an image using the image compression model can also be used for various image processing tasks, such as image classification, object detection, semantic segmentation, etc.
[0053] Step S103, sending the compressed image to the control end so that the control end can obtain the image to be transmitted according to the compressed image and the image decompression model, where the image decompression model includes a first sub-model, a second sub-model, and a third sub-model.
[0054] Specifically, after the control end receives the compressed image, it needs to restore it to the image to be transmitted. First, the compressed image is upsampled to obtain the upsampling result. For example, the upsampling is implemented by transposed convolution to restore the original size of the image. The upsampling result is the feature map obtained after restoring the compressed image. The upsampling result is input into the image decompression model, and the structure of the image decompression model is the same as that of the image compression model, both including a first sub-model, a second sub-model, and a third sub-model. The sub-models in the image decompression model are used to extract the local spatial feature information, global spatial feature information, and channel feature information of the upsampling result. These feature information are concatenated, and the concatenated result is upsampled again. The upsampling result of this upsampling is input into the image decompression model, and the above operations are repeated until the image to be transmitted is restored.
[0055] The above process is as Figure 2 shown. After the compressed image is quantized and encoded, it is sent to the control end, and the control end performs decoding and inverse quantization to obtain the compressed image. Image decompression based on the deep neural network of long and short spatial attention and channel attention obtains the image to be transmitted, and the image to be transmitted is displayed at the control end. The control end can use the image to be transmitted to remotely control the server.
[0056] The image transmission method provided in this embodiment downsamples the image to be transmitted, inputs the downsampling result into the image compression model, uses the first sub-model in the image compression model to extract the first local spatial feature information of the downsampling result, uses the second sub-model to extract the first global spatial feature information of the downsampling result, and uses the third sub-model to extract the first channel feature information of the downsampling result. It has the effects of accurately extracting and retaining the important features of the image, improving the image compression quality and compression ratio, reducing the power consumption required for compressing the image, improving the image compression speed, and reducing the transmission delay. It solves the problem that the feature extraction of the image is inaccurate during the image transmission process, which affects the image compression quality.
[0057] In some alternative embodiments, downsampling the image to be transmitted and inputting the downsampling result into the image compression model to obtain the compressed image includes:
[0058] Downsample the image to be transmitted and use the downsampling result as the object to be compressed;
[0059] Divide the object to be compressed into a first sub-feature map, a second sub-feature map, and a third sub-feature map, where the ratio between the number of units of the first sub-feature map, the number of units of the second sub-feature map, and the number of units of the third sub-feature map in the target dimension is a preset ratio;
[0060] Extract the second local spatial feature information of the first sub-feature map using the first sub-model, extract the second global spatial feature information of the second sub-feature map using the second sub-model, and extract the second channel feature information of the third sub-feature map using the third sub-model;
[0061] Concatenate the second local spatial feature information, the second global spatial feature information, and the second channel feature information to obtain a second feature map;
[0062] Rearrange the information of the second feature map in the target dimension to obtain a third feature map;
[0063] Downsample the third feature map to obtain an intermediate downsampling result;
[0064] Take the intermediate downsampling result as the object to be compressed, and start the subsequent steps from dividing the object to be compressed into the first sub-feature map, the second sub-feature map, and the third sub-feature map. When the iteration termination condition is reached, end, and downsample the object to be compressed to obtain the compressed image.
[0065] Specifically, downsample the image to be transmitted, for example: 2-fold downsampling, and take the downsampling result as the object to be compressed. As Figure 3 shown, Feature Map 0 is the downsampling result. The feature map has a spatial dimension and a channel dimension. The spatial dimension is represented by height and width. The height of the feature map is H, the width is W, and the number of channels is C.
[0066] The target dimension is, for example: the spatial dimension or the channel dimension. In this embodiment, the channel dimension is taken as the target dimension for illustration. The image compression model of this embodiment includes a hybrid attention module. The hybrid attention module consists of a first sub-model created based on a CNN deep neural network, a second sub-model created based on a ViT deep neural network, and a third sub-model. This enables the hybrid attention module to have both the convolution of CNN (short attention, local features) and the long spatial attention and global features of ViT. As Figure 4 shown, the image compression process includes multiple repeated 2-fold downsamplings and the hybrid attention module, Figure 4The original image in is the downsampling result. Compared with the spatial self-attention model and the channel self-attention model, the number of parameters and the amount of computation of convolution are the least. In order to improve the inference speed of the network, reduce resource occupancy and power consumption, therefore, the feature maps of half of the channels (C / 2) of the feature map 0 can be processed by convolution, and the remaining half of the channels (C / 4) are used for spatial self-attention and half for channel self-attention (C / 4). There is also an important reason for not using the three-way division here, because in the neural network, the number of channels is usually a power of 2, and most of the channels of the feature maps cannot be divided into three equal parts. Therefore, the division strategy of C / 2, C / 4, C / 4 is adopted. Therefore, the preset ratio can be: 2:1:1, and the preset ratio can be adjusted according to actual needs, and no specific quantity limit is set here.
[0067] Slice the downsampling result with the number of channels C in the channel dimension according to the preset ratio, and divide it into 3 pieces, which are features Figure 1 , features Figure 2 and features Figure 3 . The number of channels of feature Figure 1 is C / 2, the number of channels of feature Figure 2 is C / 4, and the number of channels of feature Figure 3 is C / 4. Take feature Figure 1 as the first sub-feature map, take feature Figure 2 as the second sub-feature map, and take feature Figure 3 as the third sub-feature map. Because the target dimension is the channel dimension at this time, the number of units in the target dimension is the number of channels, and the ratio of the number of channels of the first sub-feature map, the second sub-feature map, and the third sub-feature map is the preset ratio, that is, 2:1:1.
[0068] Use the first sub-model to extract the second local spatial feature information of the first sub-feature map, use the second sub-model to extract the second global spatial feature information of the second sub-feature map, and use the third sub-model to extract the second channel feature information of the third sub-feature map. As Figure 3 shown, the first sub-model uses the convolutional layer to extract the second local spatial feature information of feature Figure 1 . The second sub-model converts feature Figure 2 into a sequence according to the spatial dimension, uses spatial multi-head self-attention to extract the features in the sequence and converts them into a feature map to obtain the second global spatial feature information. The third sub-model converts feature Figure 3 into a sequence according to the channel dimension, uses channel multi-head self-attention to extract the features in the sequence and converts them into a feature map to obtain the third global spatial feature information. It should be noted that the serialized output of the multi-head self-attention module converts the serialized output into the form of a feature map according to the reverse operation of the serialization conversion.
[0069] Concatenate the second local spatial feature information, the second global spatial feature information, and the second channel feature information in the target dimension to obtain a second feature map, which includes local spatial features, global spatial features, and channel features. Since in the next hybrid attention module, the feature map still needs to be partitioned, in order for the feature map to include global spatial features and channel features during the next partitioning, it is necessary to rearrange the information of the second feature map in the channel dimension to obtain a feature Figure 4 as the third feature map. Downsample the third feature map to obtain an intermediate sampling result.
[0070] The iteration termination conditions are, for example: the compression ratio of the object to be compressed relative to the image to be transmitted is less than a first preset value, and the first preset value is, for example: 50%, 60% or other values; the compression processing time for the image to be transmitted is greater than a second preset value, and the second preset value is, for example: 15 seconds, 20 seconds or other values; the number of repetitions of the compression processing for the image to be transmitted is greater than a third preset value, and the third preset value is, for example: 3, 4 or other values.
[0071] Use the intermediate sampling result as the object to be compressed, and re - execute the above process until any of the above iteration termination conditions is reached, then end. Downsample the object to be compressed to obtain the compressed image.
[0072] As Figure 4 shown, the image decompression network corresponds to the compression network and consists of an upsampling layer and a hybrid attention module. Among them, the hybrid attention module is the same as the hybrid attention module of the compression network.
[0073] The above process is as Figure 4 shown. First, downsample the original image by a factor of 2, and use the hybrid attention module in the image compression model to process the downsampling result. Repeat the above process until image compression is completed to obtain the compressed image. Quantize and encode the compressed image and send it to the control end. The control end decodes and inverse - quantizes it to obtain the compressed image. The control end upsamples the compressed image by a factor of 2, and uses the hybrid attention module in the image decompression model to process the upsampling result. Repeat the above process until image decompression is completed to restore the image and obtain the compressed image.
[0074] In this embodiment, dividing the downsampling result into three sub - feature maps and respectively extracting local spatial feature information, global spatial feature information, and channel feature information from the sub - feature maps can reduce the parameters and computational amount of the model, reduce the demand for hardware resources, improve the computational speed, and reduce power consumption. At the same time, it can effectively extract the effective important information of the image and achieve high - quality image compression.
[0075] In some alternative embodiments, the second local spatial feature information of the first sub-feature map is extracted by using a first sub-model, the second global spatial feature information of the second sub-feature map is extracted by using a second sub-model, and the second channel feature information of the third sub-feature map is extracted by using a third sub-model, including:
[0076] Using the first sub-model to perform a convolution operation on the first sub-feature map with a convolution kernel of a preset size to obtain the second local spatial feature information;
[0077] Using the second sub-model to convert the information of the second sub-feature map in the spatial dimension into a first vector, extracting a first intermediate feature from the first vector, and converting the first intermediate feature into the second global spatial feature information in the spatial dimension;
[0078] Using the third sub-model to convert the information of the third sub-feature map in the channel dimension into a second vector, extracting a second intermediate feature from the second vector, and converting the second intermediate feature into the second channel feature information in the channel dimension.
[0079] Specifically, the preset size is, for example, k×k, and the value of k can be 2, 3, or other values. The first sub-feature map is processed by the convolutional layer of the first sub-model, and a k×k convolution kernel performs a convolution operation on the feature map. The convolution operation is mainly used to extract local spatial features to obtain the second local spatial feature information.
[0080] Converting the second sub-feature map into the serialized input form required by the multi-head self-attention module in the second sub-model, including: Since it is necessary to extract the global spatial features of the second sub-feature map, it is necessary to flatten all the channels of each item point in the spatial dimension of the feature map into a vector form to obtain the first vector. The first vector is used as the serialized input of the multi-head self-attention, and then enters the multi-head self-attention module for global spatial feature extraction to output the first intermediate feature. The first intermediate feature is converted into the feature map form according to the reverse operation of the serialized conversion to obtain the second global spatial feature information. The above process is as Figure 5 shown, taking the feature Figure 5 as the second sub-feature map as an example, taking the feature Figure 5Convert it into a sequence according to the spatial dimension, and input the sequence into the multi-head self-attention to obtain the second global spatial feature information. The sequence includes three inputs: Q, K, and V. The multi-head self-attention uses matrix multiplication (MatMul) to process Q and K to obtain the first processing result. The matrix multiplication is used to calculate the association strength between each position in the sequence and other positions; perform a scaling operation (Scale) on the first processing result to obtain the second processing result. The scaling operation is used to avoid the vanishing gradient caused by the increase in dimension of the dot product result; perform a masking operation (Optional) on the second processing result to obtain the third processing result; perform a normalization process on the third processing result. For example, normalize each row of the third processing result with Softmax to obtain the fourth processing result. Perform a weighted sum of the fourth processing result and the input V to obtain the final output, that is, the second global spatial feature information.
[0081] Convert the third sub-feature map into the serialized input form required by the multi-head self-attention module in the third sub-model, including: Since it is necessary to extract the global channel dimension features of the third sub-feature map, it is necessary to flatten all the pixel points of each channel of the feature map into a vector form to obtain the second vector. Use the second vector as the serialized input of the multi-head self-attention, and then enter the multi-head self-attention module for global channel feature extraction, and output the second intermediate feature. Convert the serialized output into a feature map form according to the reverse operation of the serialized conversion to obtain the second channel feature information. The above process is as Figure 5 shown, taking the feature Figure 6 as the third sub-feature map as an example, the feature Figure 6 Convert it into a sequence according to the channel dimension, and input the sequence into the multi-head self-attention to obtain the second channel feature information. The specific process is as described above and will not be repeated here.
[0082] In addition, after creating the hybrid attention module in the image compression model and the image decompression model, it is necessary to train the hybrid attention module, specifically including: First, it is necessary to construct a data set. The present invention is mainly aimed at the images on the server system desktop, so it is necessary to collect enough server desktop images, and more types of pictures need to be added when constructing the data set. For example: The training data set collected 20,000 common images in the server desktop system, and the test set collected 10,000 images, including text content images, web page content images, information monitoring content images, etc., ensuring the diversity of the data set. Use the constructed data set to train the neural network in the hybrid attention module on a deep learning server. During the training process, continuously test the image compression performance of the network on the test set. The compression quality evaluation indicators can be mean square error, peak signal-to-noise ratio, structural similarity, etc. Stop training when the compression quality target is reached and save the weights.
[0083] In this embodiment, the first sub-model is used to extract the second local spatial feature information, the second sub-model is used to extract the second global spatial feature information, and the third sub-model is used to extract the second channel feature information, so as to more accurately extract and retain the important and effective information of the image, thereby improving the image compression quality and compression ratio, reducing the power consumption at the same time, and increasing the image compression speed.
[0084] In some alternative embodiments, sending the compressed image to the control end includes:
[0085] Determine the product of the data in the compressed image and a preset quantization coefficient, and round the product to obtain intermediate data;
[0086] Encode the intermediate data using a preset encoding method to obtain encoded data, where the encoded data contains the information of the compressed image;
[0087] Transmit the encoded data to the control end so that the control end can obtain the compressed image according to the encoded data.
[0088] Specifically, first, perform a quantization operation on the data in the compressed image. Quantization is used to multiply a numerical value by a quantization coefficient and then round it to make all numerical values become integers, which is convenient for candidate encoding.
[0089] The preset quantization coefficient is, for example: S, and S can be 2, 3 or other values. Calculate the product of the floating-point numbers and other data in the compressed image and the preset quantization coefficient respectively, as shown in formula (1). Round the product to obtain intermediate data. For example: perform a rounding operation on the product to obtain the corresponding integer as the intermediate data.
[0090] Q = round(F × S) (1)
[0091] Where F is a floating-point number, S is the quantization coefficient, round is the rounding operation, and Q is the intermediate data. After formula (1), the floating-point number is quantized into an integer, which is conducive to encoding.
[0092] The preset encoding method is, for example: using Huffman coding as the entropy coding method, arithmetic coding method, Shannon-Fano coding method, etc. Encode the intermediate data using the preset encoding method to obtain encoded data, where the encoded data contains the information of the compressed image. Transmit the encoded data to the control end so that the control end can obtain the compressed image according to the encoded data. For example: as Figure 2 shown, after performing data quantization and encoding operations on the compressed image, it is sent to the control end, and the control end performs decoding and inverse quantization operations to obtain the compressed image.
[0093] In this embodiment, during the process of transmitting the compressed image, the compressed image is quantized and encoded, the data of the compressed image is streamlined, the resources occupied during the transmission process are reduced, and the data is encoded to ensure that the data can be correctly transmitted.
[0094] In this embodiment, an image transmission method is provided, and this method is applied to the control end. Figure 6 It is a flowchart of the image transmission method according to the embodiment of the present invention, as Figure 6 shown, and this process includes the following steps:
[0095] Step S601, obtain the compressed image from the image compression component, where the compressed image is obtained by splicing the first local spatial feature information, the first global spatial feature information, and the first channel feature information. The first local spatial feature information is extracted from the downsampling result by the first sub-model, the first global spatial feature information is extracted from the downsampling result by the second sub-model, the first channel feature information is extracted from the downsampling result by the third sub-model. The first sub-model, the second sub-model, and the third sub-model are included in the image compression model, and the downsampling result is obtained by downsampling the image to be transmitted.
[0096] Specifically, the image compression component obtains the server desktop image from the server through the PCIe interface as the image to be transmitted. The first sub-model is a convolutional layer created based on the convolutional neural network, and the first local spatial feature information of the downsampling result can be extracted by using the first sub-model. The second sub-model is a spatial self-attention model created based on the visual self-attention mechanism model, and the second sub-model is used to extract the first global spatial feature information of the downsampling result. The third sub-model is a channel self-attention model created based on the visual self-attention mechanism model, and the third sub-model is used to extract the first channel feature information of the downsampling result. After the extraction of the feature information is completed, the first local spatial feature information, the first global spatial feature information, and the first channel feature information are spliced to obtain the compressed image. The image compression component transmits the compressed image to the control end, and the control end obtains the compressed image from the image compression component.
[0097] Step S602, upsample the compressed image, and input the upsampling result into the image decompression model to obtain the image to be transmitted. The image decompression model is used to generate the image to be transmitted according to the local spatial feature information, the global spatial feature information, and the channel feature information of the upsampling result. The image decompression model includes the first sub-model, the second sub-model, and the third sub-model.
[0098] Specifically, after receiving the compressed image, the control end needs to restore it to the image to be transmitted. First, the compressed image is upsampled to obtain the upsampling result. For example, the upsampling is implemented by transposed convolution and is used to restore the original size of the image. The upsampling result is the feature map obtained after restoring the compressed image.
[0099] The upsampling result is input into the image decompression model. The structure of the image decompression model is the same as that of the image compression model, both including a first sub-model, a second sub-model, and a third sub-model. The sub-models in the image decompression model are used to extract the local spatial feature information, global spatial feature information, and channel feature information of the upsampling result. These feature information are spliced, and the splicing result is upsampled again. The upsampling result of this upsampling is input into the image decompression model, and the above operations are repeated until the image to be transmitted is restored.
[0100] In the image transmission method provided in this embodiment, the image compression component uses the image compression model to extract the local spatial information, global spatial information, and channel information of the image to be transmitted, and completes the compression of the image to be transmitted. The control end receives the compressed image and uses the image decompression model to decompress the compressed image, accurately restores the image, and obtains the image to be transmitted. This enables the image compression process to retain the important and effective information of the image, improve the image compression quality and compression ratio, reduce power consumption at the same time, increase the image compression speed, and reduce the image transmission delay. It solves the problem that the feature extraction of the image is inaccurate during the image transmission process, which affects the image compression quality.
[0101] In some alternative embodiments, obtaining the compressed image includes:
[0102] Obtain the encoded data from the image compression component, where the encoded data is obtained by the image compression component encoding the intermediate data using a preset encoding method. The intermediate data is obtained by the image compression component rounding the product, and the product is the product of the data in the compressed image and the preset quantization coefficient;
[0103] Decode the encoded data using a preset decoding method to obtain the decoded data;
[0104] Determine the ratio of the decoded data to the preset quantization coefficient, and obtain the compressed image according to the ratio.
[0105] Specifically, the image compression component first performs quantization operation on the data in the compressed image and then performs encoding operation to obtain the encoded data. The specific process can refer to the above embodiments and will not be elaborated here. The image compression component sends the encoded data to the control end.
[0106] In this embodiment, in the decoding and inverse quantization stages, the inverse quantization and decoding are performed using the inverse quantization and decoding methods corresponding to the quantization and encoding parts. Preset encoding methods include, for example, using Huffman coding as the entropy coding method, arithmetic coding method, Shannon-Fano coding method, etc. The preset decoding method is to decode the encoded data according to the preset decoding method.
[0107] After the control end receives the encoded data, it decodes the encoded data using the preset decoding method to obtain the decoded data. Perform inverse quantization processing on the decoded data. As shown in formula (2), determine the ratio of the decoded data to the preset quantization coefficient, and obtain the compressed image according to the ratio.
[0108]
[0109] Among them, S is the same quantization coefficient as S in formula (1).
[0110] In some alternative embodiments, since the server desktop can be divided into multiple partitions and it is not necessarily required to manage all partition data for the server, the specific process of obtaining the image to be transmitted may further include steps A1 to A4.
[0111] Step A1, divide the server desktop.
[0112] Specifically, according to the network service quality information and the image content of the desktop image, perform block division on the desktop image to obtain multiple blocks, such as: block 1, block 2, block 3, and block 4.
[0113] Step A2, receive the image requirements sent by the control end.
[0114] Specifically, the image requirements may specify the time period of image sampling, the sampling interval, and the block where the required image is located.
[0115] Step A3, obtain the image and image information corresponding to the image requirements.
[0116] Specifically, the image requirements are, for example, the images of block 1 and block 2 from 8 o'clock to 9 o'clock, and the sampling interval is 10 minutes. Obtain the corresponding images according to the image requirements, and record the image information of the images, such as: sampling time, block information. According to the image information, the images of block 1 and block 2 at the same sampling time can be spliced together.
[0117] Step A4, compress the image in step A3 using the image compression model, and send the compressed image and image information to the control end.
[0118] In this embodiment, the server desktop is divided into multiple blocks, so that the control end can only obtain the picture of a certain block, reducing the amount of data for image transmission and image compression and improving the transmission efficiency.
[0119] In this embodiment, an image compression component is provided. The image compression component includes: a compression module, a storage module, and a controller; the compression module includes an image compression model.
[0120] The compression module is used to obtain the image to be transmitted, downsample the image to be transmitted, input the downsampling result into the image compression model to obtain the compressed image. Among them, the image compression model includes a first sub-model, a second sub-model, and a third sub-model. The first sub-model is used to extract the first local spatial feature information of the downsampling result, the second sub-model is used to extract the first global spatial feature information of the downsampling result, and the third sub-model is used to extract the first channel feature information of the downsampling result. The compressed image is obtained by splicing the first local spatial feature information, the first global spatial feature information, and the first channel feature information.
[0121] The compression module is used to store the compressed image in the storage module; the controller is used to read the compressed image from the storage module and send the compressed image to the control end, so that the control end can obtain the image to be transmitted according to the compressed image and the image decompression model. Among them, the image decompression model includes a first sub-model, a second sub-model, and a third sub-model.
[0122] Specifically, as Figure 7 shown, the image compression component is a baseboard management controller (BMC) including multiple functional modules. The image compression component includes: a compression module, a DDR (memory) as the storage module, and an Ethernet controller as the controller.
[0123] In this actual example, the image compression is implemented by a deep neural network, and the amount of calculation is large. A dedicated hardware module for neural network calculation, that is, the compression module, needs to be designed in the BMC. Moreover, the intermediate data generated during the compression process and the compressed image data are stored in the external DDR of the BMC system. This DDR is the memory space used by the BMC system.
[0124] The compression module is used to obtain the image to be transmitted, downsample the image to be transmitted, input the downsampling result into the image compression model to obtain the compressed image. For the specific process, refer to the above embodiment and will not be elaborated here. The compression module stores the compressed image in the storage module.
[0125] After quantization and encoding, the compressed image is read from the DDR through an Ethernet controller and transmitted to the control end via Ethernet. After receiving the compressed data, the control end performs inverse quantization and decoding on it, and then decompresses it using an image decompression model and displays it on the screen, thus realizing the remote display and control functions.
[0126] In addition, the image compression component further includes: a PCIe controller, a VGA IP (Video Graphics Array IP), a DDR controller, a quantization and encoding module, and a CPU (Central Processing Unit). In actual application, the compression module and the quantization and encoding module are deployed in the BMC. The host of the server is connected to the baseboard management controller. The BMC obtains the server desktop image (video / image data) from the server host through a PCIe link, and then transfers the server desktop image to the VGA IP. The VGA IP processes the video / image data (such as 2D acceleration, rotation, etc.) and stores it in the DDR through the DDR controller, or directly transfers it to the compression module for image compression.
[0127] In this embodiment, a dedicated image compression computing module is designed in the image compression component, which can achieve fast compression of images. The compression module uses the first sub-model to extract the first local spatial feature information of the downsampling result, uses the second sub-model to extract the first global spatial feature information of the downsampling result, and uses the third sub-model to extract the first channel feature information of the downsampling result. It realizes accurate extraction and retention of important features of the image, improves the image compression quality and compression ratio, reduces the power consumption required for compressing the image at the same time, improves the image compression speed, and reduces the transmission delay. It solves the problem that the feature extraction of the image is inaccurate during the image transmission process, which affects the image compression quality.
[0128] In some alternative embodiments, the compression module includes: a data rearrangement unit, a storage unit, a matrix calculation unit, and a data recovery unit;
[0129] The data rearrangement unit is used to convert the data of the image to be transmitted into a preset data form to obtain the converted data;
[0130] The storage unit is used to store the converted data and model parameters;
[0131] The matrix calculation unit is used to obtain the converted data and model parameters from the storage unit, extract the first local spatial feature information from the converted data according to the model parameters and the first sub-model, extract the first global spatial feature information from the converted data according to the model parameters and the second sub-model, and extract the first channel feature information from the converted data according to the model parameters and the third sub-model;
[0132] The data recovery unit is used to convert the first local spatial feature information, the first global spatial feature information, and the first channel feature information into a compressed image;
[0133] The storage unit is further used to store the compressed image and transmit the compressed image to the storage module.
[0134] Specifically, the compression module is the core for image compression and is also the main feature that differentiates this embodiment from traditional image compression techniques. As Figure 8 shown, the compression module includes: a data rearrangement unit, a storage unit for on-chip storage, a matrix calculation unit, and a data recovery unit. The matrix calculation unit is a systolic array and is used to calculate the long and short spatial attention and channel attention.
[0135] To complete the entire image compression process, the cooperation of the CPU and the DDR is required. The CPU is responsible for controlling the entire inference process. The image data is first read from the off-chip storage of the DDR to the data rearrangement unit. The data rearrangement unit is used to convert the data of the image to be transmitted into a preset data form to obtain the converted data. The preset data form is the data form required by the matrix calculation unit, such as general matrix multiplication. The converted data is stored in the storage unit. Model parameters such as neural network weights can be read from the DDR to the on-chip storage unit or directly stored in the storage unit. Therefore, the storage unit is used to store the converted data and model parameters.
[0136] After completing the data rearrangement, the matrix calculation unit obtains the converted data and model parameters from the storage unit and enters the matrix calculation unit to calculate the long and short spatial attention and channel attention respectively, including: extracting the first local spatial feature information from the converted data according to the model parameters and the first sub-model, extracting the first global spatial feature information from the converted data according to the model parameters and the second sub-model, and extracting the first channel feature information from the converted data according to the model parameters and the third sub-model.
[0137] The data recovery unit needs to convert the feature information into the original feature map form, including: converting the first local spatial feature information, the first global spatial feature information, and the first channel feature information into a compressed image. The compressed image is temporarily stored in the storage unit and then stored in the DDR, or directly stored in the DDR. This process is looped for each layer for calculation. After all layers of calculation are completed, the compressed feature map is stored in the DDR.
[0138] In this embodiment, an image transmission device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0139] This embodiment provides an image transmission device, which is deployed in an image compression component, such as Figure 9 shown, and includes: a first image acquisition module 901 for acquiring an image to be transmitted; an image compression module 902 for downsampling the image to be transmitted, inputting the downsampling result into an image compression model to obtain a compressed image, wherein the image compression model includes a first sub-model, a second sub-model, and a third sub-model. The first sub-model is used to extract the first local spatial feature information of the downsampling result, the second sub-model is used to extract the first global spatial feature information of the downsampling result, and the third sub-model is used to extract the first channel feature information of the downsampling result. The compressed image is obtained by splicing the first local spatial feature information, the first global spatial feature information, and the first channel feature information; an image sending module 903 for sending the compressed image to a control end so that the control end can obtain the image to be transmitted according to the compressed image and an image decompression model, and the image decompression model includes a first sub-model, a second sub-model, and a third sub-model.
[0140] In some alternative embodiments, the image compression module 902 includes: a first sampling unit for downsampling the image to be transmitted and using the downsampling result as an object to be compressed; a partitioning unit for partitioning the object to be compressed into a first sub-feature map, a second sub-feature map, and a third sub-feature map, wherein the ratio among the number of units of the first sub-feature map in the target dimension, the number of units of the second sub-feature map in the target dimension, and the number of units of the third sub-feature map in the target dimension is a preset ratio; a feature extraction unit for using the first sub-model to extract the second local spatial feature information of the first sub-feature map, using the second sub-model to extract the second global spatial feature information of the second sub-feature map, and using the third sub-model to extract the second channel feature information of the third sub-feature map; a splicing unit for splicing the second local spatial feature information, the second global spatial feature information, and the second channel feature information to obtain a second feature map; an arranging unit for rearranging the information of the second feature map in the target dimension to obtain a third feature map; a second sampling unit for downsampling the third feature map to obtain an intermediate sampling result; a looping unit for using the intermediate sampling result as the object to be compressed and starting to execute subsequent steps from partitioning the object to be compressed into a first sub-feature map, a second sub-feature map, and a third sub-feature map, and ending until the iteration termination condition is reached, and then downsampling the object to be compressed to obtain a compressed image.
[0141] In some alternative embodiments, the feature extraction unit includes: a first feature extraction sub-module, configured to perform a convolution operation on a first sub-feature map by using a convolution kernel of a preset size with a first sub-model to obtain second local spatial feature information; a second feature extraction sub-module, configured to use a second sub-model to convert the information of a second sub-feature map in the spatial dimension into a first vector, extract a first intermediate feature from the first vector, and convert the first intermediate feature into second global spatial feature information in the spatial dimension; a third feature extraction sub-module, configured to use a third sub-model to convert the information of a third sub-feature map in the channel dimension into a second vector, extract a second intermediate feature from the second vector, and convert the second intermediate feature into second channel feature information in the channel dimension.
[0142] In some alternative embodiments, the image sending module 903 includes: a first data processing unit, configured to determine a product of data in the compressed image and a preset quantization coefficient, and round the product to obtain intermediate data; an encoding unit, configured to perform encoding processing on the intermediate data by using a preset encoding method to obtain encoded data, where the encoded data includes information of the compressed image; a sending unit, configured to transmit the encoded data to a control end so that the control end obtains the compressed image according to the encoded data.
[0143] This embodiment provides an image transmission device, which is deployed at a control end, as Figure 10 shown, and includes: a second image acquisition module 1001, configured to acquire a compressed image from an image compression component, where the compressed image is obtained by splicing first local spatial feature information, first global spatial feature information, and first channel feature information, the first local spatial feature information is extracted by a first sub-model from a downsampling result, the first global spatial feature information is extracted by a second sub-model from the downsampling result, the first channel feature information is extracted by a third sub-model from the downsampling result, the first sub-model, the second sub-model, and the third sub-model are included in an image compression model, and the downsampling result is obtained by performing downsampling on an image to be transmitted; an image decompression module 1002, configured to perform upsampling on the compressed image, input the upsampling result into an image decompression model to obtain the image to be transmitted, where the image decompression model is configured to generate the image to be transmitted according to local spatial feature information, global spatial feature information, and channel feature information of the upsampling result, and the image decompression model includes a first sub-model, a second sub-model, and a third sub-model.
[0144] In some alternative embodiments, the second image acquisition module 1001 includes: an acquisition unit configured to acquire encoded data from an image compression component, where the encoded data is obtained by the image compression component encoding intermediate data using a preset encoding method, the intermediate data is obtained by the image compression component rounding a product, and the product is the product of data in the compressed image and a preset quantization coefficient; a decoding unit configured to perform decoding processing on the encoded data using a preset decoding method to obtain decoded data; and a second data processing unit configured to determine a ratio of the decoded data to the preset quantization coefficient and obtain the compressed image based on the ratio.
[0145] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding foregoing embodiments, and will not be elaborated herein.
[0146] The image transmission device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0147] An embodiment of the present invention further provides a computer device having the above-mentioned Figure 9 and Figure 10 shown image transmission device.
[0148] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As shown in Figure 11 , the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system). Figure 11 In
[0149] The processor 10 may be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device may be a complex programmable logic device, a field-programmable gate array, a generic array logic, or any combination thereof.
[0150] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.
[0151] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may further include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0152] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above-mentioned types of memories.
[0153] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks.
[0154] The embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention may be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-temporary machine-readable storage medium and to be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium may be a magnetic disk, an optical disc, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium may further include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0155] A part of the present invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can call or provide the methods and / or technical solutions according to the present invention through the operations of the computer. Those skilled in the art should understand that the forms in which computer program instructions exist in a computer-readable medium include but are not limited to source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0156] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the present invention.
Claims
1. An image transmission method, characterized in that: The method is applied to an image compression component, and the method comprises: Acquire the image to be transmitted; Downsampling the image to be transmitted, and inputting the downsampling result into an image compression model to obtain a compressed image, wherein the image compression model includes a first sub-model, a second sub-model, and a third sub-model, the first sub-model is used to extract first local spatial feature information of the downsampling result, the second sub-model is used to extract first global spatial feature information of the downsampling result, and the third sub-model is used to extract first channel feature information of the downsampling result, and the compressed image is obtained by splicing the first local spatial feature information, the first global spatial feature information, and the first channel feature information; The compressed image is sent to a control end so that the control end obtains the image to be transmitted according to the compressed image and an image decompression model, wherein the image decompression model includes the first sub-model, the second sub-model and the third sub-model.
2. The method according to claim 1, characterized in that The downsampling of the image to be transmitted and inputting the downsampling result into an image compression model to obtain a compressed image includes: Down-sampling the image to be transmitted, and using the down-sampling result as the object to be compressed; Dividing the object to be compressed into a first sub-feature map, a second sub-feature map, and a third sub-feature map, wherein a ratio among the number of units of the first sub-feature map on the target dimension, the number of units of the second sub-feature map on the target dimension, and the number of units of the third sub-feature map on the target dimension is a preset ratio; Extracting second local spatial feature information of the first sub-feature map using the first sub-model, extracting second global spatial feature information of the second sub-feature map using the second sub-model, and extracting second channel feature information of the third sub-feature map using the third sub-model; splicing the second local space feature information, the second global space feature information and the second channel feature information to obtain a second feature map; Rearranging information of the second feature map on the target dimension to obtain a third feature map; Downsampling the third feature map to obtain an intermediate sampling result; The intermediate sampling result is used as the object to be compressed, and subsequent steps are performed starting from dividing the object to be compressed into a first sub-feature map, a second sub-feature map, and a third sub-feature map until an iteration termination condition is reached, then the object to be compressed is down-sampled to obtain the compressed image.
3. The method according to claim 2, characterized in that The step of extracting second local spatial feature information of the first sub-feature map by using the first sub-model, extracting second global spatial feature information of the second sub-feature map by using the second sub-model, and extracting second channel feature information of the third sub-feature map by using the third sub-model includes: Using the first sub-model to perform a convolution operation on the first sub-feature map using a convolution kernel of a preset size to obtain the second local spatial feature information; Using the second sub-model, converting information of the second sub-feature map in the spatial dimension into a first vector, extracting a first intermediate feature from the first vector, and converting the first intermediate feature into the second global spatial feature information in the spatial dimension; The third sub-model is used to convert information of the third sub-feature map in the channel dimension into a second vector, a second intermediate feature is extracted from the second vector, and the second intermediate feature is converted into the second channel feature information in the channel dimension.
4. The method according to claim 3, characterized in that The step of sending the compressed image to the control terminal comprises: Determining the product of the data in the compressed image and a preset quantization coefficient, and rounding the product to obtain intermediate data; Encoding the intermediate data using a preset encoding method to obtain encoded data, wherein the encoded data includes information of the compressed image; The encoded data is transmitted to the control end, so that the control end obtains the compressed image according to the encoded data.
5. An image transmission method, characterized in that: The method is applied to a control end, and the method comprises: Acquire a compressed image from an image compression component, wherein the compressed image is obtained by splicing first local spatial feature information, first global spatial feature information, and first channel feature information, the first local spatial feature information is extracted from a downsampling result by a first sub-model, the first global spatial feature information is extracted from the downsampling result by a second sub-model, the first channel feature information is extracted from the downsampling result by a third sub-model, the first sub-model, the second sub-model, and the third sub-model are included in an image compression model, and the downsampling result is obtained by downsampling the image to be transmitted; The compressed image is upsampled, and the upsampling result is input into an image decompression model to obtain the image to be transmitted, wherein the image decompression model is used to generate the image to be transmitted according to local spatial feature information, global spatial feature information and channel feature information of the upsampling result, and the image decompression model includes the first sub-model, the second sub-model and the third sub-model.
6. The method according to claim 5, characterized in that The step of obtaining the compressed image comprises: Acquire coded data from the image compression component, wherein the coded data is obtained after the image compression component encodes the intermediate data using a preset encoding method, and the intermediate data is obtained after the image compression component rounds the product, and the product is the product of the data in the compressed image and a preset quantization coefficient; Decoding the encoded data using a preset decoding method to obtain decoded data; A ratio of the decoded data to the preset quantization coefficient is determined, and the compressed image is obtained according to the ratio.
7. An image compression component, characterized in that: The image compression component includes: a compression module, a storage module and a controller; the compression module includes an image compression model; The compression module is used to obtain an image to be transmitted, downsample the image to be transmitted, and input the downsampled result into an image compression model to obtain a compressed image, wherein the image compression model includes a first sub-model, a second sub-model and a third sub-model, the first sub-model is used to extract first local spatial feature information of the downsampled result, the second sub-model is used to extract first global spatial feature information of the downsampled result, and the third sub-model is used to extract first channel feature information of the downsampled result, and the compressed image is obtained by splicing the first local spatial feature information, the first global spatial feature information and the first channel feature information; The compression module is used to store the compressed image in the storage module; The controller is used to read the compressed image from the storage module and send the compressed image to the control end, so that the control end obtains the image to be transmitted based on the compressed image and the image decompression model, wherein the image decompression model includes the first sub-model, the second sub-model and the third sub-model.
8. The image compression component according to claim 7, characterized in that: The compression module includes: a data rearrangement unit, a storage unit, a matrix calculation unit and a data recovery unit; The data rearrangement unit is used to convert the data of the image to be transmitted into a preset data format to obtain converted data; The storage unit is used to store the converted data and model parameters; The matrix calculation unit is used to obtain the converted data and the model parameters from the storage unit, extract the first local space feature information from the converted data according to the model parameters and the first sub-model, extract the first global space feature information from the converted data according to the model parameters and the second sub-model, and extract the first channel feature information from the converted data according to the model parameters and the third sub-model; The data recovery unit is used to convert the first local space feature information, the first global space feature information and the first channel feature information into the compressed image; The storage unit is also used to store the compressed image and transmit the compressed image to the storage module.
9. An image transmission device, characterized in that: The device is deployed in an image compression component, and the device includes: A first image acquisition module, used to acquire an image to be transmitted; An image compression module, used for downsampling the image to be transmitted, inputting the downsampling result into an image compression model, and obtaining a compressed image, wherein the image compression model includes a first sub-model, a second sub-model, and a third sub-model, the first sub-model is used to extract first local spatial feature information of the downsampling result, the second sub-model is used to extract first global spatial feature information of the downsampling result, the third sub-model is used to extract first channel feature information of the downsampling result, and the compressed image is obtained by splicing the first local spatial feature information, the first global spatial feature information, and the first channel feature information; An image sending module is used to send the compressed image to the control end so that the control end obtains the image to be transmitted according to the compressed image and the image decompression model, wherein the image decompression model includes the first sub-model, the second sub-model and the third sub-model.
10. An image transmission device, characterized in that: The device is deployed at a control end, and the device includes: A second image acquisition module is used to acquire a compressed image from the image compression component, wherein the compressed image is obtained by splicing the first local spatial feature information, the first global spatial feature information and the first channel feature information, the first local spatial feature information is extracted by the first sub-model from the downsampling result, the first global spatial feature information is extracted by the second sub-model from the downsampling result, the first channel feature information is extracted by the third sub-model from the downsampling result, the first sub-model, the second sub-model and the third sub-model are included in the image compression model, and the downsampling result is obtained by downsampling the image to be transmitted; An image decompression module is used to upsample the compressed image and input the upsampling result into an image decompression model to obtain the image to be transmitted, wherein the image decompression model is used to generate the image to be transmitted based on the local spatial feature information, global spatial feature information and channel feature information of the upsampling result, and the image decompression model includes the first sub-model, the second sub-model and the third sub-model.
11. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the image transmission method according to any one of claims 1 to 6 by executing the computer instructions.