Panoramic video transmission method, panoramic video transmission device and electronic equipment

By using semantic extraction and weight attention processing technology in the immersive communication in the 6G era, panoramic videos are encoded and transmitted, and the problems of large transmission bandwidth and low user experience quality in the existing technology are solved, and efficient and low-latency panoramic video transmission is achieved.

CN120238673APending Publication Date: 2025-07-01BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311845324.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In immersive communication in the 6G era, it is difficult for the prior art to effectively reduce the bandwidth required for panoramic video transmission while ensuring low latency and high-quality video transmission, especially when the user's perspective changes rapidly, resulting in a decline in user experience quality.

Method used

By performing semantic extraction on the sending end, the semantic features of the video are obtained, and the weighted attention process is used to generate the weighted attention feature map. The transmission rate of feature points is determined based on the entropy model and the latitude adaptive model, and then coded and sent. The receiver obtains high-quality panoramic video through decoding and semantic recovery.

Benefits of technology

With less bandwidth consumption, effectively restore panoramic video frames, improve users' immersive experience quality, reduce information redundancy in panoramic video during transmission, and realize efficient transmission of panoramic video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238673A_ABST
    Figure CN120238673A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of communication, artificial intelligence and the like, in particular to a panoramic video transmission method, a panoramic video transmission device and electronic equipment. The specific implementation scheme is as follows: a sending end performs semantic extraction on a to-be-transmitted first panoramic video to obtain corresponding semantic features; performing weight attention processing on the semantic features through the weight map to obtain a weight attention feature map; determining the transmission rate of each feature point in the weight attention feature map according to an entropy model and a latitude adaptive model; and encoding the weight attention feature map according to the transmission rate to obtain semantic encoding information, decoding the semantic encoding information by the receiving end, and performing semantic recovery to obtain a second panoramic video. Under the condition of consuming less bandwidth, the panoramic video frame is better recovered, and the immersive experience quality of the user is improved; the information redundancy of the panoramic video in the transmission process is reduced, and the efficient transmission of the panoramic video is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to technical fields such as communications and artificial intelligence, and in particular to a panoramic video transmission method, a panoramic video transmission device, and an electronic device. Background Art

[0002] In the vision of the 6G era, immersive communication is a core application scenario that allows users to experience 360-degree panoramic videos and interact with the virtual world. In order to ensure the quality of the user's immersive experience, the entire transmission system is required to provide low latency of less than 20 milliseconds and high-quality video transmission services. However, due to the limitations of the user's actual field of view, even for 4K panoramic videos, the actual visible resolution of the user is only about 960×540. This requires the system to not only transmit high-resolution videos, but also ensure the user experience, which also means a substantial increase in data transmission volume.

[0003] In order to improve the transmission efficiency of panoramic videos and enhance the quality of users' immersive experience, the first existing solution proposes a dynamic video transmission strategy that can adjust the transmission content in real time according to the changes in the user's viewing angle. This solution divides the video frames watched by the user into multiple continuous viewing areas. When the user's viewing direction changes, the system can respond immediately and selectively transmit high-resolution video clips within the user's current viewing angle, rather than the entire panoramic video, effectively reducing the demand for network bandwidth while ensuring the continuity and clarity of video viewing. However, this solution only selectively transmits the entire panoramic video frame based on the user's field of view in combination with the characteristics of panoramic videos. However, when the user's field of view changes too quickly, the area that cannot be transmitted in the future will also be displayed in front of the user, which will greatly reduce the quality of the user's immersive experience.

[0004] The second existing solution adds a response mechanism for viewing angle switching based on the first solution. The system can determine the magnitude of the user's viewing angle change and decide whether to request a new panoramic video file or continue to transmit the auxiliary FoV (field of view) video file under the current viewing angle. This technology not only further reduces the delay that may be caused by the change in viewing angle, but also improves the transmission efficiency and the quality of user experience. However, with the increase in user demand, the resolution of panoramic videos required by video transmission services is getting higher and higher. At the same time, when the user's viewing angle changes too quickly, it is not enough to just transmit the video picture within the field of view. Selectively transmitting video information according to the magnitude of the user's viewing angle change temporarily solves the problem of too fast viewing angle switching, but with the advancement of society and technology, users are increasingly in need of higher-definition videos. At the same time, if the viewing angle switches too much, the user's experience quality will also decrease. These bit-level error-free transmission methods are no longer suitable for the transmission of massive data such as high-resolution panoramic videos. Summary of the invention

[0005] The present disclosure provides a panoramic video transmission method, a panoramic video transmission device, an electronic device, and a storage medium.

[0006] According to a first aspect of the present disclosure, there is provided a panoramic video transmission method, including:

[0007] The sending end performs semantic extraction on a first panoramic video to be transmitted to obtain corresponding semantic features;

[0008] The sending end performs weighted attention processing on the semantic features through a weight map to obtain a weighted attention feature map;

[0009] The sending end determines the transmission rate of each feature point in the weighted attention feature map according to an entropy model and a latitude adaptation model;

[0010] The sending end encodes the weighted attention feature map according to the transmission rate to obtain semantic encoding information;

[0011] The sending end sends the semantic encoding information to a receiving end; wherein, the semantic encoding information is used to be provided to the receiving end to obtain a second panoramic video through decoding and semantic restoration.

[0012] According to a second aspect of the present disclosure, there is provided a panoramic video transmission method, including:

[0013] The receiving end receives semantic encoding information; wherein, the semantic encoding information is obtained by the sending end performing semantic extraction on a first panoramic video to be transmitted to obtain semantic features, performing weighted attention processing on the semantic features through a weight map, determining the transmission rate of each feature point in the weighted attention feature map according to an entropy model and a latitude adaptation model, and then encoding the weighted attention feature map according to the transmission rate;

[0014] The receiving end decodes the semantic encoding information to obtain a third semantic feature;

[0015] The receiving end performs semantic restoration on the third semantic feature to obtain a second panoramic video.

[0016] According to a third aspect of the present disclosure, there is provided a panoramic video transmission device, including:

[0017] A semantic extraction module configured to perform semantic extraction on a first panoramic video to be transmitted to obtain corresponding semantic features;

[0018] A weighted attention module configured to perform weighted attention processing on the semantic features through a weight map to obtain a weighted attention feature map;

[0019] A rate allocation module, configured to determine the transmission rate of each feature point in the weighted attention feature map according to an entropy model and a latitude adaptation model;

[0020] An encoding module, configured to encode the weighted attention feature map according to the transmission rate to obtain semantic encoding information;

[0021] A transmission module, configured to send the semantic encoding information to a receiving end; wherein, the semantic encoding information is used to be provided to the receiving end to obtain a second panoramic video through decoding and semantic restoration.

[0022] According to a fourth aspect of the present disclosure, there is provided a panoramic video transmission device, including:

[0023] A receiving module, configured to receive semantic encoding information; wherein, the semantic encoding information is obtained by a sending end performing semantic extraction on a first panoramic video to be transmitted to obtain semantic features, performing weighted attention processing on the semantic features through a weight map, determining the transmission rate of each feature point in the weighted attention feature map according to an entropy model and a latitude adaptation model, and then encoding the weighted attention feature map according to the transmission rate;

[0024] A decoding module, configured to decode the semantic encoding information to obtain a third semantic feature;

[0025] A semantic restoration module, configured to perform semantic restoration on the third semantic feature to obtain a second panoramic video.

[0026] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0027] At least one processor; and

[0028] A memory communicatively connected to the at least one processor; wherein,

[0029] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of the above technical solutions.

[0030] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to any one of the above technical solutions.

[0031] According to a seventh aspect of the present disclosure, there is provided a computer program product, including a computer program, and the computer program implements the method according to any one of the above technical solutions when executed by a processor.

[0032] The present disclosure provides a panoramic video transmission method, a panoramic video transmission device, an electronic device, and a storage medium. The weight attention module can associate the implicit information existing between the weight map and the semantic feature map, so as to better restore the panoramic video frame with less bandwidth consumption according to the characteristics of the panoramic video, and improve the quality of the user's immersive experience. Based on the entropy model and combined with the latitude adaptive model, the information entropy of each feature point of the semantic feature is appropriately scaled, so as to reduce the information redundancy in the transmission process of the panoramic video and achieve the efficient transmission of the panoramic video.

[0033] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0035] Figure 1 is an effect diagram of converting a spherical video into a two-dimensional video using the cylindrical projection method in the prior art;

[0036] Figure 2 is a schematic diagram of the steps of the panoramic video transmission method for the sender in the embodiment of the present disclosure;

[0037] Figure 3 is a system architecture diagram for implementing the panoramic video transmission method in the embodiment of the present disclosure;

[0038] Figure 4 is a network structure diagram of the weight attention module in the embodiment of the present disclosure;

[0039] Figure 5 is a network structure diagram of the latitude adaptive model in the embodiment of the present disclosure;

[0040] Figure 6 is a comparison diagram of the simulation experiment results of the panoramic video transmission method in the present disclosure;

[0041] Figure 7 is a schematic diagram of the steps of the panoramic video transmission method for the receiver in the present disclosure;

[0042] Figure 8 is a principle block diagram of the panoramic video transmission device for the sender in the embodiment of the present disclosure;

[0043] Figure 9 is a principle block diagram of the panoramic video transmission device for the receiver in the embodiment of the present disclosure;

[0044] Figure 10It is a block diagram of an electronic device for implementing the panoramic video transmission method of the embodiments of the present disclosure. Detailed implementation manners

[0045] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0046] As one of the key technologies of 6G, semantic communication pursues higher semantic fidelity rather than error-free transmission of symbols. This technology uses a deep neural network to extract semantic features of the information to be transmitted, sends the semantic features to the receiving end, and the receiving end can obtain the information transmitted by the sending end through semantic restoration. Since the data volume of the semantic features obtained by semantic extraction is significantly reduced, the amount of data to be transmitted is reduced, and semantic-level compression of various types of information such as text, images, and videos is achieved. For this reason, the existing third solution proposes a semantic communication video transmission method, which extracts corresponding semantic features according to the current video frame to be transmitted and the context information formed between the front and back frames, calculates the corresponding channel bandwidth cost in combination with the information entropy of the extracted semantic features, performs joint source-channel coding in combination with the semantic features and the channel bandwidth cost, and sends the encoded information into the channel for transmission. This method can adaptively extract the semantic information of the current video frame and perform variable-length coding, thereby realizing efficient video transmission.

[0047] Although the third solution improves the transmission efficiency of traditional videos, it is not applicable to panoramic videos. The main reasons are as follows: (1) For panoramic videos, it is also necessary to convert them from three-dimensional spherical videos into two-dimensional plane videos for transmission. Usually, the equirectangular projection (ERP) method is used to convert spherical videos into two-dimensional videos, similar to Figure 1As shown in the figure. However, this will cause different degrees of stretching in different latitudes of the panoramic video. The stretching is more severe towards the two poles, and the resulting information redundancy is greater. For example, the flat world maps we usually see are obtained through cylindrical projection. Upon careful observation, it can be found that the South Pole may be just a point on the sphere, but after projection, it is stretched into many pixel points, which causes information redundancy; (2) The quality of the panoramic video obtained by the receiving end through semantic restoration directly determines the quality of the user's immersive experience. Since the data to be transmitted is a panoramic video, different from the evaluation metrics of traditional videos, the evaluation metric for panoramic videos uses weighted peak signal-to-noise ratio (WS-PSNR), which can better reflect the quality of the user's immersive experience. If the semantic video transmission method of the third solution is used to transmit the panoramic video, the information entropy is inaccurately allocated, resulting in low utilization rate of channel resources, which will inevitably lead to a decline in the quality of the user's immersive experience.

[0048] The present disclosure provides a panoramic video transmission method, as Figure 2 shown, including:

[0049] Step S201, the sending end performs semantic extraction on the first panoramic video to be transmitted to obtain corresponding semantic features.

[0050] Step S202, the sending end performs weighted attention processing on the semantic features through a weight map to obtain a weighted attention feature map.

[0051] Step S203, the sending end determines the transmission rate of each feature point in the weighted attention feature map according to an entropy model and a latitude adaptation model.

[0052] Step S204, the sending end encodes the weighted attention feature map according to the transmission rate to obtain semantic encoded information.

[0053] Step S205, the sending end sends the semantic encoded information to the receiving end; wherein, the semantic encoded information is used to provide the receiving end with the second panoramic video obtained through decoding and semantic restoration.

[0054] Specifically, as Figure 3 shown, the transmission process of the main link of the sending end includes: at the sending end, the panoramic video frame x t at the t-th moment (i.e., the first panoramic video) is input into the semantic extraction module 301 for semantic extraction to obtain semantic features y t ; the semantic features y t are input into the weighted attention module 302, and the weighted attention module 302 performs weighted attention processing on the semantic features y tPerform weighted attention operation to obtain a weighted attention feature map, achieving an improvement in WS-PSNR; input the weighted attention feature map into the encoding module 303 (which can use Deep Joint Source-Channel Coding, Deep JSCC) for encoding. During the encoding process, the guidance of an entropy model and a dimension adaptation model is required to allocate the transmission rate of each feature point of the weighted attention feature map; then send the encoded semantic coding information into the channel 304 for transmission, and Gaussian white noise can be added to the information to be transmitted in the channel; similarly, at the receiving end, use the decoding module 305 (which can use a deep joint source-channel decoder) for decoding and the semantic recovery module 306 for semantic recovery to obtain the panoramic video frame at time t. (i.e., the second panoramic video).

[0055] The entropy model 307 in this embodiment is used to calculate the information entropy of each feature point of the semantic features. In information theory, probability is used to calculate the information entropy. Here, the entropy model is equivalent to predicting a probability for each feature point, and thus the information entropy is calculated. The information entropy of each feature point forms the information entropy map e. t . The dimension adaptation model 308 is used to adjust the information entropy map, jointly realizing the distribution of the information entropy of the semantic features, and guiding the encoding of the semantic features by adjusting the loss function of the dimension adaptation model, so as to reasonably utilize the channel resources to transmit information, reduce the redundancy of the transmitted information of the panoramic video, improve the utilization rate of the channel resources, and improve the quality of the panoramic video recovered at the receiving end, and improve the quality of the user's immersive experience.

[0056] As an optional implementation manner, the sender performs semantic extraction on the first panoramic video to be transmitted, and the obtained corresponding semantic features include:

[0057] Obtain the current frame of the first panoramic video and obtain the context information between the current frame and the previous frame;

[0058] Input the current frame and the context information into the semantic extraction model to obtain the corresponding semantic features.

[0059] Specifically, as Figure 3 shown, by performing motion estimation between frames (i.e., calculating the difference data between the two) on the panoramic video frame x at time t t and the panoramic video frame recovered at time t-1 , the context information c t between the current frame (the panoramic video frame x at time t ) and the previous frame (the panoramic video frame at time t-1 t ) is obtained. This transmission link is called the motion link. In this embodiment, the input in the semantic extraction module 301 includes the panoramic video frame x at time t t and the context information c t, when the sender performs semantic extraction, it combines context information, which can improve the quality of the panoramic video obtained after semantic extraction and semantic restoration.

[0060] As an optional implementation, the context information obtained by the sender for the current frame and the previous frame includes:

[0061] Calculate the difference data between the current frame of the first panoramic video and the previous frame of the current frame;

[0062] Perform semantic extraction on the difference data to obtain the second semantic feature;

[0063] Perform weighted attention processing on the second semantic feature through the second weight map to obtain the second weighted attention feature map;

[0064] Determine the second transmission rate corresponding to each feature point in the second weighted attention feature map according to the second entropy model and the second latitude adaptive model;

[0065] Encode the second weighted attention feature map according to the second transmission rate to obtain the second semantic coding information;

[0066] Decode the second semantic coding information and perform semantic restoration to obtain the second difference data;

[0067] Calculate the context information based on the second difference data and the previous frame of the current frame.

[0068] Figure 3 The action link shown is used to obtain context information. The action link calculates the panoramic video frame x at time t through the action estimation network 309 t and the panoramic video frame restored at time t - 1 The difference data m between t , for the difference data m t Perform semantic extraction, weighted attention processing, and encoding according to the entropy model and the latitude adaptive model, and obtain Finally and the panoramic video frame at time t - 1 Input the context generation module 310 to obtain the context information c t . It should be noted that Figure 3 The structure of the action link shown is similar to the structure of the main link, both including a semantic extraction module, an entropy model and a latitude adaptive model, an encoding module, a decoding module, a semantic restoration module, etc., which will not be elaborated below.

[0069] As an optional implementation, the sender performs weighted attention processing on the semantic feature through the weight map to obtain the weighted attention feature map, including:

[0070] Perform n times of average pooling downsampling on the weight map w to obtain the first sampling result;

[0071] Perform n times of max pooling downsampling on the weight map w to obtain the second sampling result;

[0072] Concatenate (Concat) the first sampling result and the second sampling result with the semantic features in the channel depth to obtain the concatenated semantic feature map;

[0073] Perform the first convolution on the semantic feature map to learn the relationship between the semantic feature map and the weight map;

[0074] Perform the second convolution on the semantic feature map to adjust the number of channels of the semantic feature map;

[0075] Input the semantic feature map after the second convolution into the activation function, and output the weight attention map through the activation function;

[0076] Multiply the weight attention map and the semantic features point by point to obtain the weight attention feature map.

[0077] Specifically, the structure of the weight attention module is as Figure 4 shown, including an average pooling downsampling layer (AvgPool), a max pooling downsampling layer (MaxPool), two 3x3 convolutional layers, and an activation function Sigmiod. The input of the weight attention module includes the weight map w and the semantic features y t , perform 4 times of average pooling downsampling on the weight map w through the average pooling downsampling layer, and perform 4 times of max pooling downsampling on the weight map w through the max pooling downsampling layer l, and concatenate the sampling results and the semantic features y t , then through two convolutions, and finally map to 0 to 1 through the Sigmiod function to obtain the spatial weight attention map at the current moment, and multiply the semantic features y t point by point with the spatial weight attention map to obtain the output of the weight attention module, that is, the weight attention feature map.

[0078] As an optional implementation manner, as Figure 3 shown, the sending end determines the transmission rate of each feature point in the weight attention feature map according to the entropy model and the latitude adaptive model, including:

[0079] Calculate the information entropy corresponding to each feature point of the semantic features through the entropy model 307 to obtain the information entropy map e t ; among them, the input of the entropy model 307 is the semantic features y obtained by the semantic extraction module 301 for semantic extraction t .

[0080] Input the information entropy map e tInput the latitude adaptive model 308 with the weight map w, and calculate the scaling parameter ω t ;

[0081] According to the scaling parameter ω t and the information entropy map e t Determine the transmission rate of each feature point in the weight attention feature map.

[0082] In this embodiment, the scaling parameter ω calculated by the entropy model 307 and the latitude adaptive model 308 t acts on the information entropy map e t , can adjust the information entropy of each feature point in the information entropy map, and guide the encoding of semantic feature information by participating in the design of the loss function of the encoder, so as to reasonably utilize the channel resources to transmit information. Among them, the loss function used by the encoder can be expressed as:

[0083] ∑ m ∑ n max(0, ω t,mn × max(e t ) - e t,mn )

[0084] Among them, ω t,mn represents the value of ω at the coordinates (m, n) at time t t ; ω t represents the scaling parameter; e t represents the information entropy map.

[0085] The scaling parameter ω calculated by the entropy model 307 and the latitude adaptive model 308 t participates in the design of the loss function in the encoder, and adjusts the transmission rate of each feature point during the encoding process to improve the utilization rate of channel resources. It should be noted that the loss function is not fixed, and the above loss function is only one implementation method of allocating the transmission rate through the scaling parameter.

[0086] As an optional implementation method, input the information entropy map and the weight map into the latitude adaptive model, and the calculated scaling parameter includes:

[0087] Calculate the average value of the information entropy map in the longitude direction, and reduce the dimension of the information entropy map to obtain the second information entropy map;

[0088] Input the second information entropy map into the first fully connected layer of the latitude adaptive model for expansion processing, and then input it into the second fully connected layer of the latitude adaptive model for scaling processing to obtain the third information entropy map;

[0089] Expand the third information entropy map in the longitude direction to obtain the fourth information entropy map;

[0090] Adaptive average pooling downsampling is performed on the weight map to obtain a weight map with the same dimension as the fourth information entropy map;

[0091] The scaling parameter is calculated based on the fourth information entropy map and the weight map.

[0092] Specifically, the network structure of the latitude adaptive model is as Figure 5 shown, including a longitude average layer 501, a first fully connected layer 502, a second fully connected layer 503, a longitude expansion layer 504, a calculation layer 505, and an adaptive average pooling downsampling layer 506. The input of the latitude adaptive model is the information entropy map e t , e t The longitude average layer 501, the first fully connected layer 502, the second fully connected layer 503, and the longitude expansion layer 504 obtain the fourth information entropy map η t , and the adaptive average pooling downsampling layer 506 is used to perform adaptive average pooling downsampling on the input weight map w to obtain w AAP , making its dimension the same as η t . Finally, the scaling parameter ω t is calculated according to η AAP and w t . Among them, the calculation formula can be expressed as:

[0093]

[0094] The present disclosure utilizes the weight attention module, which can improve the immersive experience quality metric WS-PSNR of the panoramic video. It utilizes the powerful fitting ability of the neural network to associate the implicit information existing between the weight map and the semantic feature map, so as to better restore the panoramic video frame with less bandwidth consumption according to the characteristics of the panoramic video and improve the immersive experience quality of users. Based on the entropy model and combined with the latitude adaptive model, the information entropy of each feature point of the semantic features is appropriately scaled, so as to reduce the information redundancy in the transmission process of the panoramic video and achieve the efficient transmission of the panoramic video.

[0095] The calculation formula of the weighted peak signal-to-noise ratio WS-PSNR is as follows:

[0096]

[0097] Among them, MAX represents the maximum value of the pixel point; WMSE represents the weighted mean square error, and its calculation formula is expressed as:

[0098]

[0099] x ij represents the pixel point value at the coordinate (i, j) on the original image, Denote the value of the pixel at coordinates (i, j) on the restored image, w ij Denote the weight value corresponding to the pixel at coordinates (i, j).

[0100] The simulation results of the entire system are as Figure 6 shown. The abscissa is the channel bandwidth overhead ratio, and the ordinate is the metric WS-PSNR. Among them, DVST represents a transmission network that only uses the entropy model to calculate the information entropy and does not set a weight attention module; APVST (w / o WA) represents a network improved on the basis of DVST by adding a latitude adaptive model; APVST represents a network improved on the basis of DVST by adding a latitude adaptive model and a weight attention module. It can be seen from Figure 6 the simulation results that the technical solution proposed in this disclosure has a gain for the transmission of panoramic videos, and can approximately save 20% of the channel bandwidth consumption.

[0101] This disclosure also provides another panoramic video transmission method, which can be applied to the receiving end, as Figure 7 shown, including:

[0102] Step S701, the receiving end receives semantic coding information; wherein, the semantic coding information is obtained by the sending end extracting semantic features from the first panoramic video to be transmitted, performing weight attention processing on the semantic features through a weight map, determining the transmission rate of each feature point in the weight attention feature map according to the entropy model and the latitude adaptive model, and then encoding the weight attention feature map according to the transmission rate.

[0103] Step S702, the receiving end decodes the semantic coding information to obtain the third semantic feature.

[0104] Step S703, the receiving end performs semantic restoration on the third semantic feature to obtain the second panoramic video.

[0105] Specifically, the semantic coding information of this disclosure passes through the weight attention module of the sending end, which can associate the implicit information existing between the weight map and the semantic feature map. Thus, the receiving end can better restore the panoramic video frames with less bandwidth consumption according to the characteristics of the panoramic video, improving the quality of the user's immersive experience. Based on the entropy model of the sending end, combined with the latitude adaptive model, the information entropy of each feature point of the semantic features is appropriately scaled, so as to reduce the information redundancy in the transmission process of the panoramic video and achieve the efficient transmission of the panoramic video.

[0106] This disclosure also provides a panoramic video transmission device, as Figure 8 shown, including:

[0107] The semantic extraction module 301 is configured to perform semantic extraction on the first panoramic video to be transmitted, and obtain corresponding semantic features.

[0108] The weight attention module 302 is configured to perform weight attention processing on the semantic features through a weight map, and obtain a weight attention feature map.

[0109] The rate allocation module 311 is configured to determine the transmission rate of each feature point in the weight attention feature map according to the entropy model 307 and the dimension adaptation model 308.

[0110] The encoding module 303 is configured to encode the weight attention feature map according to the transmission rate to obtain semantic encoding information.

[0111] The transmission module 312 is configured to send the semantic encoding information to the receiving end through the channel 304; wherein, the semantic encoding information is used to provide the receiving end with the second panoramic video obtained through decoding and semantic restoration.

[0112] Specifically, as Figure 3 shown, the transmission process of the main link at the sending end includes: at the sending end, the panoramic video frame x t (i.e., the first panoramic video) at the t-th moment is input into the semantic extraction module 301 for semantic extraction to obtain semantic features y t ; the semantic features y t are input into the weight attention module 302, and the weight attention module 302 performs weight attention operation on the semantic features y t through the weight map w to obtain a weight attention feature map, realizing the improvement of WS-PSNR; the weight attention feature map is input into the encoding module 303 (which can adopt Deep Joint Source Channel Coding, Deep JSCC) for encoding, and during the encoding process, it is necessary to be guided by the entropy model and the dimension adaptation model to allocate the transmission rate of each feature point of the weight attention feature map; then the encoded semantic encoding information is sent into the channel 304 for transmission, and Gaussian white noise can be added to the information to be transmitted in the channel; similarly, at the receiving end, the decoding module 305 (which can adopt a deep joint source channel decoder) is used for decoding, and the semantic restoration module 306 is used for semantic restoration to obtain the panoramic video frame (i.e., the second panoramic video) at the t-th moment.

[0113] The entropy model 307 in this embodiment is used to calculate the information entropy of each feature point of the semantic features. In information theory, probability is used to calculate the information entropy. Here, the entropy model is equivalent to predicting a probability for each feature point, so as to calculate the information entropy, and the information entropy of each feature point forms an information entropy map e t。The latitude adaptive model 308 is used to adjust the information entropy map, jointly implement the information entropy distribution of semantic features, and guide the encoding of semantic features through the loss function adjusted by the latitude adaptive model, so as to reasonably utilize the channel resources to transmit information, reduce the redundancy of the transmitted information of the panoramic video, improve the utilization rate of channel resources, and improve the quality of the panoramic video restored at the receiving end, as well as improve the quality of the user's immersive experience.

[0114] As an alternative implementation, the semantic extraction module 301 includes:

[0115] An acquisition unit configured to acquire the current frame of the first panoramic video and the context information between the current frame and the previous frame.

[0116] A semantic extraction unit configured to input the current frame and the context information into a semantic extraction model to obtain corresponding semantic features.

[0117] Specifically, as Figure 3 shown, by performing frame-to-frame motion estimation on the panoramic video frame x at time t t and the panoramic video frame restored at time t-1 (that is, calculating the difference data between the two), the context information c t between the current frame (the panoramic video frame x at time t ) and the previous frame (the panoramic video frame at time t-1 t ) is obtained. This transmission link is called the motion link. In this embodiment, the input in the semantic extraction module 301 includes the panoramic video frame x at time t t and the context information c t . When performing semantic extraction at the sending end, the context information is combined, which can improve the quality of the panoramic video obtained after semantic extraction and semantic restoration.

[0118] As an alternative implementation, the acquisition unit's acquisition of the context information between the current frame and the previous frame includes:

[0119] Calculating the difference data between the current frame of the first panoramic video and the previous frame of the current frame;

[0120] Performing semantic extraction on the difference data to obtain a second semantic feature;

[0121] Performing weighted attention processing on the second semantic feature through a second weight map to obtain a second weighted attention feature map;

[0122] Determining the second transmission rate corresponding to each feature point in the second weighted attention feature map according to a second entropy model and a second latitude adaptive model;

[0123] Encoding the second weighted attention feature map according to the second transmission rate to obtain second semantic encoding information;

[0124] Decode the second semantic coding information and perform semantic restoration to obtain the second difference data;

[0125] Calculate the context information based on the second difference data and the previous frame of the current frame.

[0126] Figure 3 The action link shown is used to obtain context information. The action link calculates the panoramic video frame x at time t through the action estimation network 309 t and the panoramic video frame restored at time t-1 The difference data m between t , perform semantic extraction, weight attention processing on the difference data m t and encode according to the entropy model and the latitude adaptive model, and obtain through decoding and semantic restoration at the receiving end Final and the panoramic video frame at time t-1 Input the context generation module 310 to obtain the context information c t . It should be noted that Figure 3 The structure of the action link shown is similar to that of the main link, and both include a semantic extraction module, an entropy model and a latitude adaptive model, an encoding module, a decoding module, a semantic restoration module, etc., which will not be elaborated below.

[0127] As an optional implementation manner, the weight attention module 302 performs weight attention processing on the semantic features through a weight map, and the obtained weight attention feature map includes:

[0128] Perform n times of average pooling downsampling processing on the weight map w to obtain the first sampling result;

[0129] Perform n times of max pooling downsampling processing on the weight map w to obtain the second sampling result;

[0130] Concatenate (Concat) the first sampling result and the second sampling result with the semantic features in the channel depth to obtain the concatenated semantic feature map;

[0131] Perform the first convolution processing on the semantic feature map to learn the relationship between the semantic feature map and the weight map;

[0132] Perform the second convolution processing on the semantic feature map to adjust the number of channels of the semantic feature map;

[0133] Input the semantic feature map after the second convolution processing into the activation function, and output the weight attention map through the activation function;

[0134] Multiply the weight attention map and the semantic features point by point to obtain the weight attention feature map.

[0135] Specifically, the structure of the weight attention module is as follows Figure 4 shown, including an average pooling downsampling layer AvgPool, a max pooling downsampling layer MaxPool, two 3x3 convolutional layers, and an activation function Sigmiod. The input of the weight attention module includes a weight map w and a semantic feature y t , the weight map w is subjected to 4 times of average pooling downsampling by the average pooling downsampling layer AvgPool, and the weight map w is subjected to 4 times of max pooling downsampling by the max pooling downsampling layer MaxPool, and the sampling results and the semantic feature y t are concatenated, and then through two convolutions, and finally mapped to 0 to 1 through the Sigmiod function to obtain the spatial weight attention map at the current moment, and the semantic feature y t is multiplied point by point with the spatial weight attention map to obtain the output of the weight attention module, that is, the weight attention feature map.

[0136] As an alternative implementation, as shown in Figure 3 , the rate allocation module 311 determines the transmission rate of each feature point in the weight attention feature map according to the entropy model 307 and the latitude adaptive model 308, including:

[0137] Calculating the information entropy corresponding to each feature point of the semantic feature through the entropy model 307 to obtain the information entropy map e t ;

[0138] Inputting the information entropy map e t and the weight map w into the latitude adaptive model 308 to calculate and obtain the scaling parameter ω t ;

[0139] Determining the transmission rate of each feature point in the weight attention feature map according to the scaling parameter ω t and the information entropy map e t .

[0140] In this embodiment, the scaling parameter ω t calculated by the entropy model 307 and the latitude adaptive model 308 acts on the information entropy map e t , which can adjust the information entropy of each feature point in the information entropy map, and by participating in the design of the loss function of the encoder, guide the encoding of semantic feature information, so as to reasonably utilize the channel resources to transmit information. Among them, the loss function used by the encoder can be expressed as:

[0141] ∑ m ∑ n max(0, ω t,mn × max(e t ) - e t,mn )

[0142] Among them, ω t,mn represents the value of ω at the coordinate (m, n) at time t t ; ω t represents the scaling parameter; e t represents the information entropy map.

[0143] The scaling parameter ω calculated by the entropy model 307 and the latitude adaptive model 308 t is involved in the design of the loss function in the encoder. During the encoding process, the transmission rate of each feature point is adjusted to improve the utilization rate of channel resources. It should be noted that the loss function is not fixed, and the above loss function is only one implementation method of allocating the transmission rate through the scaling parameter.

[0144] The scaling parameters calculated by the latitude adaptive model include:

[0145] Calculate the average value of the information entropy map in the longitude direction, and reduce the dimension of the information entropy map to obtain the second information entropy map;

[0146] Input the second information entropy map into the first fully connected layer of the latitude adaptive model for expansion processing, and then input it into the second fully connected layer of the latitude adaptive model for scaling processing to obtain the third information entropy map;

[0147] Expand the third information entropy map in the longitude direction to obtain the fourth information entropy map;

[0148] Perform adaptive average pooling downsampling on the weight map to obtain a weight map with the same dimension as the fourth information entropy map;

[0149] Calculate the scaling parameter according to the fourth information entropy map and the weight map.

[0150] Specifically, the network structure of the latitude adaptive model is as Figure 5 shown, including a longitude average layer 501, a first fully connected layer 502, a second fully connected layer 503, a longitude expansion layer 504, a calculation layer 505, and an adaptive average pooling downsampling layer 506. The input of the latitude adaptive model is the information entropy map e t , e t The longitude average layer 501, the first fully connected layer 502, the second fully connected layer 503, and the longitude expansion layer 504 obtain the fourth information entropy map η t , and the adaptive average pooling downsampling layer 506 is used to perform adaptive average pooling downsampling on the input weight map w to obtain w AAP , making its dimension the same as η t , and finally calculating the scaling parameter ω t according to η AAP and w t . Among them, the calculation formula can be expressed as:

[0151]

[0152] The present disclosure utilizes a weighted attention module to improve the immersive experience quality metric WS - PSNR for panoramic videos. By leveraging the powerful fitting ability of neural networks, it can associate the implicit information between the weight map and the semantic feature map, thereby better recovering panoramic video frames with less bandwidth consumption according to the characteristics of panoramic videos and enhancing the immersive experience quality of users. Based on the entropy model and combined with the latitude adaptive model, the information entropy of each feature point of the semantic features is appropriately scaled, so as to reduce the information redundancy during the transmission of panoramic videos and achieve efficient transmission of panoramic videos.

[0153] The present disclosure also provides a panoramic video transmission device, as Figure 9 shown, which can be set at the receiving end and includes:

[0154] A receiving module 313, configured to receive semantic coding information; wherein, the semantic coding information is obtained by the sending end through semantic extraction of the first panoramic video to be transmitted to obtain semantic features, performing weighted attention processing on the semantic features through a weight map, determining the transmission rate of each feature point in the weighted attention feature map according to the entropy model and the latitude adaptive model, and then encoding the weighted attention feature map according to the transmission rate;

[0155] A decoding module 305, configured to decode the semantic coding information to obtain a third semantic feature.

[0156] A semantic recovery module 306, configured to perform semantic recovery on the third semantic feature to obtain a second panoramic video.

[0157] Specifically, the semantic coding information of the present disclosure passes through the weighted attention module at the sending end, which can associate the implicit information between the weight map and the semantic feature map, so that the receiving end can better recover panoramic video frames with less bandwidth consumption according to the characteristics of panoramic videos and enhance the immersive experience quality of users. Based on the entropy model at the sending end and combined with the latitude adaptive model, the information entropy of each feature point of the semantic features is appropriately scaled, so as to reduce the information redundancy during the transmission of panoramic videos and achieve efficient transmission of panoramic videos.

[0158] In the technical solution of the present disclosure, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0159] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0160] Figure 10 FIG. Figure 10 shows a schematic block diagram of an exemplary electronic device 1000 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementations of the present disclosure described and / or claimed herein.

[0161] As Figure 10 shown, the device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0162] A plurality of components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as, for example, a keyboard, a mouse, etc.; an output unit 1007, such as, for example, various types of displays, speakers, etc.; a storage unit 1008, such as, for example, a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as, for example, a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0163] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning objective function algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as the panoramic video transmission method. For example, in some embodiments, the panoramic video transmission method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the panoramic video transmission method described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the panoramic video transmission method by any other suitable means (e.g., by means of firmware).

[0164] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0165] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0166] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0167] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0168] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0169] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0170] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of this disclosure can be achieved, and no limitation is imposed herein.

[0171] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A panoramic video transmission method, characterized in that, Including: The sending end performs semantic extraction on the first panoramic video to be transmitted to obtain corresponding semantic features; The sending end performs weighted attention processing on the semantic features through a weight map to obtain a weighted attention feature map; The sending end determines the transmission rate of each feature point in the weighted attention feature map according to an entropy model and a latitude adaptation model; The sending end encodes the weighted attention feature map according to the transmission rate to obtain semantic encoding information; The sending end sends the semantic encoding information to the receiving end; wherein, the semantic encoding information is used to enable the receiving end to obtain a second panoramic video through decoding and semantic restoration.

2. The method according to claim 1, wherein The sending end performs semantic extraction on the first panoramic video to be transmitted to obtain corresponding semantic features, including: Obtain the current frame of the first panoramic video and obtain the context information between the current frame and the previous frame; Input the current frame and the context information into a semantic extraction model to obtain the corresponding semantic features.

3. The method according to claim 2, wherein The sending end obtains the context information between the current frame and the previous frame, including: Calculate the difference data between the current frame of the first panoramic video and the previous frame of the current frame; Perform semantic extraction on the difference data to obtain second semantic features; Perform weighted attention processing on the second semantic features through a second weight map to obtain a second weighted attention feature map; Determine the second transmission rate corresponding to each feature point in the second weighted attention feature map according to a second entropy model and a second latitude adaptation model; Encode the second weighted attention feature map according to the second transmission rate to obtain second semantic encoding information; Decode the second semantic encoding information and perform semantic restoration to obtain second difference data; Calculate the context information according to the second difference data and the previous frame of the current frame.

4. The method according to claim 1, characterized in that, The sending end performs weighted attention processing on the semantic features through a weight map to obtain a weighted attention feature map, including: Perform n times of average pooling downsampling processing on the weight map to obtain a first sampling result; Perform n times of max pooling downsampling processing on the weight map to obtain a second sampling result; Concatenate the first sampling result and the second sampling result with the semantic features in the channel depth to obtain a concatenated semantic feature map; Perform a first convolution process on the semantic feature map to learn the relationship between the semantic feature map and the weight map; Perform a second convolution process on the semantic feature map to adjust the number of channels of the semantic feature map; Input the semantic feature map after the second convolution process into an activation function, and output a weighted attention map through the activation function; Multiply the weighted attention map and the semantic features point by point to obtain the weighted attention feature map.

5. The method according to any one of claims 1-4, characterized in that The sending end determines the transmission rate of each feature point in the weighted attention feature map according to an entropy model and a latitude adaptation model, including: Calculate the information entropy corresponding to each feature point of the semantic features through the entropy model to obtain an information entropy map; Input the information entropy map and the weight map into the latitude adaptation model to calculate a scaling parameter; Determine the transmission rate of each feature point in the weight attention feature map according to the scaling parameter and the information entropy map.

6. The method according to claim 5, wherein The calculating the scaling parameter by inputting the information entropy map and the weight map into the latitude adaptive model includes: Calculate the average value of the information entropy map in the longitude direction, and reduce the dimension of the information entropy map to obtain a second information entropy map; Input the second information entropy map into the first fully connected layer of the latitude adaptive model for expansion processing, and then input it into the second fully connected layer of the latitude adaptive model for scaling processing to obtain a third information entropy map; Expand the third information entropy map in the longitude direction to obtain a fourth information entropy map; Perform adaptive average pooling downsampling on the weight map to obtain the weight map with the same dimension as the fourth information entropy map; Calculate the scaling parameter according to the fourth information entropy map and the weight map.

7. A panoramic video transmission method, characterized in that, It includes: The receiving end receives the semantic coding information; wherein, the semantic coding information is obtained by the sending end extracting semantic features from the first panoramic video to be transmitted, performing weight attention processing on the semantic features through a weight map, determining the transmission rate of each feature point in the weight attention feature map according to an entropy model and a latitude adaptive model, and then encoding the weight attention feature map according to the transmission rate; The receiving end decodes the semantic coding information to obtain a third semantic feature; The receiving end performs semantic restoration on the third semantic feature to obtain a second panoramic video.

8. A panoramic video transmission device, characterized in that, It includes: A semantic extraction module configured to extract semantic features from the first panoramic video to be transmitted to obtain corresponding semantic features; A weight attention module configured to perform weight attention processing on the semantic features through a weight map to obtain a weight attention feature map; A rate allocation module configured to determine the transmission rate of each feature point in the weight attention feature map according to an entropy model and a latitude adaptive model; An encoding module configured to encode the weight attention feature map according to the transmission rate to obtain semantic coding information; A transmission module configured to send the semantic coding information to the receiving end; wherein, the semantic coding information is used to enable the receiving end to obtain a second panoramic video through decoding and semantic restoration.

9. The device according to claim 8, characterized in that, The semantic extraction module includes: An acquisition unit configured to acquire the current frame of the first panoramic video and acquire the context information between the current frame and the previous frame; A semantic extraction unit configured to input the current frame and the context information into a semantic extraction model to obtain the corresponding semantic features.

10. The device according to claim 9, characterized in that, The acquisition unit acquiring the context information between the current frame and the previous frame includes: Calculating the difference data between the current frame of the first panoramic video and the previous frame of the current frame; Performing semantic extraction on the difference data to obtain a second semantic feature; Performing weight attention processing on the second semantic feature through a second weight map to obtain a second weight attention feature map; Determining the second transmission rate corresponding to each feature point in the second weight attention feature map according to a second entropy model and a second latitude adaptive model; Encode the second weighted attention feature map according to the second transmission rate to obtain second semantic encoding information; Decode the second semantic encoding information and perform semantic restoration to obtain second difference data; Calculate the context information according to the second difference data and the previous frame of the current frame.

11. The device according to claim 8, characterized in that The weighted attention module performs weighted attention processing on the semantic features through a weight map to obtain a weighted attention feature map, including: Perform n times of average pooling downsampling processing on the weight map to obtain a first sampling result; Perform n times of max pooling downsampling processing on the weight map to obtain a second sampling result; Concatenate the first sampling result and the second sampling result with the semantic features in the channel depth to obtain a concatenated semantic feature map; Perform a first convolution process on the semantic feature map to learn the relationship between the semantic feature map and the weight map; Perform a second convolution process on the semantic feature map to adjust the number of channels of the semantic feature map; Input the semantic feature map after the second convolution process into an activation function, and output a weighted attention map through the activation function; Multiply the weighted attention map and the semantic features point by point to obtain the weighted attention feature map.

12. The device according to any one of claims 8-11, characterized in that, The rate allocation module determines the transmission rate of each feature point in the weighted attention feature map according to an entropy model and a latitude adaptive model, including: Calculate the information entropy corresponding to each feature point of the semantic features through the entropy model to obtain an information entropy map; Input the information entropy map and the weight map into the latitude adaptive model to calculate a scaling parameter; Determine the transmission rate of each feature point in the weighted attention feature map according to the scaling parameter and the information entropy map.

13. The device according to claim 12, characterized in that, The latitude adaptive model calculates the scaling parameter, including: Calculate the average value of the information entropy map in the longitude direction, and reduce the dimension of the information entropy map to obtain a second information entropy map; Input the second information entropy map into the first fully connected layer of the latitude adaptive model for expansion processing, and then input it into the second fully connected layer of the latitude adaptive model for scaling processing to obtain a third information entropy map; Expand the third information entropy map in the longitude direction to obtain a fourth information entropy map; Perform adaptive average pooling downsampling processing on the weight map to obtain the weight map with the same dimension as the fourth information entropy map; Calculate the scaling parameter according to the fourth information entropy map and the weight map.

14. A panoramic video transmission device, characterized in that, Including: A receiving module configured to receive semantic encoding information; wherein, the semantic encoding information is obtained by the sending end performing semantic extraction on the first panoramic video to be transmitted to obtain semantic features, performing weighted attention processing on the semantic features through a weight map, determining the transmission rate of each feature point in the weighted attention feature map according to an entropy model and a latitude adaptive model, and then encoding the weighted attention feature map according to the transmission rate; A decoding module configured to decode the semantic encoding information to obtain third semantic features; A semantic restoration module configured to perform semantic restoration on the third semantic features to obtain a second panoramic video.

15. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

17. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-7.