Panoramic video semantic transmission method and device

Through the panoramic video semantic transmission method, semantic coding and resource allocation models are used to optimize panoramic video transmission in immersive communication, solving the problem of signal-to-noise ratio reduction and video quality degradation in high-load scenarios by multiple access technology, achieving more efficient and robust transmission performance.

CN120017811APending Publication Date: 2025-05-16BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510172441.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-01-03
Filing Date
2025-02-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In immersive communication scenarios, the prior art is difficult to effectively manage the interference and noise of multiple access technology, resulting in a decrease in the signal-to-noise ratio and the inability to effectively separate the original signal, which seriously affects the video quality, especially when data is congested.

Method used

A panoramic video semantic transmission method is proposed. Through the combination of semantic coding, field-angle prediction, semantic mapping and resource allocation models, public subsemantic flow and private semantic flow are diverted, and encoding and resource allocation are performed to optimize transmission efficiency and robustness.

Benefits of technology

In high load scenarios, the transmission performance and user experience of panoramic videos are improved, the robustness and resource utilization of the system are enhanced, and the problems of signal-to-noise ratio reduction and video quality degradation are effectively solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017811A_ABST
    Figure CN120017811A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a panoramic video semantic transmission method and device, and the method comprises the steps: carrying out the semantic coding of panoramic video data of all users, and obtaining a semantic information flow; according to the historical field angle information of each user, predicting corresponding current field angle information by using a field angle prediction model; converting the current field angle information into current field angle semantic information; based on the current view field angle semantic information, dividing the semantic information flow of each user into a public sub-semantic flow and a private semantic flow of each user by using a message divider; combining the common sub-semantic streams into a common semantic stream, and coding the common semantic stream based on a common semantic stream codebook to obtain common semantic code words; coding the private semantic stream based on the private semantic stream codebook to obtain private semantic code words; and determining a resource allocation strategy by using a resource allocation model according to the current field angle semantic information, the current channel state information and the historical resource allocation strategy information of each user. According to the invention, resources can be reasonably allocated and transmission performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of communication technology, and in particular to a method and device for transmitting panoramic video semantics. Background Art

[0002] In the 6G vision, immersive communication will become a key application scenario, where users will watch 360-degree panoramic videos and interact with the virtual world. To ensure the quality of service (QoS) of users, base stations need to provide low latency (less than 20ms) and high-quality video transmission. However, for 4K video (resolution of 3840×1920), due to the limitation of the field of view (FoV), users can actually only see images with a resolution of 960×540. Therefore, in order to provide an immersive experience, the network needs to transmit higher-resolution panoramic videos, and the amount of data required to be transmitted increases dramatically.

[0003] Current panoramic video technologies mostly use different multiple access technologies, and make reasonable resource allocation and bit rate selection based on the characteristics of panoramic videos. The interference and noise management of multiple access technologies depends on a good communication environment. When network traffic increases sharply and data congestion occurs, the signal-to-noise ratio drops sharply, and multiple access technologies cannot effectively separate the original signal from the received signal, resulting in a cliff effect on the quality of the immersive experience, which seriously affects the video quality. Summary of the invention

[0004] In view of this, an object of the embodiments of the present application is to provide a panoramic video semantic transmission method and device.

[0005] Based on the above purpose, the embodiment of the present application provides a panoramic video semantic transmission method, which is applied to a sending end, including:

[0006] Perform semantic encoding on the panoramic video data requested by each user to obtain a semantic information stream corresponding to the panoramic video data of each user;

[0007] Inputting the historical viewing angle information of each user into a pre-built viewing angle prediction model, and the viewing angle prediction model outputs the current viewing angle information of each user;

[0008] Mapping the current viewing angle information of each user into the semantic space to obtain the current viewing angle semantic information of each user;

[0009] Based on the current field of view semantic information of each user, a preset message splitter is used to split the semantic information flow of each user into a public sub-semantic flow and a private semantic flow of each user;

[0010] The public sub-semantic streams of each user are merged into a public semantic stream, and the public semantic stream is encoded based on a preset public semantic stream codebook to obtain a public semantic codeword;

[0011] Based on a preset private semantic stream codebook, encode each user's private semantic stream to obtain each user's private semantic codeword;

[0012] The current field of view semantic information, current channel state information, and historical resource allocation strategy information of each user are input into a pre-built resource allocation model, and the resource allocation model outputs the resource allocation strategy of the public semantic codeword and the private semantic codeword.

[0013] Optionally, the current viewing angle information of each user is mapped to the semantic space to obtain the current viewing angle semantic information of each user, including:

[0014] For each user, define a coefficient matrix having the same size as the panoramic video data;

[0015] Matching the coefficient matrix with the user's current field of view angle information, setting the coefficients in the coefficient matrix corresponding to the field of view angle to a first predetermined value, and setting the coefficients outside the field of view angle to a second predetermined value, to obtain an updated coefficient matrix; wherein the first predetermined value is greater than the second predetermined value;

[0016] Performing feature extraction on the updated coefficient matrix to obtain current field of view angle features;

[0017] The current field of view angle feature is vectorized to obtain the current field of view angle semantic information.

[0018] Optionally, based on the current field of view semantic information of each user, a preset message splitter is used to split the semantic information stream of each user into a public sub-semantic stream and a private semantic stream of each user, including:

[0019] The semantic information stream corresponding to the overlapping part of the current field of view semantic information of each user is taken as the public sub-semantic stream, and the semantic information stream corresponding to the non-overlapping part of the current field of view semantic information of each user is taken as the private semantic stream of each user.

[0020] Optionally, the training method of the resource allocation model includes:

[0021] Constructing training samples; wherein the training samples include a certain amount of state information within a certain period of time, the obtained resource allocation strategy, the calculated virtual rewards and real rewards; the state information includes field of view semantic information, channel state information and historical resource allocation strategy information;

[0022] The parameters of the resource allocation model are updated based on the training samples.

[0023] Optionally, the real reward is determined based on the transmission delay score and the video quality score calculated after executing the resource allocation strategy.

[0024] Optionally, the resource allocation strategy includes: a power allocation ratio of the public semantic codeword and the private semantic codeword of each user, a transmission rate allocation ratio, and a channel bandwidth allocation ratio of each user.

[0025] Optionally, the method further includes:

[0026] According to the resource allocation strategy, the public semantic codeword and the private semantic codeword of each user are normalized, and the normalized codewords are superimposed, and the superimposed semantic codewords are transmitted via a channel.

[0027] The embodiment of the present application also provides a panoramic video semantic transmission method, which is applied to a receiving end, and includes receiving a semantic codeword via a channel;

[0028] Decoding the semantic codeword based on a preset public semantic codebook to obtain a decoded public semantic stream;

[0029] Using a preset message splitter, splitting the public sub-semantic stream of the current user from the decoded public semantic stream;

[0030] Performing continuous interference elimination processing based on the semantic codeword and the decoded public semantic stream, removing the decoded public semantic stream, and obtaining the private semantic codeword of each user;

[0031] Decoding the private semantic codewords of each user based on a preset private semantic codebook to obtain a decoded private semantic stream of the current user;

[0032] Merge the public sub-semantic stream of the current user and the decoded private semantic stream of the current user to obtain the semantic information stream of the current user;

[0033] The semantic information stream of the current user is semantically decoded to obtain restored panoramic video data.

[0034] The embodiment of the present application also provides a panoramic video semantic transmission device, which is applied to a sending end, including:

[0035] A video semantic coding module is used to semantically code the panoramic video data requested by each user to obtain a semantic information stream corresponding to the panoramic video data of each user;

[0036] A viewing angle prediction module, used to input the historical viewing angle information of each user into a pre-built viewing angle prediction model, and the viewing angle prediction model outputs the current viewing angle information of each user;

[0037] A semantic mapping module is used to map the current viewing angle information of each user into a semantic space to obtain the current viewing angle semantic information of each user;

[0038] A segmentation module, for dividing the semantic information stream of each user into a public sub-semantic stream and a private semantic stream of each user by using a preset message segmentor based on the current field of view semantic information of each user;

[0039] A merging and encoding module, used to merge the public sub-semantic streams of each user into a public semantic stream, and encode the public semantic stream based on a preset public semantic stream codebook to obtain a public semantic codeword;

[0040] A private semantic encoding module, used to encode the private semantic stream of each user based on a preset private semantic stream codebook to obtain the private semantic codeword of each user;

[0041] The resource allocation module is used to input the current field of view semantic information, current channel state information, and historical resource allocation strategy information of each user into a pre-built resource allocation model, and the resource allocation model outputs the resource allocation strategy of the public semantic codeword and the private semantic codeword.

[0042] The embodiment of the present application further provides a panoramic video semantic transmission device, which is applied to a receiving end, comprising:

[0043] A receiving module, used for receiving semantic codewords via a channel;

[0044] A public codeword decoding module, used for decoding the semantic codeword based on a preset public semantic codebook to obtain a decoded public semantic stream;

[0045] A public segmentation module, used to segment the public sub-semantic stream of the current user from the decoded public semantic stream using a preset message segmentor;

[0046] A private segmentation module, used for performing continuous interference elimination processing based on the semantic codeword and the decoded public semantic stream, removing the decoded public semantic stream, and obtaining the private semantic codeword of each user;

[0047] A private codeword decoding module, used to decode the private semantic codewords of each user based on a preset private semantic codebook to obtain a decoded private semantic stream of the current user;

[0048] A merging module, used to merge the public sub-semantic stream of the current user and the decoded private semantic stream of the current user to obtain the semantic information stream of the current user;

[0049] The restoration module is used to perform semantic decoding on the semantic information stream of the current user to obtain restored panoramic video data.

[0050] From the above description, it can be seen that the panoramic video semantic transmission method and device provided in the embodiment of the present application, by combining semantic communication with rate division multiple access technology, can achieve reasonable resource allocation for public semantic streams and private semantic streams in complex scenarios with multiple concurrent users, thereby improving the overall transmission efficiency and robustness in high-load scenarios and improving the transmission performance of panoramic videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0052] Figure 1 This is a schematic diagram of a method flow chart applied to a sending end according to an embodiment of the present application;

[0053] Figure 2 This is a schematic diagram of a method flow chart applied to a receiving end according to an embodiment of the present application;

[0054] Figure 3 A system structure block diagram of an embodiment of the present application;

[0055] Figure 4 This is a structural block diagram of a device applied to a transmitting end according to an embodiment of the present application;

[0056] Figure 5 This is a structural block diagram of a device applied to a receiving end according to an embodiment of the present application;

[0057] Figure 6 This is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0058] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0059] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should be the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present application do not represent any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing in front of the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connecting" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to represent relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0060] like Figure 1 , 3 As shown, the embodiment of the present application provides a panoramic video semantic transmission method, which is applied to a sending end and includes:

[0061] S101: semantically encode the panoramic video data requested by each user to obtain a semantic information stream corresponding to the panoramic video data of each user;

[0062] In this embodiment, the sending end can be regarded as the server end, and the receiving end can be regarded as the client end. Multiple users use the client end to request the server to obtain corresponding panoramic video data. The server performs semantic encoding according to the panoramic video data requested by each user to obtain the semantic information stream of each user. Among them, the semantic encoding of the panoramic video data is realized based on the semantic communication technology, and the panoramic video data is processed by the semantic communication to obtain the corresponding semantic information stream.

[0063] S102: inputting the historical viewing angle information of each user into a pre-built viewing angle prediction model, and the viewing angle prediction model outputs the current viewing angle information of each user;

[0064] In this embodiment, the sending end can obtain the user's historical field of view information from the client, input the historical field of view information into a preset field of view prediction model, and the model predicts the current field of view information based on the historical field of view information. Wherein, the field of view information includes pitch angle (Pitch), yaw angle (Yaw) and roll angle (Roll), etc. The pitch angle is used to describe the rotation of the user's perspective up and down. When the user raises or lowers his head, the perspective changes in pitch. The yaw angle is used to describe the rotation of the user's perspective left and right. When the user turns his head or moves his sight to the left or right, the perspective changes in yaw. The roll angle is used to describe the rotation of the user's perspective, which is usually a rotation around the horizontal axis. The pitch angle, yaw angle and roll angle determine the center coordinates of the content viewed by the user. The size of the user's field of view can be determined according to the size range of the three angles. For example, the field of view is the range covered by the center coordinate extending 30 degrees up and down (a total range of 60 degrees) and extending 60 degrees left and right (a total range of 120 degrees). Outside this range is outside the field of view. The field of view angle prediction model can be implemented based on a neural network model. This embodiment does not specifically describe the specific structure and implementation principle of the field of view angle prediction model.

[0065] S103: Mapping the current viewing angle information of each user into the semantic space to obtain the current viewing angle semantic information of each user;

[0066] In this embodiment, in order to realize the diversion of the semantic information flow in the semantic space, it is necessary to map the current field of view information of each user into the semantic space, and use the current field of view semantic information of the semantic space to divert the semantic information flow into a public semantic flow and a private semantic flow.

[0067] In some embodiments, the current viewing angle information of each user is mapped to the semantic space to obtain the current viewing angle semantic information of each user, including:

[0068] For each user, define a coefficient matrix with the same size as the panoramic video data;

[0069] Matching the coefficient matrix with the user's current field of view angle information, setting the coefficients in the coefficient matrix corresponding to the field of view angle to a first predetermined value, and setting the coefficients outside the field of view angle to a second predetermined value, to obtain an updated coefficient matrix; wherein the first predetermined value is greater than the second predetermined value;

[0070] Perform feature extraction on the updated coefficient matrix to obtain the current field of view angle features;

[0071] The current field of view angle features are vectorized to obtain the current field of view angle semantic information.

[0072] This embodiment provides a method for converting current field of view information into current field of view semantic information of a semantic space. A coefficient matrix is ​​defined for each user, and the coefficients in the coefficient matrix are set according to the current field of view information of the user, and the coefficients corresponding to the field of view are set to a first predetermined value, and the coefficients corresponding to the field of view are set to a second predetermined value, that is, in the coefficient matrix, the coefficients within the user's field of view (the area the user is watching, which directly affects the user experience) are set to a first predetermined value, and the coefficients not within the user's field of view (the area that the user cannot see, which does not affect the user experience) are set to a second predetermined value, and then the coefficient matrix with the set coefficient values ​​is subjected to feature extraction to obtain the current field of view feature, and then the current field of view feature is vectorized to obtain the current field of view semantic information of the semantic space.

[0073] In some embodiments, the first predetermined value is greater than the second predetermined value, for example, the first predetermined value is 1, and the second predetermined value is 0.1 or a negative number. The coefficient values ​​of the initialized coefficient matrix can all be set to 1, and the coefficients outside the field of view are reset to 0.1, that is, the coefficient matrix updated according to the current field of view is obtained; then, the updated coefficient matrix is ​​continuously averaged and downsampled to obtain the current field of view feature, and the size of the current field of view feature is equal to the size of the semantic feature obtained after the semantic feature extraction of the panoramic video data, and then the current field of view feature is vectorized to obtain the current field of view semantic information.

[0074] S104: Based on the current field of view semantic information of each user, using a preset message splitter, split the semantic information stream of each user into a public sub-semantic stream and a private semantic stream of each user;

[0075] In this embodiment, after the user's current viewing angle information is converted into current viewing angle semantic information, a message splitter is used to perform unified diversion processing on the semantic information streams of all users according to the current viewing angle semantic information.

[0076] Among them, the message segmenter processes the current field of view semantic information of all users, determines the overlapping parts in the semantic information stream corresponding to the current field of view semantic information of all users, and the non-overlapping parts in the semantic information stream corresponding to the current field of view semantic information of each user, and uses the semantic information stream corresponding to the overlapping parts as the public sub-semantic stream of each user, and uses the semantic information stream corresponding to the non-overlapping parts as the private semantic stream of each user. In other words, the message segmenter is used to divide the same video content being watched by all users into the public sub-semantic stream of each user, and the different video content being watched by each user into the private semantic stream of each user. When there is a large overlap in the semantic information between users, more resources can be allocated to the public semantic stream to improve the transmission efficiency of the public part. If there are fewer overlapping parts, the resource allocation of the private semantic stream can be emphasized to meet the personalized viewing needs of users.

[0077] In some embodiments, the message splitter is implemented based on a neural network model, and the current field of view semantic information of multiple users and the semantic information stream of each user can be input into the model, and the model is used to extract the public sub-semantic stream and private semantic stream of each user from the semantic information stream of each user according to the current field of view semantic information of each user. This embodiment does not explain the specific structure and implementation of the message splitter in detail.

[0078] S105: merging the public sub-semantic streams of each user into one public semantic stream, and encoding the public semantic stream based on a preset public semantic stream codebook to obtain a public semantic codeword;

[0079] In this embodiment, after the common sub-semantic streams of each user are separated by the message splitter, the common sub-semantic streams of all users are merged into one common semantic stream, and the common semantic stream is encoded using the common semantic stream codebook to obtain a common semantic codeword. The common semantic stream codebook is implemented based on the codebook of the encoder of the communication system, and the specific codebook content and encoding principle of the encoder are not limited.

[0080] S106: Encode the private semantic stream of each user based on a preset private semantic stream codebook to obtain a private semantic codeword of each user;

[0081] In this embodiment, the private semantic stream codebook is used to encode the private semantic stream of each user to obtain the private semantic codeword of each user. The private semantic stream codebook is implemented based on the codebook of the encoder of the communication system, and the specific codebook content and encoding principle of the encoder are not limited.

[0082] S107: Inputting each user's current field of view semantic information, current channel state information, and historical resource allocation strategy information into a pre-built resource allocation model, and the resource allocation model outputs resource allocation strategies for public semantic codewords and private semantic codewords.

[0083] In this embodiment, after determining the public semantic codewords and the private semantic codewords of each user, the current resource allocation strategy is determined based on the current field of view semantic information, current channel state information, and historical resource allocation strategy of each user using a preset resource allocation model. Among them, the channel state information includes the specific conditions and performance of the channel, and the current channel state information includes the user's current channel matrix and current channel gain, etc. The historical resource allocation strategy includes the power allocation ratio of the public semantic codewords and the private semantic codewords of each user at the last moment, the transmission rate allocation ratio of the public semantic codewords and the private semantic codewords at the last moment, and the channel bandwidth allocation ratio of each user at the last moment.

[0084] In some implementations, the resource allocation strategy determined by the model includes a power allocation ratio of public semantic codewords and private semantic codewords of each user, a transmission rate allocation ratio, and a channel bandwidth allocation ratio of each user.

[0085] In some embodiments, after using a resource allocation model to determine the resource allocation strategy for public semantic codewords and private semantic codewords, the public semantic codewords and each user's private semantic codewords are normalized according to the resource allocation strategy, and the normalized codewords are superimposed, and the superimposed semantic codewords are transmitted via a channel.

[0086] In some implementations, considering that the channel state and the user's field of view angle semantic information have a significant temporal relationship, the resource allocation model is implemented based on a long short-term memory network model (LSTM). This network model can capture the temporal relationship and improve the rationality of resource allocation.

[0087] In some embodiments, the training process of the resource allocation model includes: initializing state information, including initial field of view semantic information of multiple users, initial channel state information, and initial resource allocation strategy, inputting the initialized state information into the resource allocation model based on the long short-term memory network model, and the model outputting the current resource allocation strategy. Inputting the state information and the current resource allocation strategy into a preset discriminant model, and using the discriminant model to calculate a virtual reward for evaluating the quality of the resource allocation strategy determined based on the state information.

[0088] Based on the status information and the determined resource allocation strategy, the real reward is calculated. The calculation process includes:

[0089] According to the channel status of each user and the power allocation of the public semantic codeword and the private semantic codeword, the transmission rate upper limit of the public semantic codeword and the private semantic codeword of each user is calculated, which is expressed as:

[0090]

[0091] in, is the upper limit of the transmission rate of the public semantic codeword of user k, is the upper limit of the transmission rate of the private semantic codeword of user k, B is the channel bandwidth, is the channel gain of user k, p c and are the power of the public semantic codeword and the power of the private semantic codeword of user k, respectively. is the noise power of user k, is the power of the private semantic codeword of user i.

[0092] According to the transmission rate allocation ratio of the public semantic codewords of user k, the transmission rate upper limit of the public semantic codewords and the transmission rate upper limit of the private semantic codewords, the semantic codeword transmission rate R of user k is calculated. k , expressed as:

[0093]

[0094] Among them, α k is the transmission rate allocation ratio of the public semantic codeword of user k. To ensure that all users can successfully decode the public semantic stream, To take the minimum transmission rate upper limit among all users.

[0095] According to the semantic codeword transmission rate of user k, calculate the corresponding transmission delay T k , and converted into a transmission delay score It is expressed as:

[0096]

[0097] Among them, T max is the maximum tolerable delay of the user, is the smoothing coefficient, S k is the amount of codeword data sent to user k.

[0098] According to the channel bandwidth allocation ratio of the user, the restoration quality Q of the panoramic video is calculated using relevant technologies k , and converted into a video quality score It is expressed as:

[0099]

[0100] Among them, Q max , Q min They are respectively the maximum and minimum requirements of users for video quality.

[0101] According to the transmission delay score and video quality score, the real reward of k users is calculated, which is expressed as:

[0102]

[0103] κ is the equilibrium coefficient.

[0104] According to the above process, a certain amount of state information, corresponding resource allocation strategies, virtual rewards and real rewards obtained within a certain period of time are accumulated as training samples for updating the resource allocation model and the discriminant model, and the model parameters are updated based on the training samples, specifically including:

[0105] According to the accumulated training samples of δ moments, the cumulative discounted reward is calculated, which is expressed as:

[0106] G i =R i +γG i+1 ,i=δ-1,δ-2,…,1,0 (8)

[0107] Among them, γ is the discount factor, G i is the discounted reward at time i, R i is the real reward at time i.

[0108] According to the cumulative discount reward and resource allocation strategy information, the advantage value is calculated and expressed as:

[0109] A i =G i -V i ,i=0,1,…,δ-2,δ-1 (9)

[0110] Among them, V i is the virtual reward at time i.

[0111] All state information in the training samples is spliced ​​in the time dimension, and the spliced ​​results are input into the resource allocation model to obtain the corresponding resource allocation strategy, which is expressed as an action probability set. The state information and resource allocation strategy information in the training samples are spliced ​​in the time dimension, and the spliced ​​results are input into the discriminant model to obtain the corresponding virtual rewards, which are expressed as the virtual reward set

[0112] The gradient update algorithm is used to update the parameters of the resource allocation model, which is expressed as:

[0113]

[0114] Among them, P i is the resource allocation strategy information at time i in the training sample, and clip means Clip to the range of (1-∈, 1+∈), where ∈ is the clipping factor.

[0115] The gradient update algorithm is used to update the parameters of the discriminant model, which is expressed as:

[0116]

[0117] The resource allocation model is trained using different field of view semantic information and channel state information. With transmission delay and video recovery quality as rewards, the model can learn the optimal resource allocation decisions under different states, reasonably allocate resources for public semantic streams and private semantic streams under different states, and maximize the user experience quality.

[0118] like Figure 2 , 3 As shown, the embodiment of the present application also provides a panoramic video semantic transmission method, which is applied to a receiving end, including:

[0119] S201: receiving a semantic codeword via a channel;

[0120] S202: Decoding the semantic codeword based on a preset public semantic codebook to obtain a decoded public semantic stream;

[0121] In this embodiment, the receiving end, i.e., the client, receives semantic codewords via a channel and decodes the semantic codewords using a public semantic codebook. Since the received semantic codewords include public semantic codewords and private semantic codewords of all users, the private semantic codewords of all users are used as interference items during decoding, and a public semantic stream can be obtained after decoding.

[0122] S203: using a preset message splitter, splitting the public sub-semantic stream of the current user from the decoded public semantic stream;

[0123] In this embodiment, after decoding the public semantic stream, a message splitter is used to split the public sub-semantic stream corresponding to the current user from one public semantic stream. Specifically, the message splitter first detects the user identifier or tag pre-embedded in the public semantic stream, locates the data portion corresponding to the current user according to the identifier of the current user, separates the data portion from the entire public semantic stream, and reassembles it into the public sub-semantic stream corresponding to the current user.

[0124] S204: performing continuous interference elimination processing based on the semantic codeword and the decoded public semantic stream, removing the decoded public semantic stream, and obtaining the private semantic codeword of each user;

[0125] In this embodiment, based on the received semantic codewords and the decoded public semantic stream, the decoded public semantic stream is regarded as an interference item, and the decoded public semantic stream part is removed from the semantic codeword using continuous interference elimination technology to obtain private semantic codewords for all users.

[0126] S205: decoding the private semantic codewords of each user based on a preset private semantic codebook to obtain a decoded private semantic stream of the current user;

[0127] In this embodiment, after separating the private semantic codewords of all users from the semantic codewords, the private semantic codewords of all users are decoded using the private semantic codebook. Except for the private semantic codeword of the current user, the private semantic codewords of other users are used as interference items. After decoding, the private semantic stream of the current user is obtained.

[0128] S206: merging the public sub-semantic stream of the current user and the decoded private semantic stream of the current user to obtain the semantic information stream of the current user;

[0129] S207: semantically decode the semantic information stream of the current user to obtain restored panoramic video data.

[0130] In this embodiment, after processing the public sub-semantic stream and private semantic stream of the current user, the two are merged into the semantic information stream of the user, and the semantic information stream is semantically decoded based on semantic communication technology to obtain restored panoramic video data, that is, the panoramic video required by the current user.

[0131] In some application scenarios, such as live broadcasts of large-scale events, the server performs semantic encoding on the panoramic video of the event requested by multiple users, generates the semantic information stream required by each user, predicts the current field of view information of each user based on the historical field of view information of each user, and converts it into the current field of view semantic information. Based on the current field of view semantic information, the message splitter is used to split the semantic information stream into a public sub-semantic stream and a private semantic stream for each user, and the public sub-semantic streams of each user are merged into a public semantic stream. The public semantic stream and the private semantic stream are encoded respectively to obtain public semantic codewords and private semantic codewords. The resource allocation model is used to allocate power, rate, and bandwidth resources for the public semantic codewords and the private semantic codewords of each user, and the codewords are sent according to the allocated resources. The client receives the semantic codeword, first decodes the public semantic stream, uses the message splitter to split the public sub-semantic stream of the current user from the public semantic stream, eliminates the public semantic stream as interference from the semantic codeword, obtains the private semantic codeword of each user, and then decodes the private semantic stream of the current user, merges the public sub-semantic stream with the private semantic stream to restore the semantic information stream of the current user, and restores the panoramic video of the event through semantic decoding, presenting a personalized panoramic live broadcast of the event on the client.

[0132] In other application scenarios, in remote travel or virtual scenic spot visits, users hope to view scenic spots around the world without leaving home, and explore from multiple angles according to their personal interests. Traditional communication systems cannot meet users' personalized viewing needs. This application combines semantic communication and rate division multiple access technology, so that the client can obtain a public semantic stream and a personalized private semantic stream. By merging the two parts, the complete and personalized high-quality panoramic guide content can be restored, providing a personalized cultural and travel experience, and ensuring the playback quality and effect in multi-user concurrent scenarios.

[0133] The panoramic video semantic transmission method provided in this embodiment combines semantic communication and rate division multiple access technology in a panoramic video transmission system, and can reasonably allocate the system's power, bandwidth, and rate resources according to the time-varying channel environment based on the understanding of the relationship between different semantic information flows. By utilizing the generalization capabilities of semantic transmitters and semantic receivers based on semantic communication, when multiple users access panoramic videos at the same time, the system can maintain efficient and robust transmission performance in a complex interference and noise environment, significantly improving user experience and network resource utilization.

[0134] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the described method.

[0135] It should be noted that the above is a description of a specific embodiment of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0136] like Figure 4 As shown, the embodiment of the present application also provides a panoramic video semantic transmission device, which is applied to a sending end, including:

[0137] A video semantic coding module is used to semantically code the panoramic video data requested by each user to obtain a semantic information stream corresponding to the panoramic video data of each user;

[0138] A viewing angle prediction module is used to input the historical viewing angle information of each user into a pre-built viewing angle prediction model, and the viewing angle prediction model outputs the current viewing angle information of each user;

[0139] A semantic mapping module is used to map the current viewing angle information of each user into a semantic space to obtain the current viewing angle semantic information of each user;

[0140] A segmentation module, for dividing the semantic information stream of each user into a public sub-semantic stream and a private semantic stream of each user by using a preset message segmentor based on the current field of view semantic information of each user;

[0141] A merging and encoding module, used to merge the public sub-semantic streams of each user into a public semantic stream, and encode the public semantic stream based on a preset public semantic stream codebook to obtain a public semantic codeword;

[0142] A private semantic encoding module, used to encode the private semantic stream of each user based on a preset private semantic stream codebook to obtain the private semantic codeword of each user;

[0143] The resource allocation module is used to input the current field of view semantic information, current channel state information, and historical resource allocation strategy information of each user into a pre-built resource allocation model, and the resource allocation model outputs the resource allocation strategy of the public semantic codeword and the private semantic codeword.

[0144] like Figure 5 As shown, the embodiment of the present application also provides a panoramic video semantic transmission device, which is applied to a receiving end, including:

[0145] A receiving module, used for receiving semantic codewords via a channel;

[0146] A public codeword decoding module, used for decoding the semantic codeword based on a preset public semantic codebook to obtain a decoded public semantic stream;

[0147] A public segmentation module, used to segment the public sub-semantic stream of the current user from the decoded public semantic stream using a preset message segmentor;

[0148] A private segmentation module, used for performing continuous interference elimination processing based on the semantic codeword and the decoded public semantic stream, removing the decoded public semantic stream, and obtaining the private semantic codeword of each user;

[0149] A private codeword decoding module, used to decode the private semantic codewords of each user based on a preset private semantic codebook to obtain a decoded private semantic stream of the current user;

[0150] A merging module, used to merge the public sub-semantic stream of the current user and the decoded private semantic stream of the current user to obtain the semantic information stream of the current user;

[0151] The restoration module is used to perform semantic decoding on the semantic information stream of the current user to obtain restored panoramic video data.

[0152] For the convenience of description, the above devices are described in terms of functions divided into various modules. Of course, when implementing the embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0153] The apparatus of the above-mentioned embodiment is used to implement the corresponding method in the above-mentioned embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0154] Figure 6 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.

[0155] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0156] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0157] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0158] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0159] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0160] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0161] The electronic device of the above embodiment is used to implement the corresponding method in the above embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0162] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0163] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0164] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present application difficult to understand, the known power supply / ground connection with the integrated circuit (IC) chip and other components may or may not be shown in the provided drawings. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform to be implemented in the embodiments of the present application (that is, these details should be fully within the scope of understanding of those skilled in the art). In the case of elaborating specific details (e.g., circuits) to describe exemplary embodiments of the present disclosure, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0165] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0166] The embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of the present disclosure.

Claims

1. A panoramic video semantic transmission method, applied to a sending end, characterized in that: include: Perform semantic encoding on the panoramic video data requested by each user to obtain a semantic information stream corresponding to the panoramic video data of each user; Inputting the historical viewing angle information of each user into a pre-built viewing angle prediction model, and the viewing angle prediction model outputs the current viewing angle information of each user; Mapping the current viewing angle information of each user into the semantic space to obtain the current viewing angle semantic information of each user; Based on the current field of view semantic information of each user, a preset message splitter is used to split the semantic information flow of each user into a public sub-semantic flow and a private semantic flow of each user; The public sub-semantic streams of each user are merged into a public semantic stream, and the public semantic stream is encoded based on a preset public semantic stream codebook to obtain a public semantic codeword; Based on a preset private semantic stream codebook, encode each user's private semantic stream to obtain each user's private semantic codeword; The current field of view semantic information, current channel state information, and historical resource allocation strategy information of each user are input into a pre-built resource allocation model, and the resource allocation model outputs the resource allocation strategy of the public semantic codeword and the private semantic codeword.

2. The method according to claim 1, characterized in that: Map the current viewing angle information of each user to the semantic space to obtain the current viewing angle semantic information of each user, including: For each user, define a coefficient matrix having the same size as the panoramic video data; Matching the coefficient matrix with the user's current field of view angle information, setting the coefficients in the coefficient matrix corresponding to the field of view angle to a first predetermined value, and setting the coefficients outside the field of view angle to a second predetermined value, to obtain an updated coefficient matrix; wherein the first predetermined value is greater than the second predetermined value; Performing feature extraction on the updated coefficient matrix to obtain current field of view angle features; The current field of view angle feature is vectorized to obtain the current field of view angle semantic information.

3. The method according to claim 1, characterized in that: Based on the current field of view semantic information of each user, the preset message splitter is used to split the semantic information flow of each user into a public sub-semantic flow and a private semantic flow of each user, including: The semantic information stream corresponding to the overlapping part of the current field of view semantic information of each user is taken as the public sub-semantic stream, and the semantic information stream corresponding to the non-overlapping part of the current field of view semantic information of each user is taken as the private semantic stream of each user.

4. The method according to claim 1, characterized in that The resource allocation strategy includes: a power allocation ratio of the public semantic codeword and the private semantic codeword of each user, a transmission rate allocation ratio, and a channel bandwidth allocation ratio of each user.

5. The method according to claim 1, characterized in that The training method of the resource allocation model includes: Constructing training samples; wherein the training samples include a certain amount of state information within a certain period of time, the obtained resource allocation strategy, the calculated virtual rewards and real rewards; the state information includes field of view semantic information, channel state information and historical resource allocation strategy information; The parameters of the resource allocation model are updated based on the training samples.

6. The method according to claim 5, characterized in that The real reward is determined according to the transmission delay score and the video quality score calculated after executing the resource allocation strategy.

7. The method according to claim 4, characterized in that Also includes: According to the resource allocation strategy, the public semantic codeword and the private semantic codeword of each user are normalized, and the normalized codewords are superimposed, and the superimposed semantic codewords are transmitted via a channel.

8. A panoramic video semantic transmission method, applied to a receiving end, characterized in that: include: receiving a semantic codeword via a channel; Decoding the semantic codeword based on a preset public semantic codebook to obtain a decoded public semantic stream; Using a preset message splitter, splitting the public sub-semantic stream of the current user from the decoded public semantic stream; Performing continuous interference elimination processing based on the semantic codeword and the decoded public semantic stream, removing the decoded public semantic stream, and obtaining the private semantic codeword of each user; Decoding the private semantic codewords of each user based on a preset private semantic codebook to obtain a decoded private semantic stream of the current user; Merge the public sub-semantic stream of the current user and the decoded private semantic stream of the current user to obtain the semantic information stream of the current user; The semantic information stream of the current user is semantically decoded to obtain restored panoramic video data.

9. A panoramic video semantic transmission device, applied to a sending end, characterized in that: include: A video semantic coding module is used to semantically code the panoramic video data requested by each user to obtain a semantic information stream corresponding to the panoramic video data of each user; A viewing angle prediction module, used to input the historical viewing angle information of each user into a pre-built viewing angle prediction model, and the viewing angle prediction model outputs the current viewing angle information of each user; A semantic mapping module is used to map the current viewing angle information of each user into a semantic space to obtain the current viewing angle semantic information of each user; A segmentation module, for dividing the semantic information stream of each user into a public sub-semantic stream and a private semantic stream of each user by using a preset message segmentor based on the current field of view semantic information of each user; A merging and encoding module, used to merge the public sub-semantic streams of each user into a public semantic stream, and encode the public semantic stream based on a preset public semantic stream codebook to obtain a public semantic codeword; A private semantic encoding module, used to encode the private semantic stream of each user based on a preset private semantic stream codebook to obtain the private semantic codeword of each user; The resource allocation module is used to input the current field of view semantic information, current channel state information, and historical resource allocation strategy information of each user into a pre-built resource allocation model, and the resource allocation model outputs the resource allocation strategy of the public semantic codeword and the private semantic codeword.

10. A panoramic video semantic transmission device, applied to a receiving end, characterized in that: include: A receiving module, used for receiving semantic codewords via a channel; A public codeword decoding module, used for decoding the semantic codeword based on a preset public semantic codebook to obtain a decoded public semantic stream; A public segmentation module, used to segment the public sub-semantic stream of the current user from the decoded public semantic stream using a preset message segmentor; A private segmentation module, used for performing continuous interference elimination processing based on the semantic codeword and the decoded public semantic stream, removing the decoded public semantic stream, and obtaining the private semantic codeword of each user; A private codeword decoding module, used to decode the private semantic codewords of each user based on a preset private semantic codebook to obtain a decoded private semantic stream of the current user; A merging module, used to merge the public sub-semantic stream of the current user and the decoded private semantic stream of the current user to obtain the semantic information stream of the current user; The restoration module is used to perform semantic decoding on the semantic information stream of the current user to obtain restored panoramic video data.

Citation Information

Patent Citations

  • Encoding method, decoding method, encoding device, decoding device and electronic equipment

    CN117692094A

  • Rate-splitting multiple access (RSMA)

    GB202408312D0