Adaptive video stream real-time encoding transmission method in satellite network

CN116418387BActive Publication Date: 2026-10-09BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310024910.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2026-10-09
Estimated Expiration
2043-01-09

AI Technical Summary

Benefits of technology

[0075] This invention achieves adaptive real-time encoding of video streams. Based on satellite edge services, it models the mathematical relationships between tasks, links, and encoding strategies by exploring the interaction mechanisms, and then provides a real-time encoding strategy for video streams. Through this invention, the edge service adaptively determines the size of the encoded packets and performs overall scheduling of packet transmission, minimizing the total task latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116418387B_ABST
    Figure CN116418387B_ABST
Patent Text Reader

Abstract

The application provides a real-time coding transmission method for adaptive video stream in a satellite network. Influenced by the orbit of the satellite, the inter-satellite link and the satellite-ground link can only be maintained for a period of time, the link is unstable, is easily affected by the external environment, and data errors, packet loss and other conditions frequently occur. Meanwhile, satellite applications will generate massive data, which is difficult to be directly transmitted to the ground, and urgent need for on-orbit processing. The application is based on the fact that the source video end has reasonably sliced the video, and the task data volume and the link error rate are known, how to adaptively determine the size of the coding packet for each transmission, and overall scheduling of the transmission of the coding packet, so as to minimize the total time delay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of satellite networks. Background Technology

[0002] The construction of satellite networks can provide communication support for human interstellar exploration and effectively enhance a nation's global influence. Satellite networks have not only become a popular research area but also a focal point of great power competition. Despite the current rapid development of satellite networks, they remain a scarce resource.

[0003] In practical satellite applications, due to the influence of the satellite's own orbit, inter-satellite links and satellite-to-ground links can only be maintained for a limited time, and these links are unstable and easily affected by the external environment, resulting in frequent data errors and packet loss. Simultaneously, satellite applications generate massive amounts of data, which cannot all be directly transmitted back to the ground, urgently requiring on-orbit processing. On-orbit processing and efficient backhaul of satellite data have become major challenges in the current construction of satellite constellations and the development of satellite internet. Summary of the Invention

[0004] Based on satellite edge services, this invention delves into the relationship between data encoding schemes, link quality, and data transmission success rate. It investigates the mechanism of interaction between link quality and encoding scheme under minimum latency requirements and proposes a real-time dual-layer encoding transmission method for video streams based on adaptive grouping.

[0005] This invention addresses the challenge of adaptively determining the size of encoded packets for each transmission and performing overall scheduling of packet transmission, given that the video source has already been rationally segmented and the task data volume and link error rate are known. The invention utilizes the LT code, with the following degree function:

[0006]

[0007] Where d represents the degree, k represents the number of source symbols, d = 1, ..., k, and ∝ is the proportionality coefficient with a value of 0.4.

[0008] θ(d) is an optimized version of the binary exponential degree distribution (probability normalized):

[0009]

[0010] b′(d) performs a subtle optimization on the binary exponential degree distribution, making the probability of d=2 1 / 2, which is the highest probability when d=2.

[0011] b'(d) is an optimized version of the binary exponential degree distribution (only the degree of the maximum probability is changed):

[0012]

[0013] μ(d) represents the robust solitary wave distribution:

[0014]

[0015] Where d = 1, 2, ..., k, ρ(d) is the ideal solitary wave distribution, and τ(d) is the compensation function.

[0016] Compensation function τ(d):

[0017]

[0018] in:

[0019]

[0020] The compensation function needs to introduce two parameters, c and δ, to ensure that the number of check codes with a degree of 1 in each decoding loop is s, not 1. c is a constant greater than zero, set to 0.6 in this invention, and δ = 0.5, representing the upper bound of decoding failure.

[0021] The overall architecture diagram of this system is as follows: Figure 1 :

[0022] The coding flowchart for this system is as follows: Figure 2 :

[0023] In the figure, the part enclosed by the solid line within the dashed box on the left is a video slice. Ω(x) is the degree function during fountain code encoding. After completing one decoding operation, the receiver feeds back the decoding result to the transmitter. The transmitter updates the content of the decoding probability table based on the feedback information and modifies the size of the new packet and incremental packet for the next transmission, thereby achieving the effect of adaptive encoding.

[0024] The specific steps of this invention are as follows:

[0025] (1) Assign a default redundancy value to each video slice.

[0026] Fountain coding, as a forward error correction scheme, recovers lost data by introducing a certain amount of redundant coding. The formula for redundancy is:

[0027]

[0028] Where M represents the encoded symbol received by the receiver, and m represents the original symbol. For real-time video streams, the redundancy of encoding should be reduced when the link condition is good and increased when the link condition is poor. The first step of this method is to assign an initial default redundancy of 0.16 to all video slices.

[0029] In the following steps, the redundancy of the encoding will be dynamically updated based on the link status to improve the transmission efficiency of the system.

[0030] (2) The sending end calculates the size of the initial packet and incremental packet to be sent in this round.

[0031] The sending end counts all video segments to be sent. For a newly arrived video segment, it sends an initial probe packet M0. For previously sent video segments that have not yet received acknowledgment, it determines the incremental packets that need to be resent based on the decoding probability table. If the number of transmissions for a certain segment is greater than K... max If an error occurs, an error message is sent to the upper layer, and this video segment is skipped.

[0032] The selection of new packages and incremental packages is achieved by solving the following equation:

[0033]

[0034]

[0035] Where m l M is the incremental packet sent during the l-th transmission of this video slice. send The average number of symbols transmitted when transmitting video slices:

[0036]

[0037] Where M i p is the size of the data packet sent in the i-th transmission. i This represents the probability that the (i-1)th decoding attempt fails while the ith decoding attempt succeeds.

[0038] K send Average number of transmissions per data packet:

[0039]

[0040] Δm is the minimum value of a single incremental packet transmission. Δm is a fixed value, which can be set to 64 in a typical link. ave The number of transmissions required to transmit one data packet on average, which is limited by the system.

[0041] The number of data packets transmitted for each path is calculated based on the decoding probability table. The path that satisfies the above constraint formula and minimizes the average number of transmitted symbols is selected as the basis for selecting the size of new packets and incremental packets.

[0042] (3) Encoding video slices using LT codes

[0043] First, the degree value d is generated using the degree function ω(d). Then, d primitive symbols are selected from each video slice. Let a video slice contain L primitive symbols, and the video slice has been transmitted K times (K... <K max Instead of the previous random selection method, we set a separate selection probability for each packet. The selection probability of the j-th original symbol is:

[0044]

[0045] Where K max Let this be the maximum number of transmissions allowed for a single video slice as defined by the system. Then the probability of selecting P(j) can be simplified to:

[0046]

[0047] It can be seen that as the number of transmissions K increases, the probability of μ will increase, therefore The proportion will increase, and the larger the value of j, the greater the proportion will be. The smaller the value of e, the better. j-L The larger the value, the better. This can be viewed as selecting packets with earlier arrival times, e j*L This can be viewed as selecting packets with later arrival times, i.e., newly arrived packets.

[0048] Therefore, when the number of transmissions K is small, the probability of (1-μ) is larger, meaning it tends to choose packets with later numbers, i.e., newly arrived packets. When the number of transmissions is large, the probability of μ is larger, meaning it tends to choose packets with earlier numbers, i.e., packets that arrived earlier. (The denominator is...) This means summing the selection probabilities of all packets in a video group, which means normalizing the selection probability of each packet.

[0049] Using this probability function, the fountain code will tend to select the most recently arrived packet during packet selection encoding. Therefore, the most recently arrived packet will be decoded earlier than earlier arriving packets, thus meeting the system's real-time encoding requirements. Furthermore, as the number of transmissions K increases, the probability of selecting earlier arriving data packets will increase, thereby increasing the decoding success rate of video segments.

[0050] (4) Update the coding redundancy of the fountain code.

[0051] First, estimate the available bandwidth of the current network, using the following formula:

[0052]

[0053] Among them, S t Let T be the size of the encoded packet sent at time t. tThe transmission time of the encoded packet sent at time t.

[0054] The redundancy update function is as follows:

[0055]

[0056] Where γ is the scaling factor, set to 0.5, Z is the default redundancy, set to 0.16, and V is the video bitrate. Since most videos use H.264 encoding, V is set to a fixed value of 2.5 Mb / s.

[0057] The current network bandwidth is greater than or equal to the video bitrate, i.e., VB. t When the value is ≤0, it indicates that the current network is good and the next video slice can be encoded using the default redundancy. When the current network bandwidth is less than the video bitrate, it indicates that the current network is poor and the encoding redundancy needs to be increased to improve the recovery rate of the fountain code for lost data.

[0058] (5) The receiving end performs incremental decoding on the received data.

[0059] The receiving end extracts the new packet or incremental packet corresponding to each video slice based on the data frames sent by the sending end. For new packets, they are directly decoded; for incremental packets, the received data is accumulated. After merging the symbols, attempt to decode.

[0060] The receiver employs Enhanced Belief Propagation (EBP) to simplify decoding. Traditional BP decoding begins decoding whenever the receiver receives a coded packet with a degree of 1, decrementing the degree of all participating coded packets by one. Compared to Gaussian elimination (GE), BP decoding is faster, but its success rate is lower due to its heavy reliance on packets with a degree of 1. This invention uses EBP decoding. Even after a BP decoding failure, the generated matrix of the coded packet may be of full column rank or may contain submatrices of full column rank. The data packets corresponding to the columns of these full-rank submatrices are theoretically translatable. Therefore, the BP decoding algorithm possesses residual translatability.

[0061] The EBP algorithm will be explained with a specific example below:

[0062] Let G be the generator matrix when BP decoding fails (there are no nodes with degree 1):

[0063]

[0064] The update step of the encoded packet during the decoding iteration is called forward propagation, and the assignment step of the data packet is called backward propagation. The execution process of EBP is as follows: Figure 3 As shown:

[0065] Assuming X3 is a known packet, after one forward propagation, we arrive at (b). At this point, the first encoded packet becomes a packet with a degree value of 1. Then, we perform one backward propagation to (c). At this point, the second data packet is known. We perform a second forward propagation to update the values ​​of the second and third encoded packets. We propagate in this way in a loop until we get Figure (e). Then, we perform one more forward propagation: Y2 + Y1 + Y3 + Y1 + X1 = Y3 + Y2 + X1 = 0. Y3 and Y2 are both known values, so we can solve for X0.

[0066] Hypothesis packet selection strategy: To minimize the degree of encoded packets, the degree values ​​of data packets can be sorted from largest to smallest. The packet with the highest degree value is then added to the hypothesis set, and decoding is attempted. This strategy significantly improves the decoding success rate while maintaining low complexity. It is also insensitive to code length and performs well even with shorter fountain codes.

[0067] After decoding is complete, the feedback information of successful or failed decoding is sent back to the sending end in the form of 0 / 1 bits (0 represents success, 1 represents failure).

[0068] (6) The sending end updates the decoding probability table

[0069] After receiving feedback from the receiver, the sender updates the decoding probability table using the following strategy:

[0070] Step 1 (Initialization) N k =0, N total =0

[0071] Step 2 (Update) N k =β·N k +ΔN k

[0072] N total =β·N total +ΔN

[0073] Step 3 (obtain the probability) p k =N k / N total

[0074] Where, N k ΔN is the number of symbols used by the decoder for decoding. k To use N k The number of successfully decoded symbols, N total ΔN represents the total number of successfully decoded data packets, ΔN represents the total number of successfully decoded data packets in this transmission frame, and β represents the weighting factor, which is set to 0.3 in this invention. This update method can reduce the proportion of old feedback data and improve the real-time performance of the system.

[0075] This invention achieves adaptive real-time encoding of video streams. Based on satellite edge services, it models the mathematical relationships between tasks, links, and encoding strategies by exploring the interaction mechanisms, and then provides a real-time encoding strategy for video streams. Through this invention, the edge service adaptively determines the size of the encoded packets and performs overall scheduling of packet transmission, minimizing the total task latency. Attached Figure Description

[0076] Figure 1 System Architecture Diagram

[0077] Figure 2 Coding Flowchart

[0078] Figure 3 BP decoding process containing unknowns Detailed Implementation

[0079] The following example uses a specific scenario to explain the operation of this invention.

[0080] Specific parameters: Using a communication satellite as the source video terminal, a 20GB H.264 encoded video file is transmitted to the ground station. The default redundancy of the encoding is set to Z=0.16, the initial probe packet size is M0=1024, and the minimum value of one incremental packet transmission is Δm=64.

[0081] 1. The communication satellite divides the video into 1000 video slices. The fountain code encoder encodes these 1000 video slices and sends an initial probe packet of size 1000*M0 to the ground station (each video slice needs to be sent, so the total size is 1000*M0).

[0082] The degree function of the fountain code is:

[0083]

[0084] Where d represents degree, d = 1,...,k, and ∝ is the proportionality coefficient, set to 0.4. The meanings of the other variables have been explained when introducing the system model, and will not be introduced here again.

[0085] First, the degree value d is generated using the degree function ω(d). Then, d primitive symbols are selected from each video slice. Let a video slice contain L primitive symbols, and the video slice has been transmitted K times (K... <K max (We set a separate selection probability for each packet, replacing the previous random selection method), then the selection probability of the j-th original symbol is:

[0086]

[0087] 2. After the communication satellite transmission is completed, estimate the available bandwidth B of the current network (at time t). t ,in:

[0088]

[0089] S t T represents the size of the currently transmitted encoded packet. t This represents the transmission time of the currently transmitted encoded packet.

[0090] And update the redundancy for the next encoded transmission according to the redundancy update function:

[0091]

[0092] Where γ is the scaling factor, set to 0.5, Z is the default redundancy, set to 0.16, and V is the video bitrate. Since most videos use H.264 encoding, the bitrate is set to a fixed 2.5Mb / s.

[0093] 3. After receiving the transmitted data, the ground station uses incremental decoding, which combines previously received data with the current data for decoding. The ground station uses enhanced belief propagation to simplify decoding. Assume the decoding generator matrix at the current moment is:

[0094]

[0095] The decoding process is as follows: Figure 3 As shown:

[0096] Assuming X3 is a known packet, after one forward propagation, we reach (b). At this point, the first encoded packet has become a packet with a degree value of 1. Then, we perform a backward propagation to (c). Now, the second data packet is known. We perform a second forward propagation to update the values ​​of the second and third encoded packets. This process is repeated until we obtain Figure (e). Then, we perform another forward propagation: Y2 + Y1 + Y3 + Y1 + X1 = Y3 + Y2 + X1 = 0. Since Y3 and Y2 are both known values, we can solve for X0.

[0097] 4. After decoding is complete, the ground station sends the decoding result back to the communication satellite. The communication satellite updates its decoding probability table based on the decoding information. The update algorithm is as follows:

[0098] Step 1 (Initialization) N k =0, N total =0

[0099] Step 2 (Update) N k =β·N k +ΔN k

[0100] N total=β·N total +ΔN

[0101] Step 3 (obtain the probability) p k =N k / N total

[0102] Where, N k ΔN is the number of symbols used by the decoder for decoding. k To use N k The number of successfully decoded symbols, N total ΔN represents the total number of successfully decoded data packets in the current transmission frame, and β is the weighting factor, set to 0.3. This update method can reduce the proportion of old feedback data and improve the real-time performance of the system.

[0103] 5. The sending end calculates the size of the encoded packet to be sent in the next time slot.

[0104] Suppose the decoding probability table at a certain moment is as follows:

[0105] Decoding success probability 0 0.0001 0.0349 0.258 0.459 0.862 1

[0106] The selection of the new packet and the incremental packet at the next moment is achieved by solving the following equation:

[0107]

[0108]

[0109] Where m l M is the incremental packet sent during the l-th transmission of this video slice. send The average number of symbols transmitted when transmitting video slices:

[0110]

[0111] Where M l p is the size of the data packet sent in the l-th transmission. i This represents the probability that the (i-1)th decoding attempt fails while the ith decoding attempt succeeds.

[0112] K send Average number of transmissions per data packet:

[0113]

[0114] Δm is the minimum value of a single incremental packet transmission. Δm is a fixed value, set to 64 in this scenario. ave The system limits the average number of transmissions required to transmit one data packet, which is 16.

[0115] The number of data packets transmitted for each path is calculated based on the decoding probability table. The path that satisfies the above constraint formula and minimizes the average number of transmitted symbols is selected as the basis for selecting the size of new packets and incremental packets.

Claims

1. A real-time encoding and transmission method for adaptive video streams in satellite networks, characterized by: The fountain code used is the LT code, and its degree function is: ; Where d represents the degree, k represents the number of source symbols, and d = 1, ..., k. This is the proportionality constant, with a value of 0.4; An optimized version of the binary exponential degree distribution: ; in The binary exponential degree distribution was optimized so that the probability of d=2 is 1 / 2, which is the highest probability when d=2. ; Robust solitary wave distribution: ; Where d = 1, 2, ..., k, For an ideal solitary wave distribution, For compensation functions; Compensation function : ; in: ; The compensation function needs to introduce two parameters c and This ensures that the number of check codes with a degree of 1 is s, not 1, in each decoding loop; c is a constant greater than zero, set to 0.

6. , indicating the upper bound of decoding failure; The specific steps are as follows: (1) Assign a default redundancy value to each video slice. Fountain codes, as a forward error correction scheme, recover lost data by introducing redundant coding. The formula for redundancy is: ; Where M is the encoded symbol received by the receiver and m is the original symbol; for real-time video streams, the redundancy of encoding should be reduced when the link is in good condition and increased when the link is in poor condition. An initial default redundancy of 0.16 is assigned to all video slices. (2) The sending end calculates the size of the initial packet and incremental packet to be sent in this round. The sending end counts all video segments to be sent. For a newly arrived video segment, it sends an initial probe packet M0. For previously sent video segments that have not yet received acknowledgment, it determines the incremental packets that need to be resent based on the decoding probability table. If the number of transmissions for a certain segment is greater than K... max If the error occurs, an error message is sent to the upper layer, and this video segment is skipped. The selection of new packages and incremental packages is achieved by solving the following equation: ; s.t. ; in For the first The incremental packet sent during the next transmission of this video slice The average number of symbols transmitted when transmitting video slices: ; in Let be the size of the data packet sent in the i-th transmission. This represents the probability that the (i-1)th decoding attempt fails while the ith decoding attempt succeeds. Average number of transmissions per data packet: ; This is the minimum value for a single incremental packet transmission. It is a fixed value, set to 64. The system limits the average number of transmissions required to transmit one data packet. The number of data packets transmitted for each path is calculated based on the decoding probability table. The path that satisfies the above constraint formula and minimizes the average number of transmitted symbols is selected as the basis for selecting the size of new packets and incremental packets. (3) Encoding video slices using LT codes First, use the degree function. Generate a degree value d, then select d original symbols from each video slice. Suppose a video slice contains L original symbols, and the video slice has been transmitted K times, K <K max ; The probability of selecting the j-th original symbol is: ; Where K max Let be the maximum number of transmissions allowed for a single video slice as defined by the system; ,So The selection probability can be simplified to: ; It can be seen that as the number of transmissions K increases, The probability will increase, therefore The proportion will increase, and the larger the value of j, the greater the proportion will be. The smaller the value, the better. The larger the value, the better. This can be viewed as selecting packets with earlier arrival times. This can be viewed as selecting packets with later arrival times, i.e., newly arrived packets; Therefore, when the number of transmissions K is small... The probability is higher, meaning it tends to choose packets with later numbers, i.e., newly arrived packets. This is especially true when the number of transmissions is large. The probability of this is higher, meaning there's a greater tendency to choose the package with the earlier number, i.e., the package that arrived earlier; the denominator... This means summing the selection probabilities of all packets in a video group, which means normalizing the selection probability of each packet. By using this probability function, when performing packet selection encoding, the fountain code will tend to select the most recently arrived packet. Thus, the most recently arrived packet will be decoded earlier than the earlier arrived packet with a greater probability, in order to meet the real-time encoding requirements of the system. As the number of transmissions K increases, the probability of selecting the earlier arrived data packet will increase, thereby increasing the decoding success rate of the video slice. (4) Update the coding redundancy of the fountain code. First, estimate the available bandwidth of the current network, using the following formula: ; in, Let t be the size of the encoded packet sent at time t. The transmission time of the encoded packet sent at time t; The redundancy update function is as follows: ; in, Z is the scaling factor, set to 0.5; Z is the default redundancy, set to 0.16; V is the video bitrate. Since most videos use H.264 encoding, V is set to a fixed value of 2.5 Mb / s. The current network bandwidth is greater than or equal to the video bitrate, that is When the current network bandwidth is less than the video bitrate, it indicates that the current network is good and the next video slice can be encoded using the default redundancy. When the current network bandwidth is less than the video bitrate, it indicates that the current network is poor and the encoding redundancy needs to be increased to improve the recovery of lost data by the fountain code. (5) The receiving end performs incremental decoding on the received data. The receiving end extracts the new packet or incremental packet corresponding to each video slice based on the data frames sent by the sending end. For new packets, they are directly decoded; for incremental packets, the received data is accumulated. After merging the symbols, attempt to decode; The receiver uses Enhanced Belief Propagation (EBP) to simplify decoding. The hypothesis packet selection strategy is as follows: data packets are sorted from largest to smallest degree, and the packet with the largest degree value is selected each time and added to the hypothesis set for decoding. After decoding, the success or failure feedback information is sent back to the sender in the form of 0 / 1 bits; 0 represents success, and 1 represents failure. (6) The sending end updates the decoding probability table After receiving feedback from the receiver, the sender updates the decoding probability table using the following strategy: Step 1 Initialization: ; Step 2 Update: ; ; Step 3 yields the probability: ; in, The number of symbols used by the decoder for decoding. For use The number of symbols successfully decoded. The total number of data packets successfully decoded. This represents the total number of data packets successfully decoded in this transmission frame. The weighting factor is set to 0.3.

Citation Information

Patent Citations

  • Layer-aware forward error correction encoding and decoding method, encoding apparatus, decoding apparatus, and system thereof

    CN102469311A

  • Self-adaptive fountain coding method based on multimedia broadcast multicast service

    CN102684893A