A communication method and apparatus

By having terminal devices provide feedback on video frame parameters, network devices optimize downlink transmission strategies, solving the transmission efficiency problem of large video frames in XR services and improving user experience.

CN116114254BActive Publication Date: 2025-11-07HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080103785.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-14
Publication Date
2025-11-07
Estimated Expiration
2040-09-14

AI Technical Summary

Technical Problem

In extended reality (XR) service scenarios, the amount of video frame data transmitted downlink is large. How to optimize the downlink transmission method to improve transmission efficiency and user experience?

Method used

After receiving a video frame, the terminal device determines the video frame parameters and feeds them back to the network device so that the network device can optimize the transmission of downlink video frames. Video frame parameters include spread latency, inter-frame spacing, packet loss, latency, and base/enhancement layer parameters, etc. The network device adjusts its transmission strategy based on these parameters.

Benefits of technology

By leveraging feedback from terminal devices, network equipment can gain a more accurate understanding of downlink video frame reception, optimize transmission strategies, reduce latency and packet loss, and improve video transmission quality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116114254B_ABST
    Figure CN116114254B_ABST
Patent Text Reader

Abstract

A communication method and device, the method comprising: receiving, by a terminal device, at least one video frame from a network device; determining, by the terminal device, a video frame parameter according to the at least one video frame; and sending, by the terminal device, the video frame parameter to the network device, so that the network device can learn about a reception condition of a downlink video frame, thereby facilitating optimization of transmission of the downlink video frame.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of communication, and in particular to a communication method and device. BACKGROUND

[0002] In an extended reality (XR) service scenario, sensors in a helmet can perceive the position, motion change and the like of a user, and generate user information including a visual angle, a line of sight, a motion rate and the like. The user information can be transmitted to an XR server in an uplink transmission manner. The path of the uplink transmission can be helmet-> terminal device-> wireless network-> XR server. The XR server can generate new video data according to the position, line of sight and the like of the user, in combination with a scene in which the user is in a game or a real scene. The new video data can be transmitted to the helmet in a downlink transmission manner. The path of the downlink transmission can be XR server-> wireless network-> terminal device-> helmet, and finally the above-mentioned video data is displayed to the user through the helmet. The wireless network is, for example, a 3rd generation partnership project (3GPP) network, such as a long term evolution (LTE) or a 5th generation (5G) network and the like. Since the data volume of a video frame of the downlink transmission is usually large, how to optimize the downlink transmission manner is a technical problem to be solved by the present application. SUMMARY

[0003] The present application provides a communication method and device. After receiving a downlink video frame, a terminal device can feed back a frame parameter of the downlink video frame to a network device, so as to facilitate optimization of transmission of the downlink video frame.

[0004] In a first aspect, a communication method is provided. The execution subject of the method is a terminal device, and can also be a component (chip, circuit or the like) arranged in the terminal device. The method comprises: receiving, by the terminal device, at least one video frame from a network device; determining, by the terminal device, a video frame parameter according to the at least one video frame; and sending, by the terminal device, the video frame parameter to the network device. Optionally, the terminal device can report the reception of the video frame, i.e. the above-mentioned video frame parameter, in the granularity of the video frame.

[0005] By implementing the above-mentioned method, the terminal device determines a video frame parameter according to a received video frame, and sends the video frame parameter to the network device, so that the network device can know the reception of the downlink video frame, and facilitate optimization of downlink transmission of the video frame.

[0006] In a possible design, the video frame parameter comprises at least one of the following: an extended delay parameter, an interframe distance parameter, a packet loss parameter, a late arrival parameter, a base layer and an enhanced layer parameter.

[0007] By implementing the method, the video frame parameter reported by the terminal device intuitively and accurately shows the reception of the downlink video frame in different dimensions, facilitating the optimization of the downlink video frame transmission by the network device.

[0008] Optionally, the extended delay parameter is used to indicate at least one of the following: an extended delay of a first video frame, an average extended delay of a plurality of video frames, a maximum extended delay of the plurality of video frames, a minimum extended delay of the plurality of video frames, and a variance of the extended delay of the plurality of video frames; wherein the extended delay is a time length between a first data packet of a video frame successfully received by the terminal device and a last data packet of the video frame successfully received by the terminal device, and the first video frame or the plurality of video frames are one or more video frames in at least one video frame received by the terminal device.

[0009] By implementing the method, the terminal device can periodically report the extended delay parameter, and / or report the extended delay parameter under different conditions, and the reporting condition or the reported content of the extended delay parameter can or can not correspond. For the terminal device side, the content and / or manner of reporting the extended delay parameter can be flexibly set to meet various reporting requirements.

[0010] Optionally, the interframe distance parameter is used to indicate at least one of the following: an interframe distance of adjacent video frames, an average value of a plurality of interframe distances, a maximum value of the plurality of interframe distances, a minimum value of the plurality of interframe distances, and a variance of the plurality of interframe distances; wherein the terminal device receives a plurality of video frames from the network device.

[0011] By implementing the method, the network device can learn the interframe distance of the downlink video frame. For example, if it is found that the interframe distance of the video frame is unstable, the network device can subsequently increase the size of the buffer to try to ensure that the video frames are sent at the same interval.

[0012] Optionally, the packet loss parameter comprises at least one of the following: a number of video frames with packet loss, a proportion of video frames with packet loss, and continuous K video frames with packet loss, wherein K is a positive integer greater than or equal to 1.

[0013] By implementing the method, the network device can learn the packet loss of the downlink video frame. If the network device finds that there are fewer video frames with packet loss, but multiple packets are lost once packet loss occurs, the network device can determine that the wireless signal deep fade time is long, and can replace a new frequency band to send the video frame.

[0014] In a possible design, the terminal device determines whether the packet data convergence protocol (PDCP) sequence numbers (SNs) of the data packets comprised in the received video frame are continuous; and determines that the video frame is lost when the PDCP SNs of the data packets comprised in the received video frame are not continuous.

[0015] By implementing the method, the terminal device can determine whether the video frame is lost by determining whether the PDCP SNs are continuous, and the implementation is easy.

[0016] Optionally, the late parameter indicates at least one of the following: a number of late video frames, a proportion of late video frames, a time difference between an actual receiving moment and a correct receiving moment of one late video frame, an average of time differences between actual receiving moments and correct receiving moments of a plurality of late video frames, a maximum of time differences between actual receiving moments and correct receiving moments of the plurality of late video frames, a minimum of time differences between actual receiving moments and correct receiving moments of the plurality of late video frames, and a variance of time differences between actual receiving moments and correct receiving moments of the plurality of late video frames.

[0017] By implementing the method, the network device can learn the late situation of the downlink video frame, and optimize transmission of the downlink video frame.

[0018] In a possible design, each video frame comprises base layer data and enhancement layer data; and the base layer and enhancement layer parameters comprise at least one of the following: a time difference between receiving the base layer data and the enhancement layer data in a second video frame, the second video frame being a video frame in the at least one video frame; a number of third video frames, a proportion of third video frames, or at least one of a number of enhancement layer packet losses in the third video frames, the third video frames being video frames in which the base layer data has no packet loss and the enhancement layer data has packet loss, and the third video frames being video frames in the at least one video frame; a number of fourth video frames, a proportion of fourth video frames, or at least one of a number of base layer packet losses in the fourth video frames, the fourth video frames being video frames in which the enhancement layer data has no packet loss and the base layer data has packet loss, and the fourth video frames being video frames in the at least one video frame.

[0019] By implementing the method, the network device can find that the time difference between the base layer and the enhancement layer of the same video frame is too large, and can subsequently reduce the transmission interval of the two.

[0020] Optionally, the video frame after network coding comprises one or more data slices, and the video frame parameter further comprises at least one of the following: a number of data slices required by the terminal device to successfully decode one video frame and / or a data amount of the data slices; a number of data slices required by the terminal device to successfully decode a base layer of one video frame and / or a data amount of the data slices; a number of data slices required by the terminal device to successfully decode an enhancement layer of one video frame and / or a data amount of the data slices.

[0021] By implementing the method, the network device can adjust the redundancy rate of the base layer or the enhancement layer in the network coding process according to the report of the terminal device, and optimize the transmission of the downlink video frame.

[0022] In a second aspect, a communication method is provided, a subject of the method is a network device, and the method can also be a component (chip, circuit, or other component) configured in the network device, and the method comprises: the network device sending at least one video frame to a terminal device; and the network device receiving a video frame parameter from the terminal device, the video frame parameter being determined according to the at least one video frame. Optionally, the video frame parameter can be reported in a video frame granularity.

[0023] By implementing the method, the network device can obtain the reception of the downlink video frame, and facilitate the optimization of the transmission process of the downlink video frame.

[0024] In a possible design, the video frame parameter comprises at least one of the following: an extended delay parameter, an interframe distance parameter, a packet loss parameter, a late arrival parameter, and a base layer and enhancement layer parameter.

[0025] By implementing the method, the network device can intuitively and accurately obtain the reception of the downlink video frame in multiple dimensions, and facilitate the optimization of the transmission scheme of the downlink video frame.

[0026] Optionally, the extended delay parameter is used to indicate at least one of the following: an extended delay of a first video frame, extended delays of multiple video frames, a maximum extended delay of the multiple video frames, a minimum extended delay of the multiple video frames, and a variance of the extended delays of the multiple video frames; wherein the extended delay is a time length between a first data packet of a video frame successfully received by the terminal device and a last data packet of the video frame successfully received by the terminal device, and the first video frame or the multiple video frames are one or more video frames in the at least one video frame received by the terminal device.

[0027] By implementing the above method, after learning the above extended delay parameter, the network device can adjust the scheduling strategy according to the size of the extended delay. For example, if the extended delay is too large, more resources are scheduled for the terminal device, or the modulation and coding scheme (MCS) is adjusted, thereby improving the user experience; if the extended delay is too small, the scheduling resources can be reduced, or the MCS is adjusted, thereby saving resources, etc.

[0028] Optionally, the frame interval parameter is used to indicate at least one of the following: the frame interval of adjacent video frames, the average of a plurality of frame intervals, the maximum of a plurality of frame intervals, the minimum of a plurality of frame intervals, and the variance of a plurality of frame intervals; wherein the terminal device receives a plurality of video frames from the network device.

[0029] By implementing the above method, if the network device finds that the frame interval of the video frame is unstable, the network device can subsequently increase the size of the buffer to ensure that the video frames are sent at the same interval.

[0030] Optionally, the packet loss parameter includes at least one of the following: the number of video frames with packet loss, the ratio of video frames with packet loss, and K consecutive video frames with packet loss, wherein K is a positive integer greater than or equal to 1.

[0031] By implementing the above method, if the network device finds that there are fewer video frames with packet loss, but multiple packets are lost once packet loss occurs, the network device determines that the wireless signal deep fade time is long, and can replace the new frequency band to send the video frame.

[0032] Optionally, the late parameter indicates at least one of the following: the number of late video frames, the proportion of late video frames, the time difference between the actual reception time and the correct reception time of a late video frame, the average of the time difference between the actual reception time and the correct reception time of a plurality of late video frames, the maximum of the time difference between the actual reception time and the correct reception time of a plurality of late video frames, the minimum of the time difference between the actual reception time and the correct reception time of a plurality of late video frames, and the variance of the time difference between the actual reception time and the correct reception time of a plurality of late video frames.

[0033] In a possible design, each video frame includes basic layer data and enhancement layer data; the basic layer and enhancement layer parameters include at least one of the following: a time difference at which the terminal device receives the basic layer data and the enhancement layer data in a second video frame, the second video frame being a video frame in the at least one video frame; at least one of a number of third video frames, a proportion of third video frames, or a number of enhancement layer packet losses in third video frames, the third video frames being video frames in which the basic layer data has no packet loss and the enhancement layer data has packet loss, and the third video frames being video frames in the at least one video frame; and at least one of a number of fourth video frames, a proportion of fourth video frames, or a number of basic layer packet losses in fourth video frames, the fourth video frames being video frames in which the enhancement layer data has no packet loss and the basic layer data has packet loss, and the fourth video frames being video frames in the at least one video frame.

[0034] By implementing the method, if the network device finds that the time difference between the basic layer and the enhancement layer of the same video frame is too large, the network device can subsequently reduce the transmission interval of the basic layer and the enhancement layer.

[0035] In a possible design, after the video frame is network encoded, the video frame includes one or more data slices, and the video frame parameters further include at least one of the following: a number of data slices and / or a data amount of the data slices that need to be received by the terminal device for successfully decoding one video frame; a number of data slices and / or a data amount of the data slices that need to be received by the terminal device for successfully decoding the basic layer of one video frame; and a number of data slices and / or a data amount of the data slices that need to be received by the terminal device for successfully decoding the enhancement layer of one video frame.

[0036] By implementing the method, the network device can adjust the redundancy rate of the basic layer or the enhancement layer in the network encoding process according to the report of the terminal device, and optimize transmission of the downlink video frame.

[0037] In a third aspect, an apparatus is provided, and the advantages are as described in the first aspect. The apparatus has functions of implementing the behaviors in the method embodiments of the first aspect. The functions can be implemented by executing corresponding hardware or software. The hardware or software can include one or more units corresponding to the functions. In a possible design, the apparatus can include: a communication unit, configured to receive at least one video frame from a network device; a processing unit, configured to determine video frame parameters according to the at least one video frame; and the communication unit is further configured to send the video frame parameters to the network device. These units can perform the corresponding functions in the method embodiments of the first aspect, and details are described in the method embodiments, which are not described here.

[0038] In a fourth aspect, an apparatus is provided. The apparatus can be a terminal device or a chip configured in a terminal device. The apparatus includes a communication interface and a processor. Optionally, the apparatus further includes a memory. The memory is configured to store a computer program or instructions. The processor is coupled to the memory and the communication interface. When the processor executes the computer program or the instructions, the apparatus performs the method described in the first aspect.

[0039] In a fifth aspect, an apparatus is provided. The apparatus can perform the functions of the method described in the second aspect. The functions can be implemented by hardware or software. The hardware or software can include one or more units corresponding to the functions. In one possible design, the apparatus can include a communication unit configured to send at least one video frame to a terminal device. The apparatus can also include a communication unit configured to receive a video frame parameter from the terminal device, the video frame parameter being determined based on the at least one video frame. Optionally, the apparatus can include a processing unit configured to optimize transmission of a downlink video frame based on the video frame parameter. These units can perform the functions of the method examples described in the second aspect. Details are not described here again.

[0040] In a sixth aspect, an apparatus is provided. The apparatus can be a network device or a chip configured in a network device. The apparatus includes a communication interface and a processor. Optionally, the apparatus further includes a memory. The memory is configured to store a computer program or instructions. The processor is coupled to the memory and the communication interface. When the processor executes the computer program or the instructions, the apparatus performs the method described in the second aspect.

[0041] In a seventh aspect, a computer program product is provided. The computer program product includes computer program codes. When the computer program codes are executed, the method described in the first aspect is performed.

[0042] In an eighth aspect, a computer program product is provided. The computer program product includes computer program codes. When the computer program codes are executed, the method described in the second aspect is performed.

[0043] In a ninth aspect, a chip system is provided, which includes a processor configured to implement functionalities of the terminal device in the method of the first aspect. In a possible design of the chip system, the chip system further includes a memory configured to store program instructions and / or data. The chip system can be formed by a chip, or can include a chip and other discrete components.

[0044] In a tenth aspect, a chip system is provided, which includes a processor configured to implement functionalities of the network device in the method of the second aspect. In a possible design of the chip system, the chip system further includes a memory configured to store program instructions and / or data. The chip system can be formed by a chip, or can include a chip and other discrete components.

[0045] In an eleventh aspect, a computer readable storage medium is provided, which stores a computer program. When the computer program is run, the method performed by the terminal device in the first aspect is implemented.

[0046] In a twelfth aspect, a computer readable storage medium is provided, which stores a computer program. When the computer program is run, the method performed by the network device in the second aspect is implemented.

[0047] In a thirteenth aspect, a communication system is provided, which includes the apparatus of the third aspect or the fourth aspect, and the apparatus of the fifth aspect or the sixth aspect. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 A schematic diagram of a network architecture provided by an embodiment of the present application;

[0049] Figure 2 Another schematic diagram of a network architecture provided by an embodiment of the present application;

[0050] Figure 3 A schematic diagram of adjacent video frames provided by an embodiment of the present application;

[0051] Figure 4 A schematic diagram of network coding provided by an embodiment of the present application;

[0052] Figure 5 A flowchart of a communication method provided by an embodiment of the present application;

[0053] Figure 6 Another schematic diagram of adjacent video frames provided by an embodiment of the present application;

[0054] Figure 7 A schematic diagram of a late video frame provided by an embodiment of the present application;

[0055] Figure 8A schematic diagram of a basic layer and an enhanced layer provided for an embodiment of the present application;

[0056] Figure 9 Another schematic diagram of network coding provided for an embodiment of the present application;

[0057] Figure 10 A structural schematic diagram of an apparatus provided for an embodiment of the present application;

[0058] Figure 11 Another structural schematic diagram of an apparatus provided for an embodiment of the present application. DETAILED DESCRIPTION

[0059] Figure 1 A schematic diagram of a network architecture to which an embodiment of the present application is applied is shown, including at least one of the following: a terminal device, an access network device, a core network (CN) device, and a data network (DN). The access network device and the core network device can communicate through a next generation (NG) interface, and different access network devices can communicate through an Xn interface.

[0060] 1. Terminal device

[0061] The terminal device can be referred to as a terminal, which is a device with wireless transceiver function. The terminal device can be mobile or fixed. The terminal device can be deployed on land, including indoor or outdoor, handheld or vehicle-mounted; can also be deployed on the water surface (such as ships, etc.); can also be deployed in the air (such as airplanes, balloons and satellites, etc.). The terminal device can be a mobile phone, a pad, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in self driving, a wireless terminal device in remote medical, a wireless terminal device in smart grid, a wireless terminal device in transportation safety, a wireless terminal device in smart city, and / or a wireless terminal device in smart home. The terminal device can also be a cellular phone, a cordless phone, a session initiation protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a handheld device or a computing device with wireless communication function, a vehicle-mounted device, a wearable device, a terminal device in the 5th generation (5G) network, or a terminal device in an evolved public land mobile network (PLMN), etc. The terminal device can also be referred to as a user equipment (UE), and the terminal device can communicate with multiple access network devices of different technologies, for example, the terminal device can communicate with an access network device supporting long term evolution (LTE), can also communicate with an access network device supporting 5G, and can also communicate with dual connectivity of an access network device supporting LTE and an access network device supporting 5G. The embodiments of the present application are not limited.

[0062] 2. Access network device

[0063] The access network device can also be referred to as a radio access network (RAN) device, which is a device for accessing a terminal device to a wireless network, and can provide the terminal device with functions such as wireless resource management, quality of service management, data encryption and compression. The access network device includes but is not limited to:

[0064] A next generation nodeB (gNB) in 5G, an evolved node B (eNB), a radio network controller (RNC), a node B (NB), a base station controller (BSC), a base transceiver station (BTS), a home base station (for example, a home evolved nodeB, or home node B (HNB)), a base band unit (BBU), a transmitting and receiving point (TRP), a transmitting point (TP), a mobile switching center, and / or the like. Alternatively, the access network device can also be a radio controller in a cloud radio access network (CRAN) scenario, a centralized unit (CU), and / or a distributed unit (DU). Alternatively, the access network device can be a relay station, an access point, a vehicle-mounted device, an access network device in a 5G network, or an access network device in an evolved public land mobile network (PLMN), and the like.

[0065] In some embodiments, as Figure 2As shown, the access network device can include a central unit (CU) and a distributed unit (DU), that is, the functions of the access network device can be split, part of the functions of the access network device are placed in the CU, and the remaining part of the functions are placed in the DU, multiple DUs share one CU, cost is saved, and network expansion is easy. Optionally, the functions of the CU and the DU can be divided according to a protocol stack. For example, the radio resource control (RRC) layer, the service data adaptation protocol (SDAP) layer and the packet data convergence protocol (PDCP) layer are deployed in the CU. The remaining radio link control (RLC) layer, the medium access control (MAC) layer and the physical layer (PHY) layer are deployed in the DU. The CU and the DU can be connected through an FI interface. The CU can be connected with the core network through an NG interface on behalf of the access network device, and the CU can also be connected with other access network devices through an Xn interface on behalf of the access network device. Further, the functions of the CU can also be divided into:

[0066] 1. Central unit-control plane (CU-CP): mainly including the RRC layer in the CU, and the control plane in the PDCP layer;

[0067] 2. Central unit-user plane (CU-UP): mainly including the SDAP layer in the CU, and the user plane in the PDCP layer.

[0068] 3. Core network device

[0069] The core network device is mainly used for managing terminal devices and providing a gateway for communication with an external network. The core network device can include one or more network elements from among: an access and mobility management function (AMF) network element, a session management function (SMF) network element, a user plane function (UPF) network element, a policy control function (PCF) network element, an application function (AF) network element, a unified data management (UDM) network element, an authentication server function (AUSF) network element, and a network slice selection function (NSSF) network element. The AMF network element is mainly responsible for mobility management in a mobile network, such as user location updating, user registration network, and user switching. The SMF network element is mainly responsible for session management in a mobile network, such as session establishment, modification, and release; specific functions such as allocating an IP address for a user, selecting a UPF network element for providing message forwarding functions, and the like. The UPF network element is mainly responsible for forwarding and receiving user data; in downlink transmission, the UPF network element can receive user data from a data network (DN) and transmit the user data to a terminal device through an access network device; in uplink transmission, the UPF network element can receive user data from a terminal device through an access network device and forward the user data to a DN. Optionally, the transmission resources and scheduling functions provided by the UPF network element for terminal devices can be managed and controlled by the SMF network element. The PCF network element is mainly used to support a unified policy framework to control network behavior, provide policy rules to control plane network functions, and be responsible for obtaining user subscription information related to policy decision. The AF network element is mainly used to support interaction with a core network of a wireless network, such as a 3rd generation partnership project (3GPP) network, to provide services, such as affecting data routing decisions, policy control functions, or providing third-party services to the network side. The UDM network element is mainly used to generate authentication credentials, handle user identification (such as storing and managing user permanent identities), access authorization control, and subscription data management, and the like.The AUSF network element is mainly used for performing authentication when a terminal device accesses a network, including receiving an authentication request sent by a security anchor function (SEAF), selecting an authentication method, and requesting an authentication vector from an authentication repository and processing function (ARPF), etc. The NSSF network element is mainly used for selecting a network slice instance for a terminal device, determining allowed network slice selection assistance information (NSSAI), configuring NSSAI, and determining an AMF set serving the terminal device. In different communication systems, the network elements or network element names in the core network can be different. In the above. Figure 1 The schematic diagram shown is described by taking a fifth-generation mobile communication system as an example, and does not limit the present application.

[0070] 4, DN

[0071] The DN can be a service network providing data service for a user. For example, the DN can be an IP multimedia service (IP multimedia service) network or an Internet, etc. Among them, the terminal device can establish a protocol data unit (PDU) session from the terminal device to the DN to access the DN, etc.

[0072] In some embodiments, an extend reality (XR) service provides an immersive multimedia experience for a user by interacting with an XR application server through an application layer user device (such as a helmet). Optionally, the XR application server can be located on the DN side in the above Figure 1 The helmet can be located on the terminal device side in the above Figure 1 The helmet and the terminal device can be integrated or separated in physical form, without limitation.

[0073] In the uplink transmission process, the sensors in the helmet can perceive the position and action changes of the user, and generate user information including the perspective, line of sight, and motion rate, etc. The helmet can transmit the user information to the XR server through the terminal device, access network device, UPF network element, etc. in the above Figure 1

[0074] In the downlink transmission process, the XR server can generate a video frame in combination with the user information reported by the helmet and the scene of the user in the game or the real scene. The XR server can transmit the video frame to the helmet through the terminal device, access network device, UPF network element, etc. in the above Figure 1 ​The UPF network element, access network device, terminal device and the like in the illustrated architecture transmit video frames to the helmet, and finally show the above-mentioned video frames to the user through the helmet. In the downlink transmission process, how does the terminal device feed back the reception condition of the video frame to the network device when receiving the downlink video frame, so as to optimize the transmission of the downlink video frame, which is a technical problem to be solved by the embodiments of the present application.

[0075] The present application provides a kind of communication method and device, the method comprises: terminal equipment receives at least one video frame from access network device;Terminal equipment determines video frame parameter according to the at least one video frame;Terminal equipment sends the video frame parameter to access network device, so that access network device can know the reception condition of downlink video frame, facilitate optimizing the transmission of downlink video frame.

[0076] For the convenience of understanding, the communication terms or terms related to the embodiments of the present application are explained and described.

[0077] 1, the service model of XR. Optionally, XR service mainly includes:

[0078] Upstream service: the user information generated by the helmet, including user position and line of sight, etc., the data volume is small;

[0079] Downlink service: the video frame generated by the XR server according to the user information reported by the helmet, the data volume is large.

[0080] Among them, such as Figure 3As shown, for downlink service, the XR server generates video frames at a rate of 60 frames per second, i.e. one raw video frame is generated every 16.67 ms. After encoding and compression, the raw video frames form different types of video frames. For example, a typical video frame can include I-frame, P-frame, B-frame, etc. The size of each video frame can vary from 1000 kilobits (kbit) to 10000 kilobits (kbit). Optionally, for video service, the basic processing is to divide the video frame into N pictures per second, and each picture is encoded as a video frame. Since each video frame contains a large number of pixels, direct transmission will occupy a large bandwidth. Therefore, the video service can be compressed before transmission. Due to the nature of video service, as long as the lens does not switch, the content of most pictures between adjacent frames is usually the same, and only a small amount of content is different. Therefore, the video frames can be grouped, and the first frame of each group is a reference frame, and the subsequent frames are dependent frames. When compressing, the reference frame is intra-frame compressed, i.e. only the code stream of the frame itself is referred to during compression, without referring to other frames. In this way, when the decompression side receives a reference frame, it does not need other frames and can be independently decompressed; the dependent frames after the reference frame are inter-frame compressed, i.e. the code stream of the frame itself is referred to during compression, and other frames are also referred to, such as the reference frame. In this way, the compression rate can be greatly improved during compression, and the size of the compressed data can be reduced. Among them, the reference frame can also be referred to as an I-frame, which has the largest size. The dependent frames include P-frames and B-frames, which have smaller sizes. The P-frame refers to a frame that is only dependent on the previous frame during decoding, and the B-frame refers to a frame that is not only dependent on the previous frame but also dependent on the subsequent frame during decoding.

[0081] In some embodiments, due to network limitations, the encoded data stream is divided into multiple data packets with 1500 bytes (the actual data payload can be slightly less than 1500 bytes considering the overhead of the packet header) as the standard. From the perspective of wireless network, the downlink video data is embodied as a cluster of downlink data packets received by the UPF from the XR server every 16.67 ms. The cluster of downlink data packets can include multiple data packets with a size of about 1500 bytes. The cluster of downlink data packets is one downlink video frame.

[0082] 2. Video layered encoding

[0083] The size of the original video frame is too large to put a great pressure on the transmission network. The video frame is compressed before transmission. The current compression method can achieve a compression efficiency of 300: 1, i.e. a 300-megabit (Mbit) file can be compressed to about 1 megabit (Mbit). However, as the resolution of the video frame becomes higher and the chroma division becomes finer, the size of the video frame also becomes larger. According to the compression efficiency of 300: 1, the amount of compressed data is still too large to put a great pressure on the transmission network. Especially for wireless networks, as the capacity of the wireless network fluctuates greatly, when the channel quality is poor, the amount of data transmitted by the channel also decreases, and the high-definition video frame cannot be transmitted, and the user experiences mosaic or even image freezing.

[0084] Based on the above situation, video layered coding is introduced, i.e. in the compression coding process, the information of the video frame is divided into two categories, basic layer data and enhancement layer data. In some embodiments, the ratio of the data amount of the two types of data can be 1: 9. When the channel capacity is small, the user can only receive the basic layer data, and can obtain a low-resolution image through the display, but there is no mosaic. If the channel capacity is large, the user receives the basic layer data and the enhancement layer data at the same time, and can obtain a high-resolution image. Compared with the non-layered high-definition video, the sum of the data amount of the basic layer data and the enhancement layer data is greater than the data amount of the non-layered video. The layered coding method can better cope with the rapid change of the wireless channel, reduce the situation of freezing and mosaic, and improve the user experience.

[0085] In actual layered coding, the enhancement layer can include one layer or be divided into multiple layers, which is not limited. The more the number of layers is, the more suitable it is for wireless channels with different capacities.

[0086] 3. Network coding

[0087] Due to the unstable situation of the wireless channel, transmission errors may occur, and the retransmission method is usually used to compensate. The retransmission can solve the problem of packet error, but it will introduce additional delay and low efficiency. For real-time multimedia services, the amount of data to be transmitted is large, and almost every time slot has a data packet to be transmitted. If an additional retransmission data packet is inserted, the data of each subsequent time slot is delayed, thereby increasing the time delay. In addition, the data packet of the real-time multimedia service is usually large, and the error may be only a small part of it. If the entire data packet is retransmitted because of the transmission error of a small part of the data packet, wireless resources will be wasted, and efficiency will be reduced.

[0088] Based on the above situation, the network coding mechanism is introduced, and the core idea is to uniformly encode a group of data packets. In an example, as shown in FIG. 2, a group of data packets is uniformly encoded to obtain a new data packet. The new data packet is transmitted to the receiver, and the receiver can obtain the original data packet through decoding. The network coding mechanism can solve the problem of packet error, and the efficiency is also improved. Figure 4As shown, 10 data packets included in a video frame can be divided into a group, unified network coding is performed, a plurality of small data blocks are output after coding, and the small data blocks are transmitted in batches according to air interface conditions. After receiving all the small data blocks, the receiver decodes the entire data packet. If it is found that a certain data packet cannot be correctly decoded, the sender can be notified to send some small data blocks again. The small data blocks sent by the sender again are not the data blocks that have been transmitted before, but the small data blocks newly generated by the network coding module, and the receiver can uniformly decode all the received small data blocks.

[0089] In some embodiments, as shown, Figure 4 For each data packet, a PDCP PDU is generated after processing by the PDCP layer and transmitted to the RLC layer, and an RLC PDU is generated after processing by the RLC. The RLC PDU can be network coded at the RLC layer to obtain a plurality of small data blocks, which are transmitted to the MAC layer and processed by the MAC layer to obtain MAC PDUs. Each MAC PDU can include fields such as logical channel identity document (LCID), length (Len), and data (data). Alternatively, the small data blocks formed after network coding can also be referred to as data slices. In the following embodiments, data slices are described as an example.

[0090] By using the above network coding method, there is no need to design a real-time feedback mechanism, and the sender does not need to retransmit the entire original data packet, thereby improving the efficiency of wireless transmission and reducing the transmission delay.

[0091] In the description of the present application, unless otherwise specified, " / " represents that the objects before and after the correlation are in an "or" relationship, for example, A / B can represent A or B; "and / or" in the present application is only a description of the correlation of the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. And, in the description of the present application, unless otherwise specified, "multiple" means two or more than two. "At least one of the following" or the like means any combination of the items, including any combination of single item or multiple items. For example, at least one of a, b, or c can represent: a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, "first", "second", etc. are used to distinguish the same items or similar items with basically the same function and role. The skilled in the art can understand that "first", "second", etc. do not limit the quantity and execution order, and "first", "second", etc. also do not necessarily mean different.

[0092] In addition, the network architecture and service scenarios described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of network architecture and the appearance of new service scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0093] The network elements involved in the embodiments of the present application include access network devices, terminal devices, helmets, and XR servers, etc. Among them, the access network device can be of CU / DU architecture, at this time the access network device includes two network elements of CU and DU. Or, the access network device can also be of CP-UP architecture, at this time the access network device contains three network elements of CU-CP, CU-UP and DU. Or, the access network device can also be of open radio access network (ORAN) architecture, at this time the access network device includes four network elements of CU-CP, CU-UP, DU, and near real-time access network intelligent controller (RIC) or even more network elements. Further, the DU can also be separated into DU-H and DU-L, etc. to support low-layer separation such as remote physical layer. The network device (or base station) in the following embodiments can be one network element in the access network device, or can include multiple network elements in the access network device.

[0094] For ease of understanding, the terminal device is taken as UE, and the network device is taken as a base station for description. As shown in Figure 5 A flow of a communication method is provided, including:

[0095] Step 501: The UE receives at least one video frame from the base station. Optionally, as described above, one or more data packets can be included in each video frame, and the size of each data packet is about 1500 bytes. The data packet can also be referred to as an internet protocol (IP) packet. Hereinafter, the data packet is taken as an example for description. Optionally, the data packets included in each video frame can be video data, or video data and audio data, etc., without limitation.

[0096] Step 502: The UE determines a video frame parameter according to the at least one video frame.

[0097] Optionally, the UE can determine the video frame parameter in the granularity of a video frame.

[0098] Step 503: The UE sends the video frame parameter to the base station. Optionally, the UE can report the video frame parameter in the following ways:

[0099] Periodic reporting. The UE reports the video frame parameter once per period, and the reporting period can be configured by the base station, an SMF network element, or an XR server. When the reporting period is configured by the base station, the UE receives a configuration information element from the base station, the configuration information element indicates the reporting period, and the UE reports the video frame parameter according to the reporting period indicated by the configuration information element. When the reporting period is configured by the SMF network element or the XR server, the SMF network element or the XR server sends information of the reporting period to the base station, the base station generates a configuration information element indicating the reporting period according to the information, and sends the configuration information element to the UE, or the base station directly sends the information of the reporting period to the UE as the configuration information element, and the UE reports the video frame parameter according to the reporting period indicated by the configuration information element.

[0100] Condition-triggered reporting. The UE reports the video frame parameter once if it is found that the reporting condition is met, and the reporting condition can be configured by the base station, the SMF network element, or the XR server.

[0101] Periodic + condition-triggered reporting. The UE statistics the video frame parameter at the end of each period, and judges whether the pre-configured condition is met. If the condition is met, the UE reports the video frame parameter, otherwise, the UE does not report the video frame parameter. Similarly, the reporting period and the reporting condition can be configured by the base station, the SMF network element, or the XR server.

[0102] The configuration manner of the reporting condition is the same as the reporting period, and will not be described again. The object of the UE reporting the video frame parameter, i.e., the receiving object of the video frame parameter, can be a base station, a core network element such as an SMF or a UPF, or even an XR server. Of course, if the receiving object of the video frame parameter is a core network element or an XR server, the base station forwards the received video frame parameter to the core network element or the XR server. The above is described by taking one receiving object of the video frame parameter as an example. In fact, the receiving object of the video frame parameter can also be multiple. For example, the receiving object of the video frame parameter can include a base station, a core network element, and an XR server. After receiving the video frame parameter, each receiving object can perform corresponding optimization. For example, after receiving the video frame parameter, the base station can optimize the scheduling of radio resources. After receiving the video frame parameter, the SMF network element can optimize the service configuration of the user. After receiving the video frame parameter, the UPF network element can prioritize the forwarding priority of the data packet. After receiving the video frame parameter, the XR server can optimize the encoding and / or compression mode of the downlink video frame.

[0103] As can be seen from the above, in the embodiment of the present application, the UE obtains the video frame parameter according to the received video frame, and reports the receiving condition of the video frame, so that the network device can intuitively and accurately obtain the receiving condition of the video frame, and facilitate optimization of the video frame transmission scheme.

[0104] In one scheme, for a video service, the UE reports the receiving condition of each data packet, such as packet loss and / or packet delay, with data packet as the granularity. This reporting manner is not accurate enough for the statistical condition of the video service, for the following reasons. For example, the UE loses two data packets, and there are two cases:

[0105] Case 1: The two lost data packets are located in different video frames.

[0106] Case 2: The two lost data packets are located in the same video frame.

[0107] From the perspective of network performance, the above two cases correspond to the same network performance, and the number of lost data packets is the same, but from the perspective of user experience, the above two cases correspond to different user experiences. For the above case 1, the user can continuously see two pictures with poor quality. For the above case 2, the user sees one picture with poor quality, and the next picture is very clear. The UE reports the receiving condition of the data packet with data packet as the granularity, which is difficult to reflect the above two differences. In the embodiment of the present application, the UE reports the receiving condition of the video frame with video frame as the granularity, which can more accurately reflect the receiving condition of the video frame, and facilitate optimization of the entire video transmission scheme.

[0108] In an embodiment of the present application, the video frame parameter reported by the UE can include at least one of the following:

[0109] 1. A spread delay parameter, which can indicate at least one of the following: the spread delay of a first video frame, the average spread delay of multiple video frames, the maximum spread delay of multiple video frames, the minimum spread delay of multiple video frames, and the spread delay variance of multiple video frames. The first video frame or the multiple video frames are one or more of the at least one video frame received by the UE.

[0110] One picture can be encoded, compressed, etc., to form one or more video frames, and each video frame can be divided into multiple data packets. At this time, from the perspective of the UE, one video frame is received, i.e., multiple data packets are received. The spread delay can refer to the length of time between the successful reception of the first data packet of a video frame by the UE and the successful reception of the last data packet of the video frame. For example, as shown in Figure 6 , the video frame N is divided into three data packets, and the spread delay of the video frame N can specifically be the length of time between the successful reception of the first data packet of the video frame N and the successful reception of the third data packet of the video frame N.

[0111] In an embodiment of the present application, the UE can count the spread delay of each video frame, and obtain the average, maximum, minimum, or variance, etc., of the spread delays of multiple video frames, and report them. Alternatively, in an embodiment of the present application, the UE can also report the spread delay of a single video frame, i.e., the spread delay of the first video frame, without limitation.

[0112] For example, the UE can report the spread delay parameter in any of the following ways:

[0113] - Periodic reporting, configured by the base station or the SMF network element or the XR server to report once per period.

[0114] - Condition-triggered reporting, configured by the base station or the SMF network element or the XR server to report once if the UE finds that the reporting condition is met.

[0115] - Periodic + condition-triggered reporting, at the end of each period, the UE counts the spread delay of each video frame in the period and judges whether the pre-configured reporting condition is met; if the reporting condition is met, the UE reports; otherwise, it does not report. Similarly, the period and the reporting condition can be configured by the base station or the SMF network element or the XR server, etc.

[0116] The reporting condition can be set as needed and is not limited. The reporting condition can correspond to the reported content, for example, the reporting condition can include that when the extension delay of a single video frame is greater than or equal to threshold A1, the extension delay parameter of the single video frame is reported; or the reporting condition can include that when the extension delay of a single video frame is less than or equal to threshold B1, the extension delay parameter of the single video frame is reported. For another example, the reporting condition can include that when the average extension (i.e., the average value of the extension delays of multiple video frames) delay of multiple video frames is greater than or equal to threshold A2, the average extension delay of the multiple video frames is reported; or the reporting condition can include that when the average extension delay of multiple video frames is less than or equal to threshold B2, the average extension delay of the multiple video frames is reported. The reporting of the maximum value, the minimum value, or the variance of the extension delays of multiple video frames is similar. Alternatively, the reporting condition can not correspond to the reported content, for example, the reporting condition can include that when the extension delay of a single video frame is greater than or equal to threshold A3, the average extension delay of multiple video frames including the single video frame is reported; or the reporting condition can include that when the extension delay of a single video frame is greater than or equal to threshold B3, the average extension delay of multiple video frames including the single video frame is reported; wherein the single video frame can be the first video frame, the middle video frame, or the last video frame of the multiple video frames, and the number of the multiple video frames can be pre-set. The reporting of the maximum value, the minimum value, or the variance of the extension delays of multiple video frames is similar. The values of the thresholds A1-A3 and B1-B3 are not limited and can be the same or different.

[0117] After the network device learns the extension delay parameter, the scheduling strategy can be adjusted according to the size of the extension delay, for example, if the extension delay is too large, more resources are scheduled for the UE, or the modulation and coding scheme (MCS) is adjusted, thereby improving the user experience; if the extension delay is too small, the scheduling resources can be reduced, or the MCS is adjusted, thereby saving resources.

[0118] 2, frame spacing parameter, the frame spacing parameter is used to indicate at least one of the following: frame spacing of adjacent video frames, average value of multiple frame spacings, maximum value of multiple frame spacings, minimum value of multiple frame spacings, and variance of multiple frame spacings. Wherein, the UE receives multiple video frames from the base station.

[0119] Optionally, the frame interval can also be referred to as a video frame gap. The video encoder in the XR application server sends out each video frame, and the time difference between adjacent video frames is the same, and the specific value depends on the frame rate. For example, in the case of 60 frames per second, the time difference between two adjacent video frames is 16.67 ms. However, due to the time taken by the network to transmit each video frame is not exactly the same, especially for wireless networks. Therefore, the UE does not receive a video frame every 16.67 ms when receiving the video frame. The UE can determine the frame interval of adjacent video frames according to the reference position of the two adjacent video frames, which can be the starting position, the middle position, the end position or any other position, etc. For example, the UE can count the time of successful reception of the first data packet corresponding to the two adjacent video frames respectively, and calculate the time difference between the two reception times as the frame interval of the adjacent video frames. For example, as shown in FIG. 6, video frame N and video frame N+1 are adjacent video frames, each including 3 data packets. The UE can count the time of successful reception of the first data packet in video frame N and the time of successful reception of the first data packet in video frame N+1 respectively; the time difference between the two times is the frame interval of video frame N and video frame N+1; or the UE can also count the time of successful reception of the last data packet corresponding to the two adjacent video frames respectively, and calculate the time difference between the two reception times as the frame interval of the adjacent video frames. Figure 6

[0120] 3. A packet loss parameter, the packet loss parameter comprising at least one of: a number of video frames in which packet loss occurs, a proportion of video frames in which packet loss occurs, and K consecutive video frames in which packet loss occurs, K being a positive integer greater than or equal to 1.

[0121] ​In the above example, the UE determines that a video frame has packet loss if the UE receives multiple data packets of the video frame and finds that there is packet loss. Specifically, the UE can determine whether there is packet loss by using PDCP sequence numbers (SNs). If the PDCP SNs of the received data packets are not continuous, the UE determines that there is packet loss. For example, a video frame includes M data packets, and the PDCP SNs of the M data packets are 1 to M in sequence. When the UE receives the Mth data packet, the UE finds that the data packet with PDCP SN N has not been received, and determines that the PDCP SNs of the video frame are not continuous, and thus determines that the video frame has packet loss. Alternatively, whether the PDCP SNs are continuous can be replaced by whether the PDCP SNs are missing, and the like. Specifically, the UE can determine whether there is packet loss at the time when the UE receives an end marker of a video frame, at the time when data packets of the video frame are submitted to an IP layer or a real-time transport protocol (RTP) layer, at the time when the data packets of the video frame are played in an application layer, and the like.

[0122] In some embodiments, the UE can determine, in a period of time, a number of video frames that have packet loss, a proportion of the video frames that have packet loss in the total video frames in the period of time, and the like. For example, in a period of time, the UE receives 5 video frames from the base station. Among the 5 video frames, 3 video frames have packet loss. Thus, the number of the video frames that have packet loss is 3, and the proportion of the video frames that have packet loss is 3 / 5 = 60%. The UE can further determine, among the video frames that have packet loss, a proportion of video frames that lose 1 data packet, a proportion of video frames that lose 2 data packets, a proportion of video frames that lose 3 data packets, and the like. In the above example, in a period of time, 3 video frames have packet loss, 1 video frame loses 1 data packet, and 2 video frames lose 2 data packets. Thus, the proportion of the video frames that lose 1 data packet is 1 / 3 = 33.3%, and the proportion of the video frames that lose 2 data packets is 2 / 3 = 66.7%. The UE can further determine how many video frames have consecutive packet loss. For example, in a period of time, the UE receives 10 video frames from the base station, and the sequence numbers of the 10 video frames are 1 to 10 in sequence. Among the 10 video frames, the 1st video frame has packet loss, the 2nd to 5th video frames have no packet loss, the 6th to 8th video frames have packet loss, and the 9th and 10th video frames have no packet loss. Thus, the UE can report that the 6th to 8th video frames have consecutive packet loss.

[0123] 4. a late parameter, the late parameter is used to indicate at least one of the following: the number of late video frames, the proportion of late video frames, the time difference between the actual receiving time and the correct receiving time of a late video frame, the average of the time difference between the actual receiving time and the correct receiving time of multiple late video frames, the maximum of the time difference between the actual receiving time and the correct receiving time of multiple late video frames, the minimum of the time difference between the actual receiving time and the correct receiving time of multiple late video frames, the variance of the time difference between the actual receiving time and the correct receiving time of multiple late video frames, etc.

[0124] Since the correlation between consecutive video frames is relatively high, the encoding party in the XR application server will select a certain video frame in front of it as a reference frame for compression in the process of compressing the video frame. For example, if the video frame is a late video frame, the encoding party will select a certain video frame in front of it as a reference frame for compression in the process of compressing the video frame. Figure 7Therefore, if part of the data packets in a certain video frame are not successfully transmitted within the predefined time delay, the successful transmission is after the time delay budget. Although the video frame cannot be displayed by the display, it can still be useful as a reference frame for subsequent video frames. The UE can count the number of such video frames, the proportion of all video frames, how much time the correct receiving time is later than the time delay budget, and report to the base station. In the embodiments of the present application, the specific parameters reported by the UE can be specifically: the number of late video frames, the proportion of late video frames, the time difference between the actual receiving time and the correct receiving time of a late video frame, the maximum, minimum, or variance of the time difference between the actual receiving time and the correct receiving time of multiple late video frames, etc., without limitation. For example, in a period of time, the UE receives 10 video frames from the base station, video frames 1 to 5 arrive correctly, and video frames 6 to 10 are late. The number of late video frames is 5, and the proportion of late video frames is 5 / 10 = 50%. The actual arrival time of video frame 6 is 2020-08-28 15:05:35, and the correct arrival time of video frame 6 should be 2020-08-28 15:05:20, so the time difference between the actual arrival time and the correct arrival time of video frame 6 is 15 seconds. In a possible implementation manner, the following manner can be used to define the actual receiving time and the correct receiving time of a video frame: the actual receiving time of a video frame can refer to the time when the UE receives all data packets corresponding to the video frame. Assuming that the Nth video frame includes 10 data packets, the actual receiving time of the Nth video frame can be the time when the UE receives all 10 data packets of the Nth video frame. First, define the correctly received video frame, if all data packets of a video frame are received before the predefined time, the video frame can be considered as a correctly received video frame. Assuming that the ith video frame is correctly received, the correct receiving time of the (i+M)th video frame is: the actual receiving time of the last data packet of the ith video frame + M*the time difference between adjacent video frames. Optionally, the time difference between adjacent video frames can be 16.67 ms, and M is a positive integer greater than or equal to 1.

[0125] In the above description, the late video frame is defined as "received after the time delay budget, but still can be used as a reference frame for other video frames behind it" as an example, which is not limited to the embodiments of the present application. For example, in the embodiments of the present application, the late video frame can also be defined as "a video frame received after the first time delay budget", or "a video frame received after the first time delay budget and before the second time delay budget", or "a video frame whose late time meets the preset time requirement", etc.

[0126] 5. Base layer and enhancement layer parameters

[0127] If video is layered, for example, into base layer and enhancement layer. UE can receive two streams flow or two data radio bearers (DRB) or two data channels, respectively, for base layer data and enhancement layer data. UE can count the data reception of the two layers respectively, and report one or more of the following parameters:

[0128] 1. The time difference of receiving base layer data and enhancement layer data in the same video frame. Optionally, UE can count the time difference of reference positions in the base layer data and enhancement layer data in the same video frame. The reference position can be the start of the frame, the end of the frame, the middle position, or any position, etc. For example, as shown in the figure, for video frame N, the base layer includes 3 data packets, and the enhancement layer includes 4 data packets. The time difference of the start of the frame of the base layer and the enhancement layer of video frame N can be counted, or the time difference of the end of the frame of the base layer and the enhancement layer of video frame N can be counted, etc. It can be understood that UE can report the time difference between the base layer data and the enhancement layer data in a single video frame, and UE can also report the average, maximum, minimum, or variance of the time difference between the base layer data and the enhancement layer data in multiple video frames. Figure 8

[0129] In the above description, the video frame is divided into two layers as an example for description. In fact, the video frame can also be divided into three layers, four layers, or more layers. For example, for three layers, the entire video frame is divided into a base layer, a first enhancement layer, and a second enhancement layer. For four layers, the entire video frame is divided into a base layer, a first enhancement layer, a second enhancement layer, and a third enhancement layer. Taking the above video frame divided into three layers as an example, UE can report the time difference between the first enhancement layer data and the base layer data, the time difference between the second enhancement layer data and the base layer data, etc., or even the time difference between different enhancement layers, etc. without limitation.

[0130] 2. The number, proportion, or number of lost packets of video frames in which the base layer has no lost packets but the enhancement layer has lost packets, or multiple of the number, proportion, and number of lost packets of video frames in which the base layer has no lost packets but the enhancement layer has lost packets. Since the network adopts more robust transmission strategy for the base layer data, the probability of packet loss of the enhancement layer data is greater. UE can count one or more of the number, proportion, and number of lost packets of video frames in which the base layer has no lost packets but the enhancement layer has lost packets, and report the base station. The base station can adjust the transmission strategy of the enhancement layer according to the information.

[0131] ​The above is described by taking the example of the video frame being divided into two layers of the base layer and the enhancement layer. In practice, the video frame can be divided into three layers, four layers or more layers. Taking the example of the video frame being divided into three layers, the UE can count the number, proportion or number of lost packets of the video frame in which the base layer has no lost packet but the first enhancement layer has lost packet, the number, proportion or number of lost packets of the video frame in which the base layer has no lost packet but the second enhancement layer has lost packet, or the number, proportion or number of lost packets of the video frame in which the base layer and the first enhancement layer have no lost packet but the second enhancement layer has lost packet, or the number, proportion or number of lost packets of the video frame in which the base layer has no lost packet but the first enhancement layer and the second enhancement layer have lost packet.

[0132] 3. In the same video frame, the number, proportion or number of lost packets of the video frame in which the enhancement layer has no lost packet but the base layer has lost packet, or a plurality of the number, proportion and number of lost packets of the video frame in which the base layer has lost packet. Since the network adopts a more robust transmission strategy for the base layer data, although the probability of this situation is low, it does not mean that there is no. The UE can report this situation to the base station, and the base station can analyze the reason for the lost packet of the base layer and optimize the network deployment accordingly.

[0133] Similarly, the above is described by taking the example of the video frame being divided into two layers of the base layer and the enhancement layer. In practice, the video frame can be divided into three layers, four layers or more layers. Taking the example of the video frame being divided into three layers, the UE can count the number, proportion or number of lost packets of the video frame in which the first enhancement layer has no lost packet but the base layer has lost packet. Or the UE can count the number, proportion or number of lost packets of the video frame in which the second enhancement layer has no lost packet but the base layer has lost packet. Or the UE can count the number, proportion or number of lost packets of the video frame in which the first enhancement layer and the second enhancement layer have no lost packet but the base layer has lost packet.

[0134] Optionally, the UE specifically reports the above frame spacing parameters, packet loss parameters, late parameters or base layer and enhancement layer parameters in a manner similar to the UE reporting the extended delay parameter, which will not be described one by one.

[0135] As can be seen from the above, the UE reports various video frame parameters in the granularity of the video frame, so that the network can more intuitively understand the transmission of the video frame and optimize the network algorithm.

[0136] The above describes the reporting conditions by taking the example of the extended delay. The reporting conditions of other parameters can refer to the extended delay.

[0137] In Embodiment Two, network coding can be used due to the particularity of the video service. Specifically, the base station can perform network coding on the data of the video service and then transmit the data to the UE. The base station can generally perform network coding on a plurality of data packets included in one video frame. In this case, the video frame parameters reported by the UE can include at least one of the following:

[0138] 1. An extended delay parameter

[0139] The extended delay parameter in Embodiment One is different from that in the present application only in that the video frame in Embodiment One is not subjected to network coding, while the video frame in the present application is subjected to network coding. In some embodiments, after receiving the extended delay parameter, if the base station finds that the extended delay is too large and exceeds the decoding time of a general video decoder, the base station can subsequently select an appropriate time to send the first data packet of a subsequent video frame to ensure that the last data packet of the video frame is sent within a suitable extended delay.

[0140] 2. A frame interval parameter

[0141] The frame interval parameter in Embodiment One is different from that in the present application only in that the video frame in Embodiment One is not subjected to network coding, while the video frame in the present application is subjected to network coding. In some embodiments, if the base station finds that the frame interval of the video frame is unstable, the base station can subsequently increase the size of the buffer to try to ensure that the video frames are sent at the same interval.

[0142] 3. A packet loss parameter

[0143] The packet loss parameter in Embodiment One is different from that in the present application only in that the video frame in Embodiment One is not subjected to network coding, while the video frame in the present application is subjected to network coding. In some embodiments, if the base station finds that the video frame has a small number of lost packets but a large number of packets are lost once a packet is lost, the base station determines that the wireless signal has a long deep fade time and can change to a new frequency band to send the video frame.

[0144] 4. A base layer and enhancement layer parameter

[0145] The base layer and enhancement layer parameter in Embodiment One is different from that in the present application only in that the video frame in Embodiment One is not subjected to network coding, while the video frame in the present application is subjected to network coding. In some embodiments, if the base station finds that the time difference between the base layer and the enhancement layer of the same video frame is too large, the base station can subsequently reduce the transmission interval of the two.

[0146] 5. A data slice parameter; the video frame subjected to network coding includes one or more data slices, and the data slice parameter includes a parameter for at least one of the following:

[0147] - the number and / or amount of data slices that need to be received by the UE for successful decoding of one video frame;

[0148] During video frame encoding, the original data becomes redundant, increasing the data volume. Higher redundancy allows for greater tolerance to packet loss during network transmission, but also increases the demand on network capacity. For example, a video frame's data packet might be 1000 kilobits (kbits) before encoding and 1500 kbits after. However, if the UE receives 1200 kbits, even with some packet loss, it can still perform network decoding and successfully recover the entire video frame data. The UE can then report the amount of data fragments needed to successfully decode a video frame, such as "1200 kbits are needed for successful network decoding." Similarly, the UE can report the number of data fragments needed to successfully decode a video frame, such as "N data fragments are needed for successful network decoding." Upon receiving this information, the base station can adjust the redundancy rate of network encoding for subsequent video frames. Meanwhile, to facilitate base station control of network coding redundancy, the UE can also statistically report "the amount of data it needs to receive and successfully decode a video frame". It is understandable that the UE can report the number or amount of data fragments required to successfully decode a single video frame, or the average, maximum, minimum, or variance of the number or amount of data fragments required to successfully decode multiple video frames.

[0149] - The number of data fragments and / or the amount of data that the UE needs to receive to successfully decode a video frame at the base layer.

[0150] - The number of data fragments and / or the amount of data that the enhancement layer needs to receive to successfully decode a video frame for the UE.

[0151] like Figure 9 As shown, when using layered video coding + network coding, the base station can perform network coding separately for the base layer and the enhancement layer. After receiving base layer data, the UE can count the number or amount of data fragments required to successfully decode the base layer data and report it. After receiving enhancement layer data, the UE can count the number or amount of data fragments required to successfully decode the enhancement layer data, etc. Of course, the aforementioned number or amount of data fragments can be the number or amount of data fragments required for the base or enhancement layer of a single video frame, or it can be the average, maximum, minimum, or variance of the number or amount of data fragments required for the base or enhancement layer across multiple video frames, etc., without limitation. Optionally, the base station can adjust the redundancy rate of the base or enhancement layer during network coding based on the UE's reporting information.

[0152] The above combination Figures 1 to 9 The methods provided in the embodiments of this application are described in detail below. Figure 10 andFigure 11 The apparatus provided by the embodiments of the present application is described in detail. It should be understood that the description of the apparatus corresponds to the description of the method embodiments, and the content not described in detail in the apparatus can be referred to the description in the method embodiments.

[0153] Figure 10 is a schematic block diagram of the apparatus 1000 provided by the embodiments of the present application, which is used to implement the functions of the terminal device or the network device in the above method. For example, the apparatus 1000 can be a software unit or a chip system. The chip system can be composed of a chip, or can include a chip and other discrete devices. The apparatus includes a communication unit 1001, and can further include a processing unit 1002. The communication unit 1001 can communicate with the outside. The processing unit 1002 is used for processing. The communication unit 1001 can also be referred to as a communication interface, a transceiver unit, an input / output interface, etc.

[0154] In an example, the apparatus 1000 can implement the steps performed by the terminal device in the above embodiments. The apparatus 1000 can be a terminal device, or a chip or circuit configured in the terminal device. The communication unit 1001 performs the transceiving operations in the above embodiments, and the processing unit 1002 is used to perform the processing-related operations of the terminal device in the above method embodiments.

[0155] For example, the communication unit 1001 is configured to receive at least one video frame from a network device; the processing unit 1002 is configured to determine a video frame parameter according to the at least one video frame; and the communication unit 1001 is further configured to send the video frame parameter to the network device.

[0156] Optionally, the video frame parameter includes at least one of the following: an extended delay parameter, an inter-frame distance parameter, a packet loss parameter, a late arrival parameter, and a base layer and enhancement layer parameter.

[0157] Optionally, the extended delay parameter is used to indicate at least one of the following: an extended delay of a first video frame, an average extended delay of a plurality of video frames, a maximum extended delay of the plurality of video frames, a minimum extended delay of the plurality of video frames, and a variance of the extended delays of the plurality of video frames; wherein the extended delay is a time length between a successful reception of a first data packet of a video frame by the terminal device and a successful reception of a last data packet of the video frame, and the first video frame or the plurality of video frames are one or more video frames in the at least one video frame received by the terminal device.

[0158] Optionally, the inter-frame distance parameter is used to indicate at least one of the following: an inter-frame distance of adjacent video frames, an average value of a plurality of inter-frame distances, a maximum value of the plurality of inter-frame distances, a minimum value of the plurality of inter-frame distances, and a variance of the plurality of inter-frame distances; wherein the terminal device receives a plurality of video frames from the network device.

[0159] Optionally, the packet loss parameter comprises at least one of: a number of video frames with packet loss, a proportion of video frames with packet loss, K consecutive video frames with packet loss, where K is a positive integer greater than or equal to 1.

[0160] Optionally, the processing unit 1002 is further configured to determine whether the packet data convergence protocol (PDCP) sequence numbers (SNs) of the data packets included in the received video frame are continuous, and determine that the video frame has packet loss when the PDCP SNs of the data packets included in the received video frame are not continuous.

[0161] Optionally, the late parameter indicates at least one of: a number of late video frames, a proportion of late video frames, a time difference between an actual receiving time and a correct receiving time of one late video frame, an average of time differences between actual receiving times and correct receiving times of a plurality of late video frames, a maximum of time differences between actual receiving times and correct receiving times of a plurality of late video frames, a minimum of time differences between actual receiving times and correct receiving times of a plurality of late video frames, and a variance of time differences between actual receiving times and correct receiving times of a plurality of late video frames.

[0162] Optionally, each video frame comprises base layer data and enhancement layer data, and the base layer and enhancement layer parameter comprises at least one of: a time difference between a time at which the terminal device receives the base layer data and a time at which the terminal device receives the enhancement layer data in a second video frame, the second video frame being a video frame in the at least one video frame; a number of third video frames, a proportion of third video frames, or a number of enhancement layer packet losses in the third video frames, the third video frame being a video frame in the at least one video frame in which the base layer data has no packet loss and the enhancement layer data has packet loss; a number of fourth video frames, a proportion of fourth video frames, or a number of base layer packet losses in the fourth video frames, the fourth video frame being a video frame in the at least one video frame in which the enhancement layer data has no packet loss and the base layer data has packet loss.

[0163] Optionally, the video frame comprises one or more data shards after network coding, and the video frame parameter further comprises at least one of: a number of data shards and / or a data amount of the data shards required by the terminal device to successfully decode one video frame; a number of data shards and / or a data amount of the data shards required by the terminal device to successfully decode the base layer of one video frame; and a number of data shards and / or a data amount of the data shards required by the terminal device to successfully decode the enhancement layer of one video frame.

[0164] In an example, the apparatus 1000 can implement the steps performed by the network device in the above method embodiments. The apparatus 1000 can be a network device, or a chip or circuit configured in the network device. The communication unit 1001 is configured to perform the transceiving operations of the network device in the above method embodiments, and the processing unit 1002 is configured to perform the processing-related operations of the network device in the above method embodiments.

[0165] For example, the communication unit 1001 is configured to send at least one video frame to a terminal device, and the communication unit 1001 is further configured to receive a video frame parameter from the terminal device, the video frame parameter being determined according to the at least one video frame. Optionally, the processing unit 1002 can be configured to optimize the transmission of a downlink video frame according to the video frame parameter.

[0166] Optionally, the video frame parameter comprises at least one of an extended delay parameter, an inter-frame distance parameter, a packet loss parameter, a late frame parameter, and a base layer and enhancement layer parameter.

[0167] Optionally, the extended delay parameter is used to indicate at least one of an extended delay of a first video frame, an extended delay of a plurality of video frames, a maximum extended delay of the plurality of video frames, a minimum extended delay of the plurality of video frames, and a variance of the extended delay of the plurality of video frames. The extended delay is a time length between a time when a terminal device successfully receives a first data packet of a video frame and a time when the terminal device successfully receives a last data packet of the video frame. The first video frame or the plurality of video frames are one or more video frames in at least one video frame received by the terminal device.

[0168] Optionally, the inter-frame distance parameter is used to indicate at least one of an inter-frame distance between adjacent video frames, an average value of a plurality of inter-frame distances, a maximum value of the plurality of inter-frame distances, a minimum value of the plurality of inter-frame distances, and a variance of the plurality of inter-frame distances. The terminal device receives a plurality of video frames from the network device.

[0169] Optionally, the packet loss parameter comprises at least one of a number of video frames with packet loss, a ratio of video frames with packet loss, and a number of consecutive K video frames with packet loss, where K is a positive integer greater than or equal to 1.

[0170] Optionally, the late frame parameter indicates at least one of a number of late video frames, a proportion of late video frames, a time difference between an actual reception time and a correct reception time of a late video frame, an average value of time differences between actual reception times and correct reception times of a plurality of late video frames, a maximum value of time differences between actual reception times and correct reception times of the plurality of late video frames, a minimum value of time differences between actual reception times and correct reception times of the plurality of late video frames, and a variance of time differences between actual reception times and correct reception times of the plurality of late video frames.

[0171] Optionally, each video frame comprises base layer data and enhancement layer data; the base layer and enhancement layer parameters comprise at least one of the following: a time difference between which the terminal device receives the base layer data and the enhancement layer data in a second video frame, the second video frame being a video frame in the at least one video frame; a number of third video frames, a proportion of third video frames, or at least one of a number of enhancement layer packet losses in third video frames, the third video frames being video frames in which the base layer data has no packet loss and the enhancement layer data has packet loss, and the third video frames being video frames in the at least one video frame; a number of fourth video frames, a proportion of fourth video frames, or at least one of a number of base layer packet losses in fourth video frames, the fourth video frames being video frames in which the base layer data has packet loss and the enhancement layer data has no packet loss, and the fourth video frames being video frames in the at least one video frame.

[0172] Optionally, the video frame after network coding comprises one or more data slices, and the video frame parameters further comprise at least one of the following: a number of data slices and / or a data amount of the data slices required for the terminal device to successfully decode one video frame; a number of data slices and / or a data amount of the data slices required for the terminal device to successfully decode the base layer of one video frame; a number of data slices and / or a data amount of the data slices required for the terminal device to successfully decode the enhancement layer of one video frame.

[0173] The division of units in the embodiments of the present application is illustrative, and is merely a logical functional division. Actual implementation can have another division manner. In addition, each functional unit in each embodiment of the present application can be integrated in one processor, or can be physically separated, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0174] It can be understood that the functions of the communication unit in the above embodiments can be realized by a transceiver, and the functions of the processing unit can be realized by a processor. The transceiver can include a transmitter and / or a receiver, etc., and is used to realize the functions of the sending unit and / or the receiving unit. The functions of the processing unit will be described below in combination with the functions of the transceiver. Figure 11 The above-mentioned embodiments are described by way of example.

[0175] Figure 11 is a schematic block diagram of the apparatus 1100 provided by the embodiments of the present application, Figure 11 The apparatus 1100 shown in the figure can be a hardware circuit implementation of the apparatus shown in the figure. Figure 10 The apparatus can perform the functions of the terminal device or the network device in the above-mentioned method embodiments. For ease of illustration, Figure 11 Only the main components of the communication apparatus are shown.

[0176] Figure 11 The communication device 1100 shown includes at least one processor 1101. The communication device 1100 may also include at least one memory 1102 for storing program instructions and / or data. The memory 1102 and the processor 1101 are coupled. The coupling in this embodiment is an indirect coupling or communication connection between devices, units, or modules, and can be electrical, mechanical, or other forms, used for information exchange between devices, units, or modules. The processor 1101 can operate collaboratively with the memory 1102, and the processor 1101 can execute program instructions stored in the memory 1102. At least one of the at least one memory 1102 may be included in the processor 1101.

[0177] The device 1100 may further include a communication interface 1103 for communicating with other devices via a transmission medium, thereby enabling the communication device 1100 to communicate with other devices. In this embodiment, the communication interface may be a transceiver, a circuit, a bus, a module, or other types of communication interface. In this embodiment, when the communication interface is a transceiver, the transceiver may include an independent receiver, an independent transmitter, or a transceiver with integrated transceiver functions, or an interface circuit.

[0178] It should be understood that the connection medium between the processor 1101, memory 1102, and communication interface 1103 described above is not limited in the embodiments of this application. The embodiments of this application... Figure 11 The memory 1102, processor 1101, and communication interface 1103 are connected via a communication bus 1104. Figure 11 The connections between other components are shown in bold and are for illustrative purposes only, not as limiting information. The bus may include an address bus, data bus, control bus, etc. For ease of illustration, Figure 11 The symbol is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0179] In one example, device 1100 is used to implement the steps performed by the terminal device in the above method embodiments. Communication interface 1103 is used to perform transmit / receive related operations of the terminal device in the above method embodiments, and processor 1101 is used to perform processing related operations of the terminal device in the above method embodiments.

[0180] For example, communication interface 1103 is used to receive at least one video frame from a network device; processor 1101 is used to determine video frame parameters based on the at least one video frame; communication interface 1103 is also used to send the video frame parameters to the network device.

[0181] Optionally, the video frame parameter comprises at least one of: an extended delay parameter, an inter-frame distance parameter, a packet loss parameter, a late frame parameter, a base layer and an enhanced layer parameter.

[0182] Optionally, the extended delay parameter is used to indicate at least one of: an extended delay of a first video frame, an average extended delay of a plurality of video frames, a maximum extended delay of the plurality of video frames, a minimum extended delay of the plurality of video frames, a variance of the extended delay of the plurality of video frames; wherein the extended delay is a time length between a time when a first data packet of a video frame is successfully received by the terminal device and a time when a last data packet of the video frame is successfully received, and the first video frame or the plurality of video frames are one or more video frames in at least one video frame received by the terminal device.

[0183] Optionally, the inter-frame distance parameter is used to indicate at least one of: an inter-frame distance between adjacent video frames, an average value of a plurality of inter-frame distances, a maximum value of the plurality of inter-frame distances, a minimum value of the plurality of inter-frame distances, a variance of the plurality of inter-frame distances; wherein the terminal device receives a plurality of video frames from the network device.

[0184] Optionally, the packet loss parameter comprises at least one of: a number of video frames with packet loss, a proportion of video frames with packet loss, a number of consecutive K video frames with packet loss, wherein K is a positive integer greater than or equal to 1.

[0185] Optionally, the processor 1101 is further configured to determine whether the packet data convergence protocol (PDCP) sequence numbers (SNs) of the data packets included in the received video frame are continuous, and determine that the video frame has packet loss when the PDCP SNs of the data packets included in the received video frame are not continuous.

[0186] Optionally, the late frame parameter indicates at least one of: a number of late video frames, a proportion of late video frames, a time difference between an actual receiving time and a correct receiving time of a late video frame, an average value of time differences between actual receiving times and correct receiving times of a plurality of late video frames, a maximum value of time differences between actual receiving times and correct receiving times of the plurality of late video frames, a minimum value of time differences between actual receiving times and correct receiving times of the plurality of late video frames, a variance of time differences between actual receiving times and correct receiving times of the plurality of late video frames.

[0187] Optionally, each of the video frames comprises base layer data and enhancement layer data; the base layer and enhancement layer parameters comprise at least one of: a time difference between which the terminal device receives the base layer data and the enhancement layer data in a second video frame, the second video frame being a video frame in the at least one video frame; a number of third video frames, a proportion of third video frames, or a number of enhancement layer packet losses in third video frames, the third video frames being video frames in which the base layer data has no packet loss and the enhancement layer data has packet loss, and the third video frames being video frames in the at least one video frame; a number of fourth video frames, a proportion of fourth video frames, or a number of base layer packet losses in fourth video frames, the fourth video frames being video frames in which the base layer data has packet loss and the enhancement layer data has no packet loss, and the fourth video frames being video frames in the at least one video frame.

[0188] Optionally, the video frames, after being network encoded, comprise one or more data slices, and the video frame parameters further comprise at least one of: a number of data slices and / or a data amount of the data slices required for the terminal device to successfully decode one video frame; a number of data slices and / or a data amount of the data slices required for the terminal device to successfully decode the base layer of one video frame; a number of data slices and / or a data amount of the data slices required for the terminal device to successfully decode the enhancement layer of one video frame.

[0189] In another example, the apparatus 1100 is configured to implement the steps performed by the network device in the above method embodiments. The communication interface 1103 is configured to perform the receiving and transmitting related operations of the network device in the above method embodiments, and the processor 1101 is configured to perform the processing related operations of the network device in the above method embodiments.

[0190] For example, the communication interface 1103 is configured to send at least one video frame to a terminal device, and the communication interface 1103 is further configured to receive video frame parameters from the terminal device, the video frame parameters being determined according to the at least one video frame. Optionally, the processor 1101 is configured to optimize the transmission of the downlink video frames according to the video frame parameters.

[0191] Optionally, the video frame parameters comprise at least one of: an extended delay parameter, an inter-frame distance parameter, a packet loss parameter, a late arrival parameter, a base layer and enhancement layer parameter.

[0192] Optionally, the extended delay parameter is used to indicate at least one of: an extended delay of a first video frame, extended delays of a plurality of video frames, a maximum extended delay of the plurality of video frames, a minimum extended delay of the plurality of video frames, a variance of the extended delays of the plurality of video frames; wherein the extended delay is a time length between a time when the terminal device successfully receives a first packet of a video frame and a time when the terminal device successfully receives a last packet of the video frame, and the first video frame or the plurality of video frames are one or more of the at least one video frame received by the terminal device.

[0193] Optionally, the frame interval parameter is used to indicate at least one of: a frame interval between adjacent video frames, an average of a plurality of frame intervals, a maximum of the plurality of frame intervals, a minimum of the plurality of frame intervals, a variance of the plurality of frame intervals; wherein the terminal device receives a plurality of video frames from the network device.

[0194] Optionally, the packet loss parameter comprises at least one of: a number of video frames with packet loss, a ratio of video frames with packet loss, a number of consecutive K video frames with packet loss, wherein K is a positive integer greater than or equal to 1.

[0195] Optionally, the late parameter indicates at least one of: a number of late video frames, a proportion of late video frames, a time difference between an actual receiving time and a correct receiving time of a late video frame, an average of time differences between actual receiving times and correct receiving times of a plurality of late video frames, a maximum of time differences between actual receiving times and correct receiving times of the plurality of late video frames, a minimum of time differences between actual receiving times and correct receiving times of the plurality of late video frames, a variance of time differences between actual receiving times and correct receiving times of the plurality of late video frames.

[0196] Optionally, each video frame comprises base layer data and enhancement layer data; and the base layer and enhancement layer parameter comprises at least one of: a time difference between a time when the terminal device receives the base layer data and a time when the terminal device receives the enhancement layer data in a second video frame, wherein the second video frame is one of the at least one video frame; at least one of a number of third video frames, a proportion of third video frames, or a number of enhancement layer packet losses in the third video frames, wherein the third video frame is a video frame in which the base layer data has no packet loss and the enhancement layer data has packet loss, and the third video frame is one of the at least one video frame; at least one of a number of fourth video frames, a proportion of fourth video frames, or a number of base layer packet losses in the fourth video frames, wherein the fourth video frame is a video frame in which the base layer data has packet loss and the enhancement layer data has no packet loss, and the fourth video frame is one of the at least one video frame.

[0197] Optionally, the video frame after network coding comprises one or more data slices, and the video frame parameter further comprises at least one of the following: a number of data slices required by the terminal device to successfully decode one video frame and / or a data amount of the data slices; a number of data slices required by the terminal device to successfully decode a base layer of one video frame and / or a data amount of the data slices; and a number of data slices required by the terminal device to successfully decode an enhancement layer of one video frame and / or a data amount of the data slices.

[0198] Further, the embodiments of the present application also provide a device for executing the method in the above method embodiments. A computer readable storage medium comprises a program, when the program is run by a processor, the method in the above method embodiments is executed. A computer program product comprises computer program code, when the computer program code is run on a computer, the computer implements the method in the above method embodiments. A chip comprises a processor coupled with a memory, the memory is used to store programs or instructions, when the programs or instructions are executed by the processor, the device executes the method in the above method embodiments. A system comprises at least one of the terminal device and the network device executing the above method embodiments. Figure 5 Figure 5 Figure 5 Figure 5

[0199] In the embodiments of the present application, the processor can be a general processor, a digital signal processor, an application specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and can implement or execute the disclosed methods, steps and logic block diagrams in the embodiments of the present application. The general processor can be a microprocessor or any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as hardware processor execution or executed by a combination of hardware and software modules in the processor.

[0200] In the embodiments of the present application, the memory can be a non-volatile memory such as a hard disk drive (HDD) or a solid-state drive (SSD), and can also be a volatile memory such as a random-access memory (RAM). The memory can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but is not limited to this. The memory in the embodiments of the present application can also be a circuit or any other device capable of realizing a storage function, used for storing program instructions and / or data.​​​​

[0201] The method provided by the embodiments of the present application can be implemented by software, hardware, firmware or any combination thereof, in whole or in part. When implemented by software, the method can be implemented in the form of a computer program product, in whole or in part. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment or other programmable apparatus. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as an SSD), etc.

[0202] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the claims of the present application and their equivalents, the present application also intends to include these modifications and variations.

Claims

1. A communication method characterized by comprising: The method comprises: a terminal device receiving at least one video frame from a network device; the terminal device determining a video frame parameter according to the at least one video frame; the terminal device sending the video frame parameter to the network device; the video frame parameter is determined in a granularity of a video frame, and the video frame parameter comprises at least one of the following: an extended delay parameter, an inter-frame distance parameter, a video frame packet loss parameter, a late frame parameter, and a base layer and enhancement layer parameter.

2. The method of claim 1, wherein: the extended delay parameter is used to indicate at least one of the following: an extended delay of a first video frame, an average extended delay of a plurality of video frames, a maximum extended delay of the plurality of video frames, a minimum extended delay of the plurality of video frames, and a variance of the extended delay of the plurality of video frames; wherein the extended delay is a time length between a time when the terminal device successfully receives a first packet of a video frame and a time when the terminal device successfully receives a last packet of the video frame, and the first video frame or the plurality of video frames are one or more video frames in the at least one video frame received by the terminal device.

3. The method of claim 1, wherein: the inter-frame distance parameter is used to indicate at least one of the following: an inter-frame distance between adjacent video frames, an average value of a plurality of inter-frame distances, a maximum value of the plurality of inter-frame distances, a minimum value of the plurality of inter-frame distances, and a variance of the plurality of inter-frame distances; wherein the terminal device receives a plurality of video frames from the network device.

4. The method of claim 1, wherein, the packet loss parameter comprises at least one of the following: a number of video frames with packet loss, a proportion of video frames with packet loss, and K consecutive video frames with packet loss, wherein K is a positive integer greater than or equal to 1.

5. The method of claim 4, wherein, The method further comprises: the terminal device determining whether packet data convergence protocol (PDCP) sequence numbers (SNs) of data packets included in the received video frame are continuous; if the PDCP SNs of the data packets included in the received video frame are not continuous, the terminal device determines that the video frame has packet loss.

6. The method of claim 1, wherein, the late frame parameter indicates at least one of the following: a number of late video frames, a proportion of late video frames, a time difference between an actual receiving time and a correct receiving time of a late video frame, an average value of time differences between actual receiving times and correct receiving times of a plurality of late video frames, a maximum value of time differences between actual receiving times and correct receiving times of the plurality of late video frames, a minimum value of time differences between actual receiving times and correct receiving times of the plurality of late video frames, and a variance of time differences between actual receiving times and correct receiving times of the plurality of late video frames.

7. The method of claim 1, wherein: each video frame comprises base layer data and enhancement layer data; the base layer and enhancement layer parameter comprises at least one of the following: a time difference between the terminal device receiving the base layer data and the enhancement layer data in a second video frame, wherein the second video frame is a video frame in the at least one video frame; a number of third video frames, a proportion of third video frames, or at least one of a number of enhancement layer packet losses in third video frames, wherein the third video frames are video frames in which the base layer data has no packet loss and the enhancement layer data has packet loss, and the third video frames are video frames in the at least one video frame. a fourth video frame, a proportion of the fourth video frame, or at least one of a number of lost packets of base layer in the fourth video frame, the fourth video frame being a video frame in which there is no lost packet of enhancement layer data, there is lost packet of base layer data, and the fourth video frame is a video frame in the at least one video frame.

8. The method of any one of claims 1 to 7, wherein, The video frame after being network encoded comprises one or more data fragments, and the video frame parameter further comprises at least one of: a number of data fragments and / or a data amount of the data fragments required for the terminal device to successfully decode one video frame; a number of data fragments and / or a data amount of the data fragments required for the terminal device to successfully decode a base layer of one video frame; a number of data fragments and / or a data amount of the data fragments required for the terminal device to successfully decode an enhancement layer of one video frame.

9. A communication method characterized by comprising: comprises: the network device sending at least one video frame to the terminal device; the network device receiving a video frame parameter from the terminal device, the video frame parameter being determined according to the at least one video frame; wherein the video frame parameter is determined in a granularity of a video frame, and the video frame parameter comprises at least one of: an extended delay parameter, an inter-frame interval parameter, a lost packet parameter of a video frame, a late arrival parameter, a base layer and an enhancement layer parameter.

10. The method of claim 9, wherein the extended delay parameter is used to indicate at least one of: an extended delay of a first video frame, extended delays of a plurality of video frames, a maximum extended delay of the plurality of video frames, a minimum extended delay of the plurality of video frames, a variance of the extended delays of the plurality of video frames; wherein the extended delay is a time length between a time when a first data packet of a video frame is successfully received by the terminal device and a time when a last data packet of the video frame is successfully received, and the first video frame or the plurality of video frames is one or more video frames in the at least one video frame received by the terminal device.

11. The method of claim 9, wherein the inter-frame interval parameter is used to indicate at least one of: an inter-frame interval of adjacent video frames, an average of a plurality of inter-frame intervals, a maximum of the plurality of inter-frame intervals, a minimum of the plurality of inter-frame intervals, a variance of the plurality of inter-frame intervals; wherein the terminal device receives a plurality of video frames from the network device.

12. The method of claim 9, wherein, the lost packet parameter comprises at least one of: a number of video frames in which there is lost packet, a proportion of the video frames in which there is lost packet, a number of consecutive K video frames in which there is lost packet, the K being a positive integer greater than or equal to 1.

13. The method of claim 9, wherein, the late arrival parameter indicates at least one of: a number of late arrival video frames, a proportion of the late arrival video frames, a time difference between an actual receiving time and a correct receiving time of one late arrival video frame, an average of time differences between actual receiving times and correct receiving times of a plurality of late arrival video frames, a maximum of time differences between actual receiving times and correct receiving times of the plurality of late arrival video frames, a minimum of time differences between actual receiving times and correct receiving times of the plurality of late arrival video frames, a variance of time differences between actual receiving times and correct receiving times of the plurality of late arrival video frames.

14. The method of claim 9, wherein Each video frame comprises base layer data and enhancement layer data; The base layer and enhancement layer parameters comprise at least one of: A time difference between when the terminal device receives the base layer data and the enhancement layer data in a second video frame, the second video frame being a video frame in the at least one video frame; At least one of a number of third video frames, a proportion of third video frames, or a number of enhancement layer packet losses in third video frames, the third video frames being video frames in which the base layer data has no packet loss and the enhancement layer data has packet loss, and the third video frames being video frames in the at least one video frame; At least one of a number of fourth video frames, a proportion of fourth video frames, or a number of base layer packet losses in fourth video frames, the fourth video frames being video frames in which the base layer data has packet loss and the enhancement layer data has no packet loss, and the fourth video frames being video frames in the at least one video frame.

15. The method according to any one of claims 9 to 14, characterized in that, The video frame, after being network encoded, comprises one or more data slices, and the video frame parameters further comprise at least one of: A number of data slices and / or a data amount of the data slices that the terminal device needs to receive to successfully decode a video frame; A number of data slices and / or a data amount of the data slices that the terminal device needs to receive to successfully decode the base layer of a video frame; A number of data slices and / or a data amount of the data slices that the terminal device needs to receive to successfully decode the enhancement layer of a video frame.

16. An apparatus, comprising: A unit comprising each step of the method of any one of claims 1 to 8, or a unit comprising each step of the method of any one of claims 9 to 15.

17. An apparatus, comprising: An apparatus comprising at least one processor and interface circuitry, the at least one processor configured to communicate with other apparatuses via the interface circuitry and perform the method of any one of claims 1 to 8, or perform the method of any one of claims 9 to 15.

18. An apparatus, comprising: A processor configured to invoke a program stored in a memory to perform the method of any one of claims 1 to 8, or perform the method of any one of claims 9 to 15.

19. A computer-readable storage medium, characterized in that, A program that, when executed by a processor, causes the method of any one of claims 1 to 8 to be performed, or the method of any one of claims 9 to 15 to be performed.

20. A computer program product, characterised in that, A computer program product comprising a computer program that, when executed by a processor, causes the method of any one of claims 1 to 8 to be performed, or the method of any one of claims 9 to 15 to be performed.

Citation Information

Patent Citations

  • Video data transmission method, device and system

    CN102547376A