Video stream low-delay transmission and real-time decoding rendering method based on TCP / UDP cooperation

By adopting the TCP/UDP collaboration method in video streaming, combining the data integrity of TCP and the real-time UDP, the problems of high latency, insufficient reliability and poor multi-platform adaptability in video streaming are solved, and video streaming and rendering effects with low latency, high reliability and cross-platform compatibility are achieved.

CN119922376APending Publication Date: 2025-05-02HUNAN INSTITUTE OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510068357.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

There are problems of high latency, insufficient reliability and poor multi-platform adaptability in existing video streaming.

Method used

The video stream low-latency transmission and real-time decoding and rendering method based on TCP/UDP are adopted to transmit video stream data through TCP to ensure data integrity and reliability, and to combine UDP for connection management and signaling interaction to improve the system's real-time response capability and stability. The receiver uses an efficient decoder for real-time decoding, and uses GPU hardware acceleration technology to achieve efficient rendering and dynamic display of video frames.

Benefits of technology

It achieves low latency, high reliability, high rendering efficiency and cross-platform compatibility, and is suitable for real-time video application scenarios such as online live broadcast, video conferencing, and distance education.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119922376A_ABST
    Figure CN119922376A_ABST
Patent Text Reader

Abstract

The invention provides a TCP / UDP (Transmission Control Protocol / User Datagram Protocol) collaboration-based video stream low-delay transmission and real-time decoding rendering method. The method comprises the following steps of: carrying out connection management and signaling interaction; the sending end captures data for coding, and dynamically captures screen content; synchronizing the collected multi-source data according to a timestamp, transmitting the data to an encoder, and preprocessing the data before transmission; a packaged data packet is sent according to a designed transmission protocol, after receiving the data packet, a receiving end analyzes and restores the data packet according to the protocol to obtain data format information, creates and initializes decoder parameters, and then sends a prompt that preparatory work is ready to wait for receiving video stream data to a sending end; video stream data in the data transmission process are transmitted through a TCP, and in the process of rendering and displaying frame data after the video stream data are received by a receiving end, hardware acceleration is performed by using an OpenGL and utilizing a GPU, and decoded frames are displayed on a screen in real time. The method has the advantages of low delay, high reliability, high rendering efficiency, cross-platform compatibility and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer wireless communication, and in particular to a method for low-delay transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration. Background Art

[0002] In recent years, real-time video applications such as video conferencing, online live broadcasting, distance education, and cloud gaming have developed rapidly. These applications have extremely high requirements for the quality of video streaming and rendering, especially in terms of latency control, transmission reliability, multi-platform compatibility, and high-performance rendering. However, existing technologies still face many challenges in this field.

[0003] First, real-time video applications have extremely high latency requirements. In interactive scenarios such as video conferencing and cloud gaming, any jitter, packet loss, and encoding and decoding processing delays in network transmission will significantly affect the user experience. Traditional video transmission technologies often have difficulty achieving sufficiently low latency while ensuring video quality, thus limiting the further development of these applications.

[0004] Secondly, multi-platform compatibility is also an issue that needs to be addressed urgently. Differences in hardware architecture and operating systems of different devices make it complicated and difficult to adapt video applications on different platforms. High-resolution and high-frame-rate video streams place higher demands on network bandwidth, storage space, and rendering performance, which further exacerbates the challenge of multi-platform compatibility.

[0005] In addition, traditional transmission and processing methods often have difficulty balancing low latency and high reliability when facing unstable network environments or high performance requirements. Network instability may cause problems such as video streaming freezes and interruptions, while high performance requirements require video applications to achieve high-quality transmission and rendering with limited resources.

[0006] Therefore, optimizing transmission architecture, improving decoding and rendering efficiency, enhancing multi-platform compatibility, and intelligent network scheduling have become key directions for promoting the development of real-time video technology. These technical improvements are aimed at meeting users' needs for high-quality real-time interaction and promoting the development of real-time video applications to a higher level. Summary of the invention

[0007] In view of the above, the present invention provides a method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration, aiming to solve the problems of high latency, insufficient reliability and poor multi-platform adaptability in existing video stream transmission.

[0008] The technical solution of the present invention:

[0009] The present invention provides a method for low-delay transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration, which is characterized by comprising the following steps:

[0010] S1. Perform connection management and signaling interaction;

[0011] S2. The sender (mobile phone) captures data for encoding. Android uses the MediaProjection API, and iOS uses ReplayKit to achieve dynamic capture of screen content.

[0012] S3. The collected multi-source data is synchronized according to the timestamp and transmitted to the encoder, and pre-processed before transmission to ensure the integrity, real-time and stability of data capture;

[0013] S4. Send the packaged data packet according to the designed transmission protocol. After receiving the data packet, the receiving end (computer) parses and restores the data packet according to the protocol to obtain the data format information, create and initialize the decoder parameters, and then sends a prompt to the sending end that the preparation is ready to receive the video stream data;

[0014] S5. The video stream data in the data transmission process is transmitted via TCP. After the receiving end receives the video stream data, in the process of rendering and displaying the frame data, OpenGL is used to utilize GPU for hardware acceleration to display the decoded frame on the screen in real time.

[0015] Furthermore, in step S1, UDP is combined for connection management and signaling interaction, and the UDP protocol is used to send data packets to all devices in the same local area network. First, the client initiates a connection and sends an initial signaling packet (connection request packet), which includes a device identifier and a session ID, to notify the server to establish a TCP connection. The server then responds. After receiving the connection request, the server returns a confirmation signaling packet, which includes confirmation information and connection parameters.

[0016] Furthermore, on the Android platform, the MediaProjection API is an officially provided screen content capture mechanism. The captured content contains sensitive information and requires explicit user authorization before it can be used for real-time screen recording or screen sharing;

[0017] On the iOS platform, ReplayKit is an official screen recording and streaming media sharing API (Application Programming Interface) provided by Apple. It supports devices starting from iOS 9. When using ReplayKit, the system will automatically pop up a user authorization prompt, and the user must agree to record the screen.

[0018] Furthermore, in step S3, pre-processing is performed before transmission including resolution adjustment, frame rate control and bit rate setting.

[0019] Furthermore, in step S3, specifically: synchronization of multi-source data is a key operation based on timestamp alignment. When processing screen capture video, the timestamp records each data stream (screen content, camera video), and is accompanied by a timestamp during acquisition to mark the moment of data acquisition. The timestamp format is an offset relative to the start time. A queue buffer mechanism is used to cache the collected data of each data source separately, align each data source according to the timestamp, and eliminate advanced or lagging data frames to ensure that multi-source data can be transmitted synchronously to the encoder. The encoder is responsible for converting the original multimedia data into a specific compression format for transmission, storage and decoding. If some data frames are lost, they are supplemented by repeating the previous frame.

[0020] Furthermore, before the synchronized data is passed to the encoder, preprocessing is required to ensure transmission efficiency and streaming quality. Different devices (such as mobile devices and computers) have different support capabilities for video resolution. The resolution needs to be adjusted according to the target device to adapt to its decoding capabilities. The image processing library (OpenCV) is used to scale the resolution of the video frames. When the network bandwidth is limited, the resolution and frame rate are reduced to reduce the amount of data. The bit rate is dynamically adjusted according to the network conditions to ensure the stability of the transmission. In a multi-threaded environment, a thread-safe queue is used to cache the collected data.

[0021] Furthermore, after receiving the data packet, the receiving end parses and restores the data body, obtains the data format information, creates and initializes the decoder parameters, wherein the decoder is used to restore the encoded compressed video stream into playable video data. The decoder is used to perform reverse processing on the compressed data generated by the encoder to restore the original picture, and then sends feedback to the sending end that the decoding preparation is ready and waiting to receive the video stream data.

[0022] Furthermore, after receiving the video stream data, after the decoder decodes and obtains the frame data, the frame data is rendered and displayed, wherein rendering refers to the process of converting the decoded image data into a format that can be displayed on the screen.

[0023] Furthermore, the key points of rendering include: performing color space conversion on the decoded frame data, converting the decoded YUV (Y represents the brightness component, U and V represent the chrominance components) format image into RGB (Red, Green, Blue) format, adapting it to the screen display, and using OpenGL to utilize the GPU for hardware acceleration during the rendering process to improve rendering efficiency.

[0024] Furthermore, display is the process of presenting the rendered frames to the screen. The key points include: presenting each frame of the image to the screen in chronological order according to the screen refresh rate, and then the GPU draws the rendered texture to the display buffer (Frame Buffer). To avoid screen tearing, the buffer adopts a double buffer mechanism, one buffer (Front Buffer) displays the current frame, and the other buffer (Back Buffer) loads the next frame of the image. The buffer is switched when the screen is refreshed, and the displayed frame is synchronized with the timestamp of the video stream to avoid screen freeze or delay. The final rendering result is output to the screen by the GPU, and the entire display process is completed.

[0025] The present invention provides a method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration, which can solve the problems of high latency, insufficient reliability and poor multi-platform adaptability in existing video stream transmission. This method transmits the data in the video stream through TCP to ensure the integrity and reliability of the data by designing a collaborative transmission architecture combining TCP and UDP; at the same time, UDP is combined for connection management and signaling interaction to improve the real-time response capability and stability of the overall system. After obtaining the video stream data, the receiving end performs real-time decoding through an efficient decoder, and combines the graphics processing unit (GPU, Graphics Processing Unit) hardware acceleration technology (OpenGL) to achieve efficient rendering and dynamic display of video frames. The present invention can support mainstream mobile operating systems such as Android and Apple (iOS), and can achieve smooth operation on different mobile phone platforms. It has the advantages of low latency, high reliability, high rendering efficiency and cross-platform compatibility, and is suitable for real-time video application scenarios such as online live broadcast, video conferencing, and distance education.

[0026] The preferred embodiments of the present invention and their beneficial effects will be further described in detail in conjunction with specific implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the present invention, but should not be construed as limiting the present invention. In the accompanying drawings:

[0028] Figure 1 It is a flow chart of the method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration of the present invention;

[0029] Figure 2 It is a schematic diagram of frames, slices and macroblocks of the present invention;

[0030] Figure 3 It is a schematic diagram of the H264 structure overview of the present invention. DETAILED DESCRIPTION

[0031] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.

[0032] See also Figure 1 The present invention provides a method for low-delay transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration, which is characterized by comprising the following steps:

[0033] S1. Perform connection management and signaling interaction;

[0034] S2. The sender (mobile phone) captures data for encoding. Android uses the MediaProjection API, and iOS uses ReplayKit to achieve dynamic capture of screen content.

[0035] S3. The collected multi-source data is synchronized according to the timestamp and passed to the encoder. Before transmission, pre-processing such as resolution adjustment, frame rate control and bit rate setting is performed to ensure the integrity, real-time and stability of data capture.

[0036] S4. Send the packaged data packet according to the designed transmission protocol. After receiving the data packet, the receiving end (computer) parses and restores the data packet according to the protocol to obtain the data format information, create and initialize the decoder parameters, and then sends a prompt to the sending end that the preparation is ready to receive the video stream data;

[0037] S5. The video stream data in the data transmission process is transmitted via TCP. After the receiving end receives the video stream data, in the process of rendering and displaying the frame data, OpenGL is used to utilize GPU for hardware acceleration to display the decoded frame on the screen in real time.

[0038] The present invention provides a method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration, which can solve the problems of high latency, insufficient reliability and poor multi-platform adaptability in existing video stream transmission. This method transmits the data in the video stream through TCP to ensure the integrity and reliability of the data by designing a collaborative transmission architecture combining TCP and UDP; at the same time, UDP is combined for connection management and signaling interaction to improve the real-time response capability and stability of the overall system. After obtaining the video stream data, the receiving end performs real-time decoding through an efficient decoder, and combines the graphics processing unit (GPU, Graphics Processing Unit) hardware acceleration technology (OpenGL) to achieve efficient rendering and dynamic display of video frames. The present invention can support mainstream mobile operating systems such as Android and Apple (iOS), and can achieve smooth operation on different mobile phone platforms. It has the advantages of low latency, high reliability, high rendering efficiency and cross-platform compatibility, and is suitable for real-time video application scenarios such as online live broadcast, video conferencing, and distance education.

[0039] The present invention provides a method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration, which is specifically described below.

[0040] Depend on Figure 1As can be seen from the algorithm flow structure diagram, after initializing some necessary algorithm parameters, the device needs to be connected. First, UDP is used to broadcast, because UDP supports broadcast and multicast, and can quickly send messages to all devices in a network. At this time, other devices in the LAN can be quickly discovered without establishing a connection. In addition, UDP uses a lightweight protocol, with a small data packet header and low communication overhead, which is suitable for quickly sending service discovery or device status messages in the LAN. The client broadcasts a data packet through UDP, which contains some key information (device identification, IP address, port number, etc.). The server in the LAN listens to the broadcast port. After receiving the broadcast, it parses the data packet to confirm whether it is a valid request. If the request is confirmed to be valid, the server will return a UDP data packet containing the device's IP address and available TCP port. Then a TCP connection is established. The advantage of the TCP connection method is that it is very reliable. The client initiates a TCP connection through the IP address and port number returned by the server. TCP establishes a three-way handshake. Data transmission has a sequence control and error checking mechanism to ensure connection reliability. Key data (video stream data, key control interaction information such as starting or stopping screen projection, etc.) is transmitted between the device and the server through TCP to ensure that the data packets arrive in order and perform integrity checks. At the same time, there is also a trusted device identification bit. A trusted device is a device that has been successfully connected once. When it receives a connection request from a trusted device during subsequent connections, it will directly reply with its own device information, reducing the manual confirmation steps and further optimizing the user experience. The parameter information of the device information reply is shown in Table 1: It is used to reply to device searches and device information requests. The command code is 0x1001, and the data body is in the form of json (JavaScript Object Notation). The content is terminal information, and the fields are as follows. When UDP is sent, only information is reported. When TCP is sent, the client is connected. When requesting a device search, the reply uses TCP. The following is an example of device information:

[0041] {"ip":"169.254.159.1","cmdPort":9216,"dataPort":9226,"uuid":"7BE0BF2D12","os":0,"name":"HP"}

[0042] Table 1 Device information reply format parameter information

[0043]

[0044] The sender performs screen capture, which means dynamically capturing the content on the device screen as frame data for subsequent encoding and transmission. The following describes its implementation logic in detail, taking Android and iOS systems as examples. First, we initialize the screen capture tool, select the appropriate API according to the device operating system (such as Android's MediaProjection API or iOS's ReplayKit), and then request screen capture permission. After obtaining the permission, we set the defined capture area (full screen or partial area). In this application, we choose full screen. It should be noted that the full screen of mobile phones of different systems may not include the top status bar information (such as Xiaomi's Pengpai system). Next, on the Android platform, create a virtual display, obtain the MediaProjection object, create a virtual display to specify the capture resolution, pixel format, and frame rate. Use the ImageReader object to receive the captured frame data. ImageReader will output frame data in RGBA format. In order to facilitate the subsequent transmission of the video stream, the frame data is converted to the necessary format (from RGBA to YUV420). When the capture is stopped, the resource MediaProjection is released. On the iOS platform, frame data callback and error handling callback are provided after the ReplayKit framework is initialized. The frame data is then decoded into raw pixel data using CMSampleBuffer.

[0045] Because the frame data generated by screen capture is usually stored in RGBA or BGRA format, it needs to be converted to YUV format according to the encoder requirements. The capture resolution and frame rate will directly affect the system performance and network transmission load. The parameters are adjusted dynamically to strike a balance between smoothness and clarity. Because the acquisition of frame data involves a large amount of memory allocation and release, memory management is a key factor, especially in applications that need to process and store a large amount of data in real time. If the memory management is improper, it will cause memory leaks or the program allocates memory but fails to release it, resulting in a gradual decrease in available memory, which may eventually cause the program to crash or the system to become slow, thus affecting the projection effect. In this application, a memory pool is used to manage memory allocation. The memory pool can pre-allocate a large memory area and then allocate small blocks of memory from it, which can reduce frequent allocation and release operations and improve performance. At the same time, regularly check and clean up frame data that is no longer used to ensure that the memory usage rate remains at a reasonable level. Even if an abnormal situation occurs, when handling the abnormal situation, it is ensured that the allocated memory can be correctly released. In addition to memory management issues, there are also redundant data processing issues. Among them, redundant data refers to redundant data that may be generated at different processing stages when processing video frames. For example, the same frame is captured multiple times without releasing the previous frame data in time. In this application, a first-in-first-out (FIFO) buffer is used for effective management. If the old frames do not need to be retained, they are directly overwritten to save memory. Through the above method, memory leaks and redundant data occupation can be effectively avoided, ensuring that the process of capturing frame data is efficient and stable, and effectively reducing the delay caused by the sender.

[0046] The rendering method of the present application can be used as a screen projection method, involving multi-source data processing and real-time systems, so the synchronization of timestamps is crucial. A timestamp is a numerical value or string representing a specific point in time, usually used to record the time when an event occurs. Data from different sources may be generated or arrive at the system at different points in time. Through timestamp synchronization, it can be ensured that they are aligned in real time order. It supports multi-source data fusion. When capturing multiple sensors or signal sources (such as screen capture and user operation data), timestamp synchronization can unify data from different sources into a time base for subsequent processing or analysis. In a system with high real-time requirements such as screen projection, timestamp synchronization can ensure that data arrives and is processed in the correct time order, avoiding the transmission of redundant or outdated data. At the same time, optimize network bandwidth utilization to avoid delays or freezes.

[0047] In the initialization, the default screen resolution of the sender is used, and the frame rate is 30 frames. 30 frames (30FPS, Frames Per Second) is a common frame rate in videos and animations, which means that 30 frames of images are displayed every second. In most cases, 30FPS is a frame rate that balances fluency and file size, suitable for recording and playing various video content, and provides sufficient fluency. If the network status is good, the bandwidth utilization can be further improved, and the frame rate can be increased to 60 frames to provide users with a smoother screen projection experience. 60 frames is a common frame rate for modern video games. If the network condition is poor, the frame rate can be appropriately reduced to 24 frames, which is the standard frame rate in the film industry, providing natural and smooth visual effects. Similarly, the bit rate is also dynamically set. The default bit rate is set to 10000kbps. The bit rate (Bit Rate) refers to the amount of data transmitted in a certain period of time, usually expressed in bits per second (bps). It is an important indicator for measuring the quality and size of audio or video files. The video compression of this application adopts the H.264 standard, H.264, also known as AVC (Advanced Video Coding), is an efficient and widely used video coding standard, suitable for the compression and transmission of a variety of video content. Its high compression efficiency and good image quality make it occupy an important position in modern video technology. It is jointly formulated by the International Organization for Standardization (ISO) and the International Telecommunication Union (ITU) to provide efficient video encoding and decoding. Therefore, the bit rate selection of this application takes into account the resolution, frame rate and complexity of the content of the video. Optimizing the bit rate helps to reduce file size and bandwidth requirements while maintaining video quality. Then the data is encoded using H.264, which is mainly used to compress and encode the uncompressed original video data (usually in YUV format, where Y represents the brightness component, and U and V represent the chrominance components). In order to improve the encoding efficiency, H.264 slices each frame of the image, and each slice is processed by one or more macroblocks. Each macroblock is a 16x16 pixel area. The macroblock can be further divided into sub-blocks (such as 4x4 or 8x8) to achieve more refined encoding. The purpose of block processing is to facilitate subsequent prediction, transformation, and quantization. The relationship between frames, slices, and macroblocks is as follows: Figure 2 shown.

[0048] After the encoding is completed, it is packaged and sent according to the established transmission protocol. The custom transmission protocol is based on the standard protocol. The format and processing logic of the data packet are designed according to the actual business needs to meet specific needs. Different business data (such as video frames, control commands) are distinguished according to the identification data type. The main method is to ensure that the data packet has not been tampered with or damaged through the verification mechanism to ensure data integrity. After the original packaged data is prepared, the header information is generated according to the protocol format, and then the original data is attached to the header as the data body, and then a fixed identifier and check code are added to mark the end of the data packet. The decoding process is the reverse process of this process, but the main thing is to avoid the sticky packet problem. The sticky packet problem (StickyPacket Problem) refers to the situation that occurs in network communications, especially in TCP-based protocols. It involves the situation where multiple data packets are stuck together during transmission, resulting in the receiver being unable to correctly distinguish and process separate data packets. The decoding process is the process of the receiver parsing and restoring the received data packet according to the custom transmission protocol, which mainly includes the following steps:

[0049] 1. The receiving end receives the video stream data packet through the network socket and caches it into a receiving buffer. Due to the asynchronous nature of network transmission, packet segmentation and packet sticking may occur, so the data stream needs to be checked for integrity. The reason is that TCP is a connection-oriented protocol that treats data as a byte stream. The data received by the receiver does not necessarily correspond to the data packet sent by the sender. In addition, when the sender sends multiple small data packets in a short period of time, the receiver may receive a merged data packet. The final result is that the receiver may not be able to determine the boundaries of each data packet when reading the data, resulting in data parsing errors. Therefore, a packet header needs to be added before each data packet. The packet header contains the length information of the data packet, and the receiver can parse the data according to the length. And a specific delimiter is used to identify the start and end of the data packet, and the receiver divides the data according to the delimiter. The protocol format information of the packet header is shown in Table 2, and the protocol format information of the packet tail is shown in Table 3.

[0050] Table 2 Packet header format parameter information

[0051]

[0052] Table 3 Packet tail format parameter information

[0053]

[0054] 2. After receiving the complete data packet, parse the header, data body and tail information according to the format of the custom protocol. First, read the version number, data type, data body length and other fields and verify whether the protocol version is correct. Then read the data body according to the length specified in the header. Finally, confirm the fixed mark to determine whether the data transmission is packetized and verify the integrity of the data packet.

[0055] 3. According to the type of data body parsed (such as video stream data, control signals such as signaling information such as starting or stopping screen projection), restore the data body to the original content. If the data body is a video frame, use the H.264 decoder to decode the video frame. If the data body is control information, parse the control command according to the predefined format.

[0056] 4. Use the H.264 decoder to decode the video frame. The basic components of H.264 data are mainly composed of NALU (Network Abstraction Layer Unit), Slice and bit stream. Its structure is as follows Figure 3As shown. Common NALUs mainly include IDR frames (key frames, Intra-coded Frames), P frames (predicted frames, Predicted Frames) and B frames (bi-predictive frames, Bi-predictive Frames). Among them, the key frame is a complete frame in the video stream, which contains complete image data. Slice further divides each frame into fragments for parallel processing. The bitstream is the NAL unit arranged into a bitstream for the decoder to parse. Decoding first requires parsing the bitstream. The video stream received from the network is a continuous bitstream, which needs to be parsed into NAL units. Specifically, the start code is checked to separate the NAL units. Then the type and content of the NAL unit are extracted and unnecessary information (filling bytes) are skipped. After parsing the bitstream, the NAL unit is decoded and processed according to the type of the NAL unit, mainly based on the key frame and the predicted frame. After the NAL unit is decoded, entropy decoding is performed. H.264 uses entropy coding technology to compress video data, and decoding is required to restore the original data. The two entropy coding methods in H.264 are CAVLC (Context-Adaptive Variable Length Coding) or CABAC (Context-Adaptive Binary Arithmetic Coding). The corresponding decoder selects the decoding method based on the header information. After that, the quantized coefficients are restored to frequency domain data closer to the original value based on inverse quantization and the inverse transform is used to restore the data from the frequency domain to the spatial domain through inverse discrete cosine transform (iDCT, Inverse Discrete Cosine Transform) or integer transform. After that, motion compensation and frame reconstruction are performed. The decoder predicts the pixel value based on the motion vector and the reference frame and adds the prediction residual to the reference frame through motion compensation to reconstruct the current frame. After the above steps, the frame data has been obtained, but it should be noted that the color space of the frame after direct decoding is usually in YUV420 format (components of brightness Y and chrominance U / V), and it needs to be converted to RGB (Red, Green, Blue) format before display.

[0057] 5. Rendering refers to the use of a graphics processing unit (GPU) to convert the decoded raw frame data (YUV or RGB) into an image that can be directly presented on the screen. The core steps of rendering are mainly frame data loading, color space conversion, and texture uploading. After obtaining the decoded frame data from the decoder, it is determined whether the frame data format needs further processing (such as YUV to RGB). In this application, color space conversion is performed to convert it into RGB format to adapt to the display device. The formulas for the corresponding component color space conversion are shown in Equations 1 to 3:

[0058] R=Y+1.402(V-128) (1)

[0059] G=Y-0.344(U-128)-0.714(V-128) (2)

[0060] B=Y+1.722(U-128) (3)

[0061] The conversion can be calculated on the CPU and completed using GPU shaders to improve efficiency. After the color space conversion, texture loading is the main step, which is to load the frame data into the texture memory of the GPU. Before rendering, the texture is bound to the appropriate texture unit to make it available during the rendering process. After loading, the image is drawn, and a rectangle is drawn using OpenGL as the display area. The display window size can be manually adjusted in real time and adapted to the height and width of the frame. Then create a vertex buffer to define the position of the four corners of the screen. Configure the texture coordinates to ensure that the texture data is correctly mapped to the rectangle. Use vertex shaders and fragment shaders to process vertex and texture data. Call the drawing function to draw the image to the frame buffer to form the final image. You only need to update the texture data and call the drawing function in the main loop, and exchange the buffer display results to display continuously. Regarding the texture cache, the texture pool reuses the allocated texture to reduce the overhead of memory allocation and release. At the same time, the display buffer and the rendering buffer are used alternately to avoid flickering problems during display. To avoid screen tearing, the buffer adopts a double buffer mechanism, one buffer (Front Buffer) displays the current frame, and the other buffer (Back Buffer) loads the next frame image, and switches the buffer when the screen is refreshed. At the same time, the displayed frame is synchronized with the timestamp of the video stream to avoid screen freeze or delay. The final rendering result is output to the screen by the GPU, and the entire display process is completed.

[0062] In addition, a common screen projection method is to use Bitmap transmission. Bitmap saves the original pixel data, which can maintain the clarity of the picture and will not cause the picture quality to deteriorate due to the compression algorithm. At the same time, because the bitmap format directly saves the pixel value, no complex calculation is required during decoding. After being transmitted to the receiving end, GPU rendering can be used directly, thereby reducing the delay of encoding and decoding. However, the high data volume of Bitmap requires high network bandwidth. For example, a frame of 2400x1080 resolution 24-bit color Bitmap is about 7.7MB in size. When the network bandwidth is sufficient, the screen projection effect is good; but when the bandwidth is insufficient, the resolution and frame rate need to be reduced (such as to about 15 frames) to adapt to network conditions. If the content with small changes (such as PPT) is played, a certain degree of fluency and clarity can still be guaranteed, but for high-dynamic content such as video playback or game live broadcast, the user experience will be significantly reduced.

[0063] Compared with Bitmap bitmap transmission, the algorithm proposed in this application is less dependent on bandwidth. By optimizing video encoding and transmission methods, it can still ensure the clarity and smoothness of the screen projection of high-dynamic content (such as video playback or game live broadcast) even in poor network conditions and insufficient bandwidth. This advantage makes it particularly outstanding in scenarios with large fluctuations in network conditions, which can significantly improve the user experience while saving bandwidth resources for other network tasks (such as downloading or multi-device use). The delay performance of the algorithm in this application for projecting the same content (playing video) under different conditions is shown in Table 4:

[0064] Table 4 Delay performance information of screen projection in different situations

[0065]

[0066] The specifications, standards and performance of wireless network cards (Wi-Fi modules) and hardware configurations (such as processors) of different mobile phones may vary. Therefore, even if the network speed test is performed in the same network environment, the speed test results of different devices may still be different. As shown in Table 4, the screen projection delay fluctuates within a certain range under different conditions (such as network environment, device performance, and screen content complexity). This fluctuation is mainly because when the amount of information in the projection screen changes greatly, the amount of data encoding and transmission will increase significantly, which will have a certain impact on the delay measurement. In addition, in high frame rate scenarios, frame-by-frame delay calculation requires processing a large amount of data per second. If all frame-by-frame calculation results are retained, the memory usage will increase significantly with the running time, which may affect the projection effect and indirectly cause measurement errors. At the same time, the fluctuation of individual frame delays may also have a greater impact on the test results. To reduce these effects, the average delay per hundred frames is used as the measurement standard in the test. This method can not only smooth the abnormal fluctuations of individual frame delays and reduce measurement errors, but also reduce the pressure on memory for frame-by-frame storage and calculation, and ensure the stability of the projection effect. In addition, it can be seen from Table 4 that when the network conditions are good, that is, when all are in WiFi conditions, the delays of the three test devices are all in the range of about 20ms, which is usually very low in screen projection applications. For most daily screen projection applications (such as video playback, application interface sharing, etc.), such delays are almost imperceptible. For applications that require instant feedback (such as game live broadcasts, video conferencing, etc.), a delay of 20ms is still very ideal, users will hardly notice any delay, and the interactive experience is very smooth. Users will not feel that the picture is lagging, and smooth playback and real-time interaction can be well guaranteed. Even in poor network conditions, the delay is around 90ms, which may start to cause a slight perception of delay in games or scenarios that require higher real-time performance. Most users can still accept it, especially in video playback. This delay usually does not cause obvious freezes or lags.

[0067] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0068] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0069] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration, characterized in that: The steps include: S1. Perform connection management and signaling interaction; S2. The sender captures the data for encoding. Android uses the MediaProjection API, and iOS uses ReplayKit to achieve dynamic capture of screen content. S3. The collected multi-source data is synchronized according to the timestamp and transmitted to the encoder, and pre-processed before transmission to ensure the integrity, real-time and stability of data capture; S4. Send the packaged data packet according to the designed transmission protocol. After receiving the data packet, the receiving end parses and restores the data packet according to the protocol to obtain the data format information, create and initialize the decoder parameters, and then sends a prompt to the sending end that the preparation is ready to receive the video stream data; S5. The video stream data in the data transmission process is transmitted via TCP. After the receiving end receives the video stream data, in the process of rendering and displaying the frame data, OpenGL is used to utilize GPU for hardware acceleration to display the decoded frame on the screen in real time.

2. The method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration according to claim 1 is characterized in that: In step S1, UDP is combined for connection management and signaling interaction, and the UDP protocol is used to send data packets to all devices in the same local area network. First, the client initiates a connection and sends an initial signaling packet containing a device identifier and a session ID to notify the server to establish a TCP connection. The server then responds. After receiving the connection request, the server returns a confirmation signaling packet containing confirmation information and connection parameters.

3. The method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration according to claim 1 is characterized in that: On the Android platform, MediaProjection API is an official screen content capture mechanism. The captured content contains sensitive information and requires explicit user authorization before it can be used for real-time screen recording or screen sharing. On the iOS platform, ReplayKit is an official screen recording and streaming media sharing API provided by Apple. It supports devices starting from iOS 9. When using ReplayKit, the system will automatically pop up a user authorization prompt, and the user must agree to record the screen.

4. The method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration according to claim 1 is characterized in that: In step S3, pre-processing is performed before transmission, including resolution adjustment, frame rate control and bit rate setting.

5. The method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration according to claim 1 is characterized in that: In step S3, specifically: synchronization of multi-source data is a key operation based on timestamp alignment. When processing screen capture video, the timestamp records each data stream, and a timestamp is attached during acquisition to mark the moment of data acquisition. The timestamp format is an offset relative to the start time. A queue buffer mechanism is used to cache the acquired data of each data source separately, align each data source according to the timestamp, and remove advanced or lagging data frames to ensure that multi-source data can be transmitted synchronously to the encoder. The encoder is responsible for converting the original multimedia data into a specific compression format for transmission, storage and decoding. If some data frames are lost, they are supplemented by repeating the previous frame.

6. The method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration according to claim 5 is characterized in that: Before passing the synchronized data to the encoder, preprocessing is required to ensure transmission efficiency and streaming quality. Different devices have different support capabilities for video resolution. The resolution needs to be adjusted according to the target device to adapt to its decoding capabilities. Use the image processing library to scale the resolution of the video frames. When the network bandwidth is limited, reduce the resolution and frame rate to reduce the amount of data. Dynamically adjust the bit rate according to the network conditions to ensure the stability of the transmission. Use thread-safe queues to cache collected data in a multi-threaded environment.

7. The method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration according to claim 6 is characterized in that: After receiving the data packet, the receiver parses and restores the data body, obtains the data format information, creates and initializes the decoder parameters, where the decoder is used to restore the encoded compressed video stream into playable video data. The decoder is used to perform reverse processing on the compressed data generated by the encoder to restore the original picture, and then sends feedback to the sender that the decoding preparation is ready and waiting to receive the video stream data.

8. The method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration according to claim 7 is characterized in that: After receiving the video stream data, the frame data is obtained after being decoded by the decoder, and then rendered and displayed, wherein rendering refers to the process of converting the decoded image data into a format that can be displayed on the screen.

9. The method for low-delay transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration according to claim 8, characterized in that: The key points of rendering include: color space conversion of decoded frame data, converting the decoded YUV format image into RGB format, adapting to screen display, and using OpenGL to utilize GPU for hardware acceleration during the rendering process.

10. The method for low-latency transmission and real-time decoding and rendering of video streams based on TCP / UDP collaboration according to claim 8, characterized in that: Display is the process of presenting the rendered frames to the screen. The key points include: presenting each frame of the image to the screen in chronological order according to the screen refresh rate, and then the GPU draws the rendered texture to the display buffer. The buffer adopts a double buffer mechanism, one buffer displays the current frame, and the other buffer loads the next frame of the image. The buffer is switched when the screen is refreshed, and the displayed frame is synchronized with the timestamp of the video stream. The final rendering result is output to the screen by the GPU.

Citation Information

Cited By

  • Video information processing method and system

    CN120583200A

  • Screen recording and splicing method based on multi-region frame selection

    CN121165986A

  • DRM-based striped video display method, system and device

    CN121239805A

  • Video transmission method and device, equipment, storage medium and program product

    CN122137971A

  • Digital human interaction system, method and device, storage medium and program product

    CN122199762A