Audio and video monitoring method in IP video production and broadcasting system

By designing a blocking queue-based software device on a general-performance general computer, real-time analysis and monitoring of IP video streams is realized with low latency and low packet loss rate, and the problem of difficulty in localization and control in the existing technology is solved, and efficient and low-cost video stream monitoring effect is achieved.

CN112383771BActive Publication Date: 2025-05-23COMMUNICATION UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011256101.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-11
Publication Date
2025-05-23
Estimated Expiration
2040-11-11

AI Technical Summary

Technical Problem

The prior art is difficult to realize real-time analysis and monitoring of IP video streams with low latency and low packet loss rate on ordinary general-purpose computers, and foreign professional equipment all-in-one machines are expensive and are not convenient for domestic production and control.

Method used

A software device based on running a general computer is designed, using blocking queues as real-time data cache device, and combining the efficient form of multi-threaded concurrency to realize real-time analysis and monitoring of IP video streams. The system adapts to different video formats through an autonomously designed queue length design algorithm to achieve low latency and low packet loss rate video stream analysis.

Benefits of technology

It realizes IP video stream analysis and monitoring with low cost, low latency and low packet loss rate at the software level, breaks the limitations of foreign professional equipment all-in-one machines, has the advantages of domestic production and control, and significantly improves frame loss rate, packet loss rate and end-to-end delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112383771B_ABST
    Figure CN112383771B_ABST
Patent Text Reader

Abstract

The present invention provides an audio and video monitoring method in an IP video production and broadcasting system, which can break the limitations of foreign professional equipment integrated machines and realize control changes at the software level that are convenient for our research and data packet level. The invention adopts an efficient form of acquisition caching, parsing, and multi-threaded playback concurrency, and realizes real-time parsing and monitoring of IP video streams through five main steps: video data stream capture, calculation of initial queue lengths with adaptability to different video formats, blocking queue data caching, video data processing, and video monitoring and display. The present invention can cache video data according to blocking queues with different queue lengths of video format adaptability, as well as accurate decapsulation and data processing for video encapsulation standards, to achieve real-time video parsing and monitoring with low latency and low packet loss rate. The method is applicable to video formats of different resolutions, such as standard definition, high definition, and ultra-high definition, under the SMPTE ST 2022‑6 or SMPTE ST 2110‑20 standards.
Need to check novelty before this filing date? Find Prior Art

Description

Technical field:

[0001] The present invention mainly relates to baseband video stream analysis and monitoring in the process of IP-based media video production, and mainly considers the analysis and monitoring method of an IP baseband video stream analysis system and the device thereof.

[0002] IP baseband video stream analysis and monitoring technology is a key technology in the full IP video production. At present, only foreign manufacturers have mastered this technical method, using professional equipment all-in-one machine to monitor IP layer signal transmission, analysis accuracy and system delay, and observe the video analysis effect in real time. This professional all-in-one machine has a very high price, and is based on foreign TV media system standards for analysis and monitoring. It uses hardware boards to burn fixed programs, which is not convenient for our research and control changes at the data packet level. The standards of the SMPTE ST 2022 and SMPTE ST 2110 series protocols not only standardize the encapsulation of professional videos into IP data packet formats, but also provide methods for transmitting high-quality video signals via IP. Among them, SMPTE 2022-6 defines the encapsulation of SDI format, based on the real-time transport protocol (RTP) and the high bit rate media transport protocol (HBRMT). HBRMT is the abbreviation of high bit rate transport stream. The high bit rate here is to distinguish it from other transmission compression signals. Each IP data packet contains 1376 bytes of SDI data. In order to make the IP packets of each frame neatly arranged, the last packet of each frame is filled with zero bytes. The SMPTE ST 2110 series of standards puts the video, audio and auxiliary data parts of the signal into different streams and sends them separately, so that the receiver can directly obtain the streams they want separately. Among them, ST 2110-20 specifies the basic encapsulation and transmission format of real-time, RTP-based uncompressed active video on IP networks. The difference between ST 2110-20 and ST 2022-6 in transmitting video streams is that a special SDP (Session Description Protocol)-based signaling method is defined for receiving and interpreting the image technology metadata required for the stream. The main content transmitted includes the necessary media type parameters such as sampling structure, bit depth, frame rate, and media parameter types with default values ​​such as interlaced scanning, progressive segmentation, TCS, etc. The SMPTE ST 2022 and SMPTE ST 2110 standards make it possible to flexibly transmit uncompressed video, audio and auxiliary data through IP networks based on existing broadcasting and television.

[0003] Our research team has designed and developed a set of software devices with independent intellectual property rights that run on general-purpose computers with ordinary performance, realize the controllability of data caching and processing mechanisms, and realize the analysis and system monitoring of IP baseband video streams with low threshold and low cost. The following will propose the implementation method and monitoring indicators of IP video analysis, and use our system as an example to introduce the actual use of the device and data results. Summary of the invention:

[0004] The purpose of the present invention is to provide a new method for implementing real-time analysis and monitoring of IP video streams based on a software device running on a general-purpose computer, so as to solve the problems of transparent and controllable data processing of IP video systems and localization and generalization of video stream analysis technology.

[0005] To implement the analysis and monitoring of IP video, several complex operations are required, including video data stream capture, data storage, video data processing (including decapsulation, pixel extraction, and cache pixel processing), video monitoring, and display. Video data stream capture mainly relies on WinPcap, which can store data in local PCAP files and read local PCAP files in real time for data processing. Video data processing relies on encapsulation protocols such as ST 2022-6 and ST 2110. Video monitoring and display are implemented through SDL. However, these can only implement the analysis and playback of high-latency and high-packet-loss-rate videos in software, and the effect is poor.

[0006] Therefore, the realization of a low-delay, low-packet-loss-rate real-time IP video stream parsing and monitoring based on a software device running on a general-purpose computer is a major focus and highlight of the present invention. It requires a deeper level of complex operations on data storage, that is, a device that can cache video stream data packets in real time is required to ensure the real-time input and output of IP video streams, and the device can automatically determine whether it is a video stream data packet. The device can break the limitations of foreign professional equipment all-in-one machines and realize the convenience of our research and control changes at the data packet level at the software level.

[0007] To this end, the present invention independently designs a blocking queue with an initial queue length adaptable to different video formats as a real-time data caching device, and adopts an efficient form of multi-threaded concurrent acquisition, analysis, and playback. The blocking queue caches data to solve the rate matching problem between the three threads, and the device is used to realize real-time analysis and monitoring of IP video streams in software. The queue length design algorithm independently designed by the present invention obtains blocking queues with different initial queue lengths according to different video formats, such as clarity, scanning format, etc., and has strong adaptability to video formats. The algorithm is suitable for video formats with different clarity, such as high definition and ultra-high definition, but only has a line-by-line scanning format in the ultra-high definition format. By accurately connecting the video data stream capture and video data processing to the input and output ends of the blocking queue, respectively, and closely coordinating with video monitoring and display, the analysis and playback of IP video streams are realized, and the quality of real-time video playback is guaranteed.

[0008] At the level of video stream data processing and parsing, the present invention, based on compliance with the ST 2022-6 and ST 2110-20 standards, decapsulates the protocol header through an independently developed program design, obtains the video information in the protocol header, judges, filters, and arranges the valid video frame pixel information in the video data, splices the video frame image information, and finally transmits the video information frame by frame to the player.

[0009] To measure the performance of the ultra-high-definition IP video acquisition, analysis and monitoring system, we can propose the following five indicators based on data integrity, transmission stability and real-time performance, namely frame loss rate, packet loss rate, frame interval distribution (inter-frame jitter) / packet interval jitter, first packet processing delay, and end-to-end delay. The addition of the blocking queue and queue length design algorithm independently designed by the present invention can achieve real-time effects while also greatly improving the frame loss rate, packet loss rate, and end-to-end delay.

[0010] The following are the various operation implementation methods and principles of the present invention:

[0011] 1. Video data stream capture

[0012] Efficiently monitoring data packets in the network is the first step in parsing data. We use WinPcap, an architecture for capturing network traffic and performing network analysis, to intercept underlying data packets and complete filtering under the Windows system.

[0013] (1) First, we determine the target interface to be sniffed and use file handles to identify different interfaces to distinguish different sniffing tasks.

[0014] (2) Next, create a new sniffing task, assign the file handle of the target network card device to the variable, and set the maximum number of bytes of data to be captured. Unless an error occurs, the sniffing task will continue to execute, and the error should also be saved using the errbuf string to facilitate our problem location and analysis.

[0015] (3) Secondly, filter the traffic. Generally, the sniffing task only captures specific traffic, so filters can be used. Although a sufficient number of if / else conditional statements can achieve the same effect, filters are still the best choice because WinPcap filters directly call BPF filters, which is very efficient; in addition, the BPF driver can directly perform many operations.

[0016] (4) Capture one packet at a time, and call the callback function every time a data packet is sniffed to process and store the data packet accordingly.

[0017] 2. Data Storage

[0018] In order to facilitate data communication between multiple threads, a blocking queue is used for intermediate data storage. The characteristics of a blocking queue are: when the queue is full, it will wait for the data in the queue to be consumed, and when the queue is empty, it will wait for data to be filled. Through a shared blocking queue, the data entrance and exit of the queue are fixed, and the data enters from the entrance and is taken out from the exit. At the same time, the first-in-first-out (FIFO) mode is adopted to ensure that the order in which the data in the queue is taken out is consistent with the order in which it enters. In a multi-threaded programming environment, a blocking queue is a very efficient way to achieve thread communication. In addition, a blocking queue can also solve the problem of speed mismatch between threads that generate data and threads that consume data within a certain period of time. If the speed of data generation is greater than the speed of consumption, and it lasts for a period of time, so that the accumulated data size exceeds the length of the blocking queue, then the thread that generates data will go into sleep to wait for the data consumption thread to process the accumulated data; on the contrary, if the speed of data consumption is greater than the speed of generation, so that there is no accumulated data in the blocking queue, then the thread that consumes data will go into sleep to wait for data production.

[0019] 3. Queue Length Design Algorithm

[0020] The frame loss rate, packet loss rate, and end-to-end delay in measuring the performance of ultra-high-definition IP video acquisition, analysis, and monitoring systems are closely related to the length of the blocking queue. The longer the queue, the smaller the average packet loss rate and frame loss rate, and the longer the delay; the shorter the queue, the larger the average packet loss rate and frame loss rate, and the shorter the delay. Therefore, when the system is sensitive to packet loss but not to real-time performance, a queue unlimited storage solution can be used. Compared with the real-time write-read PCAP file data sharing solution (that is, storing video stream data in local PCAP files instead of in the blocking queue cache), this solution can achieve zero frame loss rate and zero packet loss rate; when the system is not sensitive to packet loss but has high real-time performance requirements, a queue limited storage solution should be used. Compared with the real-time write-read PCAP file data sharing solution, this solution can achieve lower end-to-end delay.

[0021] In view of the limited storage length scheme of the queue, the present invention independently designs an algorithm to design the corresponding queue length according to the input and output rates and changes of different video stream formats under network conditions to obtain low packet loss rate and low delay.

[0022] It should be emphasized again that the algorithm is applicable to video formats of different resolutions, such as HD and UHD, but in the UHD format there is only a progressive scan format.

[0023] The input rate λ of the blocking queue refers to the rate of video data packets collected per second and arriving at the queue after being transmitted through the network, which can be obtained by the computer directly dividing the number of data packets n received in a period of time by the corresponding time t; the output rate μ refers to the rate of parsing and playing data packets per second during video playback, which can be obtained by the frame rate (F) of the video playback * the number of data packets per frame. The number of data packets per frame is calculated in two cases: in the first case, the number of data packets per frame of the video encapsulated by the ST 2022-6 standard (including progressive scanning and interlaced scanning) and the progressive scanning video encapsulated by the ST 2110-20 standard is equal to the difference between the data packet sequence numbers Seq_num2 and Seq_num1 with the flag bit M in two consecutive frames. The flag bit M represents that this data packet is the last data packet of this frame. In the second case, the flag bit M in the interlaced scanning video encapsulated by the ST 2110-20 standard indicates the last data packet of a field, so the number of data packets of a frame of the interlaced scanning needs to be multiplied by 2 by the number of data packets of a field, and the video in the ultra-high-definition format does not have an interlaced scanning format. The performance of the monitoring computer we designed is high enough, and the parsing and playback of the video stream after entering the program is stable, that is, the output rate of the queue is stable. The acquisition rate of the video stream is constant, but in the process of IP network transmission, there are factors such as network jitter, so the input rate of the queue is variable, and it is less than or equal to the output rate under long-term steady-state conditions. In actual network conditions, the network status changes are complex, and the queue length needs to be adaptively adjusted according to the queue length algorithm to ensure a good packet loss rate and delay.

[0024] When the input rate is lower than the output rate, the queue length L1 is equal to the video stream output rate μ minus the input rate λ multiplied by the real-time video duration:

[0025] L1=(μ-λ)*duration (1)

[0026] First case:

[0027] L1_first=[(Seq_num2-Seq_num1)*Fn / t]*duration (2)

[0028] Second case:

[0029] L1_second=[(Seq_num2-Seq_num1)*F*2-n / t]*duration (3)

[0030] When the input rate is approximately equal to the output rate, the queue length L2 is equal to the input jitter input_jitter plus the initial parsing and playback delay delay_len. Under the premise of PTP synchronization, the input jitter of the network can be represented by the difference between the RTP timestamp carried in the data packet and the actual time of the received packet. The input jitter of the network refers to the degree of change of the packet delay, which is random. Therefore, this algorithm adopts a maximum jitter delay processing, which is specifically implemented by measuring and predicting the jitter of the time interval of the arriving data packet to obtain the maximum jitter queue length input_jitter. Or fit a period of video data history input to multiple mathematical models, find the arrival mathematical model with the highest fitting degree, and find the maximum jitter queue length input_jitter. The initial parsing and playback delay refers to the queue length required to cache a part of the data first to prevent network jitter when the video stream arrives, and then parse and play the video to ensure smooth video output. In this algorithm, the initial analysis and playback delay is optimized according to the network jitter and the real-time requirements of playback. A basic method is to design the initial analysis and playback delay as the queue length delay_len for caching a frame of data packets in the corresponding video format, that is, the difference between the sequence numbers Seq_num2 and Seq_num1 of the data packets with the flag bit M in two consecutive frames (*2 is required for interlaced video encapsulated in ST 2110-20, and ultra-high definition does not include interlaced scanning). This basically ensures that the queue will not be cleared during playback.

[0031] L2=input_jitter+delay_len (4)

[0032] First case:

[0033] L2_first=input_jitter+(Seq_num2-Seq_num1) (5)

[0034] Second case:

[0035] L2_second=input_jitter+(Seq_num2-Seq_num1)*2 (6)

[0036] The pseudo code is as follows:

[0037]

[0038]

[0039] 4. Data Processing

[0040] The SDI data format is explained by taking the virtual interface of 3G-SDI source data consisting of two 10-bit parallel data streams as an example. The virtual interface complies with the SMPTE 425M standard. The virtual interface divides the data of each line of data stream 1 and 2 into four areas: EAV (end of active video) clock reference, digital blanking area, SAV (start of active video) clock reference, and digital active line. The format is as follows: Figure 1 As shown in the figure, both EAV and SAV are clock reference signals (TRS). The TRS at the start line of each video frame and the data indicating the line number afterwards are fixed, so they can be used to search for the starting point data packet of a frame of image.

[0041] The two parallel data streams of the virtual interface are transmitted on a single channel in bit-serial form after multiplexing, parallel-to-serial conversion and scrambling have been applied. Data stream 1 and data stream 2 of the virtual interface shall be multiplexed word-by-word into a single 10-bit parallel stream in the following order: data stream 2, data stream 1, data stream 2 and so on, as Figure 2 As shown, the order of brightness and chromaticity of corresponding pixels in the image is Cb0, Y0, Cr0, Y1, Cb1, Y2...

[0042] The data packets captured from Ethernet are encapsulated according to the SMPTE ST 2022-6 standard or the ST 2110 series standards. The order of packaging from outside to inside is: Ethernet, IP, UDP, RTP, and the innermost is uncompressed video data, audio, and auxiliary data. Among them, the innermost encapsulation in the ST 2022-6 standard is the SDI data format described above, namely HBRMT (High Bit Rate Media Transport Protocol), while in the ST 2110 series standards, video, audio, and auxiliary data are separately encapsulated and transmitted according to the corresponding standards. For example, in the ST 2110-20 standard, the innermost data only encapsulates the video content, that is, the digital valid line content in the SDI data format.

[0043] After obtaining the data packet, according to the SMPTE ST2022-6 standard (or ST2110 series standard), verify whether the data packet format complies with the standard requirements, use the corresponding position value of the data packet to assign each field and array in the structure, complete the transmission of video data parameters, including video width and height, frame rate, chroma sampling, interlaced / progressive parameters, and create a new image buffer area of ​​corresponding byte size according to the width and height. In the ST 2022-6 standard, the data packet of the image start line is obtained by judging the TRS in the SDI data format, and then the data packet belonging to the same frame is read. Then, according to the digital blanking information encapsulation format in SAV, the auxiliary data, audio data, and other data are stripped off, leaving only the image data of the digital valid line. In the ST 2110-20 standard, it can be directly based on Figure 6The encapsulation standard of the RTP payload header obtains the first packet of the video frame and all data packets of the same frame by judging the length, number of lines, and offset. Then, the original sampled pixel value tuple is restored by splicing adjacent bytes, and then the original tuple pixel data is processed in combination with the characteristics of computer display images to facilitate image display. Finally, the Y, Cb, and Cr brightness and chromaticity of the corresponding pixels of the image are assigned until the image buffer is completely filled, and we get a complete image. Repeat the data processing process to continuously update the image in the image buffer.

[0044] 5. Video Monitoring and Display

[0045] Our monitoring system uses SDL (Simple DirectMedia Layer), which is an open source cross-platform multimedia development library written in C language. It provides a wealth of functions for controlling images, sounds, input and output. Other function libraries, especially those that support 10-bit or more HDR, can be used for video monitoring display.

[0046] The process of SDL video display can be divided into two parts: initialization and loop display. The initialization process can be further divided into four steps: initializing SDL, creating a window (Window), creating a renderer (Render) based on the window, and creating a texture (Texture). The loop display includes setting texture data, copying the texture to the rendering target, and displaying.

[0047] SDL initialization requires the use of the SDL_Init() function to determine the subsystem you want to activate. You can choose timers, audio, video, joysticks, touch screens, game controllers, events, or all subsystems. Use the SDL_CreateWindow function to create a video playback window (Window), specifying the window title, window position coordinates, width and height. Use the SDL_CreateRender() function to create a renderer based on the window. The parameters need to specify the target window to be rendered, the rendering device used for initialization, hardware / software rendering, and synchronized display refresh rate. If the creation is successful, the renderer ID will be returned. Then you need to use SDL_CreateTexture to create a texture based on the renderer. You need to specify the format of the texture. Each format includes the following properties: the storage method of pixel components, whether the pixel components are stored together or separately; the storage order of pixel components, that is, the big endian and small endian issues. For a small endian system like Windows, the storage order of the "ARGB" format in memory is B, G, R, A; the number of bits occupied by each component; the number of bits and bytes occupied by each pixel. Use SDL_UpdateTexture() to set the pixel data of the texture, SDL_RenderCopy() to copy the texture data to the rendering target, and when the video is played, the next frame will completely cover the previous frame. Finally, use SDL_RenderPresent() to display the picture. Description of the drawings:

[0048] Figure 1 : Parallel data stream format

[0049] Figure 2 : 10bit serial data stream

[0050] Figure 3 :System physical topology and data logic diagram solution

[0051] Figure 4 : Video display screen

[0052] Figure 5 : RTP payload header (SMPTE ST2022-6)

[0053] Figure 6 : RTP payload header (SMPTE ST2110-20)

[0054] Figure 7 : A frame of the 2022-6 standard encapsulated video frame tail data packet

[0055] Figure 8 : Comparison of three solutions for frame loss rate

[0056] Fig. 9: Comparison of three solutions for intra-frame packet loss rate

[0057] Fig.10 : Comparison of three end-to-end delay solutions Specific implementation method:

[0058] The following further describes the details of the IP baseband video stream analysis and monitoring method and device in conjunction with the accompanying drawings.

[0059] 1. Physical topology of the experimental system

[0060] A complete IP video stream acquisition and analysis system consists of three parts: video signal source (IP gateway), transmission and switching network, and signal acquisition and analysis terminal.

[0061] The physical topology of the experimental system is constructed as follows: Figure 3 The physical topology shown in the figure. The devices used in this experimental topology mainly include: HD camera (SDI format output), SMPTE 2022-6 / ST 2110 gateway, Intel 82599ES 10G network card, and general-purpose computer.

[0062] The camera outputs the captured high-definition video in real time through the SDI coaxial cable and connects to the IP gateway. According to the SMPTE ST 2022-6 / ST 2110 standard specification, the IP gateway performs packet segmentation, adds protocol headers and other encapsulation processing on the received signal, thereby converting it into an IP stream signal that can be transmitted in the packet switching network. Combined with other switches and routing nodes, uncompressed baseband video can be flexibly transmitted in a very large network, thus breaking through the limitations of traditional coaxial cable transmission, such as complex wiring and single-stream transmission. The IP gateway outputs through the 10 Gigabit network port and transmits through optical fiber, and is connected to the terminal 10 Gigabit network card port, so that the data packet can be captured by us.

[0063] 2. Experimental Framework

[0064] The experimental design of this product, based on a full analysis of IP protocol specifications such as the SMPTE ST 2022-6 / ST 2110 standard, implements IP video signal acquisition through a 10G (10 Gigabit) network card and WinPcap, completes IP signal analysis and video display in the VS2015 development environment, and realizes the functions of expensive commercial off-the-shelf hardware terminals on the market at a low cost, with controllable data processing.

[0065] In order to more clearly reflect the superiority of the IP baseband video stream parsing and monitoring method and device of the present invention, in terms of experimental framework, this section introduces two experimental frameworks for realizing real-time video parsing and playback, namely, the initial scheme of sharing data through real-time writing and reading PCAP files and the improved blocking queue data sharing scheme independently designed and developed by us.

[0066] The video content in this manual takes the HD standard 1920*1080 60i, frame rate 30Hz, 4:2:2 sampling structure, 10bit sampling, ST 2022-6 encapsulation standard as an example, and the size of each Ethernet data packet is 1442 bytes. When the video is in 4K / 8K ultra-high-definition format, the hardware performance requirements of the device will be higher, but the monitoring method can be the same.

[0067] 1. Monitoring plan

[0068] like Figure 3 As shown, the acquisition and analysis functions of the system terminal can be logically divided into three parallel tasks. Considering the actual application, we choose to implement multi-threaded concurrency. Thread 1 is the acquisition thread, and thread 2 and thread 3 are responsible for data analysis and video playback respectively. The processing within each thread is introduced in detail below.

[0069] Thread 1 uses WinPcap to sniff the network card to capture data packets. First, create a network device list to obtain the adapter of the local machine and obtain detailed information of the adapter (including name, mask, source / destination address, broadcast address). Type the number of the 10G network card adapter from the command line, open the adapter and start capturing data. Specify the capture length as 65535 (greater than the maximum MTU) to ensure that the complete packet is obtained, and set the promiscuous mode to ensure that no data packets are missed. Use the PCAP_loop function and the callback function to continuously capture data packets. In the initial scheme, open a heap file in a local disk (if the path does not exist, a new heap file will be created). Use the PCAP_dump function to write each captured data packet into the specified heap file for storage in the format of a PCAP file. At this point, thread 1 has completed the mission of real-time data collection and storage in a file. The PCAP file consists of a file header, a data packet header, a data packet, a data packet header, a data packet, a data packet header, a data packet, etc. The PCAP file has only one file header, which contains 7 fields, and the packet header contains 4 fields. The data transmitted in the actual network is located in the data packet after the packet header. In actual parsing, it takes time to parse the redundant packet header. When each data packet is written to the PCAP file separately, hundreds of thousands of data packets are written per second. Writing to the disk itself is a relatively time-consuming operation, and writing a small amount of data multiple times takes longer than writing a large amount of data at one time. Therefore, the overall data processing of the initial scheme system is quite time-consuming. In the improved scheme designed and developed independently, for real-time monitoring, the PCAP_setbuffer and PCAP_setuserbuffer functions are used to adjust the size of the kernel buffer and the user buffer respectively, both set to 10MB, and the buffer area is expanded to solve the problem of packet loss during packet capture. And using blocking queues to access data instead of file reading and writing to achieve inter-thread communication, reducing the time for writing data to the disk and parsing the packet header, greatly improving the speed of data reading and writing, and improving the convenience of data sharing.

[0070] Since the speed of thread 1 does not match that of threads 2 and 3, the length of the blocking queue as a cache is also a very important variable here. We set the length to be infinite and to a queue length that can adapt to different video formats, and verify the difference in their effects through experiments.

[0071] In this example, the high-definition video format of the 2022-6 standard package is 60i and the frame rate is 30 Hz. This is used as an example to describe the calculation method of the blocking queue length.

[0072] 1) When the input rate is less than the output rate, the formula for calculating the finite queue length is:

[0073] L=[(Seq_num2-Seq_num1)*Fn / t]*duration

[0074] For easier observation and understanding, use Wireshark to capture real-time IP video data packets to calculate the queue length. Figure 7 The captured data shown has been filtered and only displays the data packets whose flag bit Mark is True. Figure 7 All the data packets between two adjacent data packets with the flag bit M and the second data packet with the flag bit M constitute a frame of data packets (excluding Figure 7 The first data packet with a flag bit in the frame is the last data packet of the previous frame), and the number of all data packets in a frame is equal to the Seq difference: Seq_num2-Seq_num1=4691-194=4497.

[0075] The input rate can be determined by Figure 7 The sequence number and timestamp of the first and last data packets in the green box are as follows: (159902-146411) / (1.254768-1.154743) = 134876. In practice, more data packets should be taken to find the average value of the input rate.

[0076] The video duration is set to 10 minutes, that is, 600 seconds, so the actual value of the video queue length in this example is: [4497*30-134876]*600=20400 data packets.

[0077] 2) When the input rate is approximately equal to the output rate, the calculation formula for the finite queue length is:

[0078] L=input_jitter+(Seq_num2-Seq_num1)

[0079] Since there are many mathematical models for network jitter, we will not perform specific model calculations here. Instead, we will use the statistical observation method to obtain the actual length of the maximum jitter queue, input_jitter. First, set the queue length to infinite, observe the changes in the number of cached packets in the queue, record the maximum number of packets, and remove individual burst traffic. This value is input_jitter, assuming the size is 20,000 (this is just an example, the actual value should be obtained based on the actual situation). The size of the video frame is given by Figure 7 The difference between the two adjacent data packets Seq_num is: (Seq_num2-Seq_num1)=(4691-194)=4497

[0080] Therefore, the queue length of this video is: 20000+4497=24497

[0081] After completing the initialization of each data structure, thread 2 opens the cache queue or the PCAP file storing the captured data in read-only mode, reads the first data packet header, and reads the data after the data packet header with a length of caplen, thus obtaining the data of the first data packet. According to the SMPTE ST2022-6 / ST2110 standard, verify whether the data packet format complies with the standard, assign the values ​​of the corresponding positions of the data packet to each field and array in the structure, and complete the transmission of the video width and height, frame rate, chroma sampling, and interlaced / progressive parameters. The fields where the parameters are located are as follows: Figure 5 Figure 6 As shown. According to the resolution and chroma sampling of the video data, a new image buffer of 1920*1080*2 bytes is created. Then continue to read the second data packet header and data packet from the file, and obtain the clock reference signal (TRS) according to the SDI payload data format in the inner layer of the data packet to search for the starting point data packet of a frame of image. According to TRS, determine whether the second packet is the image start line. If not, continue to read the file until it matches. After the data packet where the image start line is located is determined, read the data packets belonging to the same frame of image in turn, restore the 10-bit original sampled pixel value tuple by splicing adjacent bytes, and then process the pixel value data in combination with the characteristics of SDL display image to facilitate image display. Finally, assign the Y, Cb, and Cr brightness and chromaticity of the corresponding pixels of the image until the image buffer is completely filled, and we get a complete frame of image. Repeat this process to continuously update the image in the image buffer.

[0082] Thread 3 calls the SDL multimedia development library, creates the SDL window, renderer, and texture in order, and then obtains the complete frame data stored in the image buffer from thread 2 as a texture parameter. The image buffer is updated in frames and the texture is also updated. Through the rendering of the renderer, the display effect of the original video is achieved. The video display screen plays smoothly without any lag. Figure 4 shown.

[0083] 2. Comparison of experimental results

[0084] In this section, we will measure the performance of the IP video stream analysis system from the three indicators of frame loss rate, intra-frame packet loss rate, and end-to-end delay. By comparing the initial solution, the queue unlimited storage solution, and the queue limited storage solution (the queue length is set to be able to store 5 frames of video data packets (4497*5=22485 packets) as an example), we can show the superiority of our self-developed improved solution.

[0085] 1. Frame loss rate

[0086] Figure 5It is the field content in the RTP payload header encapsulated by the ST 2022-6 protocol, where FRcount represents the value of the video frame counter. The counter value should be increased by 1 in the next data packet after the M flag bit of the video frame tail packet is 1, and will be flipped and counted again after 256 frames. The number of lost frames and the total number of frames are calculated by capturing and parsing the data difference of the FRcount of the previous and next video frame tail packets, thereby obtaining the frame loss rate.

[0087] Figure 8 The results of the three schemes are compared. The initial scheme is close to the queue limited storage scheme, with an average frame loss rate of about 6%. The frame loss rate of the queue unlimited storage scheme is 0 in 20 experimental tests, which is significantly improved compared with the initial scheme.

[0088] 2. Intra-frame packet loss rate

[0089] The sequence number field in the RTP header represents the lower 16 bits of the RTP sequence counter. When the sending node sends an RTP data packet at the same power each time, the sequence number should be increased by 1 to indicate the total number of RTP data packets sent. The number of lost packets and the total number of packets are calculated by capturing and parsing the data difference between the sequence numbers of the packets before and after a frame, thereby obtaining the packet loss rate.

[0090] Fig. 9 The results of the three schemes are compared. The initial scheme is close to the queue limited storage scheme, with an average frame packet loss rate of about 5.5%. The queue unlimited storage scheme has an average frame packet loss rate of 0 in 20 experimental tests, which is significantly improved compared with the initial scheme.

[0091] 3. End-to-end latency

[0092] First, network synchronization is achieved through PTP, and then the end-to-end delay is obtained through the real-time time difference between the video data source and the video playback end.

[0093] Fig.10 The results of the three solutions are compared. The average end-to-end delay of the initial solution is 12525ms; the average end-to-end delay of the queue unlimited storage solution is 11500ms, which is slightly improved compared with the initial solution; and the average end-to-end delay of the queue limited storage solution is 463ms, which is greatly improved compared with the initial solution.

[0094] From the comparison of various indicators, it is found that the infinite queue storage solution is far superior to the initial solution in terms of frame loss rate, average packet loss rate per frame, and first packet processing delay, but the end-to-end delay indicator is similar to the initial solution but significantly inferior to the limited queue storage solution. Therefore, the infinite queue storage solution can be used when the system is sensitive to packet loss but not to real-time performance.

[0095] In the experiment, the performance of the queue-limited storage solution in terms of frame loss rate and average packet loss rate is basically the same as that of the initial solution, while the end-to-end delay effect is significantly better than that of the initial solution and the queue-limited storage solution. The queue-limited storage solution in this experiment is set to be able to store 5 frames of video data packets as an example. When we use the queue length optimization algorithm to adjust the queue length, we can reduce the frame loss rate and average packet loss rate, slightly extend the end-to-end delay, and finally obtain a low packet loss rate and low delay, reflecting the superiority of our independently developed improvement solution.

Claims

1. Audio and video monitoring method in IP video production and broadcasting system, It is characterized in that The method specifically includes: Adopting the efficient form of acquisition cache, analysis, and multi-threaded concurrent playback, it realizes real-time analysis and monitoring of IP video streams through five main steps: video data stream capture, calculation of initial queue length adaptable to different video formats, blocking queue data cache, video data processing, video monitoring and display; 1) Use WinPcap to sniff the network card and capture data packets. First, create a network device list to obtain the adapter of the local machine and obtain detailed information of the adapter, including name, mask, source / destination address, broadcast address, open the adapter and start capturing video data packets. At the same time, set the maximum number of bytes of data to be captured and filter the data traffic; 2) Determine the length of the blocking queue through an adaptive queue length algorithm to ensure low latency and low packet loss rate for real-time video analysis and monitoring; 3) Cache the collected video data into a blocking queue; take it out from the blocking queue when parsing the video data; through a shared blocking queue, fix the data entrance and exit of the queue, the data enters from the entrance and is taken out from the exit, and adopts a first-in-first-out mode to ensure that the order of data in the queue when it is taken out is consistent with the order when it enters; 4) The order of video encapsulation protocols in ST 2022-6 and ST 2110 series standards from outside to inside is: Ethernet, IP, UDP, RTP, and the innermost is uncompressed video data, audio and auxiliary data; the innermost encapsulation in ST 2022-6 standard is the SDI data format described above, that is, HBRMT high bit rate media transmission protocol, while in ST 2110 series standards, video, audio and auxiliary data are separately encapsulated and transmitted according to the corresponding standards; in ST 2110-20 standard, the innermost data only encapsulates video content, that is, the digital valid line content in SDI data format; decapsulate the protocol header according to different standards, obtain the video information in the protocol header, judge, filter and arrange the valid video pixel information in the video data, perform splicing of video frame image information, and finally transmit the video pixel information frame by frame to the player; 5) The video monitoring and display system device uses the SDL multimedia development library; the process of SDL displaying video is mainly divided into two parts: initialization and loop display of the picture; the initialization process is further divided into four major steps: initializing SDL, creating a window, creating a renderer based on the window, and creating a texture; the loop display of the picture includes setting texture data, copying the texture to the rendering target, and displaying; The steps 2) and 4) are specifically as follows: Step 2): The input rate λ of the blocking queue refers to the rate of video data packets collected per second and arriving at the queue after being transmitted through the network. It is obtained by the computer directly dividing the number of data packets n received in a period of time by the corresponding time t; the output rate μ refers to the rate of parsing and playing data packets per second during video playback, which is obtained by the frame rate (F) of the video playback * the number of data packets per frame; the number of data packets per frame is calculated in two cases: in the first case, the number of data packets per frame of the progressive scan and interlaced scan video encapsulated by the ST 2022-6 standard and the progressive scan video encapsulated by the ST 2110-20 standard is equal to the difference between the data packet sequence numbers Seq_num2 and Seq_num1 with the flag bit M in two consecutive frames; the flag bit M represents that this data packet is the last data packet of this frame; in the second case, the number of data packets per frame encapsulated by the ST The flag M in the interlaced video encapsulated in the 2110-20 standard indicates the last data packet of a field. Therefore, the number of data packets in a frame of interlaced scanning needs to be multiplied by 2. The video in ultra-high-definition format does not have an interlaced scanning format. The acquisition rate of the video stream is constant, but during the transmission over the IP network, the input rate of the queue is variable and is less than or equal to the output rate under long-term steady-state conditions. When the input rate is lower than the output rate, the queue length L1 is equal to the video stream output rate μ minus the input rate λ multiplied by the real-time video duration: L1=(μ-λ)*duration (1) First case: L1_first=[(Seq_num2-Seq_num1)*Fn / t]*duration (2) Second case: L1_second=[(Seq_num2-Seq_num1)*F*2-n / t]*duration (3) When the input rate is approximately equal to the output rate, the queue length L2 is equal to the input jitter input_jitter plus the initial parsing and playback delay delay_len; under the premise of PTP synchronization, the input jitter of the network is represented by the difference between the RTP timestamp carried in the data packet and the actual time of the received packet; the input jitter of the network refers to the degree of change of the packet delay, and the change is random, so a maximum jitter delay processing is adopted, which is specifically implemented by measuring and predicting the jitter of the time interval of the arriving data packets to obtain the maximum jitter queue length input_jitter; or a period of video data history input is fitted with multiple mathematical models to find the arrival mathematical model with the highest fitting degree, and the maximum jitter queue length input_jitter is found; the initial parsing and playback delay refers to the queue length required for caching a part of the data when the video stream arrives, in order to prevent network jitter, and then parse and play the video to ensure smooth video output; the initial parsing and playback delay is designed to cache the queue length delay_len of a frame of data packets in the corresponding video format, that is, the difference between the sequence numbers Seq_num2 and Seq_num1 of the data packets with the flag bit M in two consecutive frames, ST *2 is required for interlaced video in 2110-20 package. Ultra HD does not include interlaced scanning, which ensures that the queue will not be cleared during playback. L2=input_jitter+delay_len (4) First case: L2_first=input_jitter+(Seq_num2-Seq_num1) (5) Second case: L2_ second=input_jitter+(Seq_num2-Seq_num1) *2 (6) Step 4): After obtaining the data packet, verify whether the data packet format complies with the standard according to the SMPTE ST2022-6 standard or the ST2110 series standard, assign the values ​​of the corresponding positions of the data packet to each field and array in the structure, complete the transmission of video data parameters, including video width and height, frame rate, chroma sampling, and interlaced / progressive parameters, and create a new image buffer area of ​​corresponding byte size according to the width and height; in the ST 2022-6 standard, the data packet of the image start line is obtained by judging the TRS in the SDI data format, and then the data packet belonging to the same frame is read; then, according to the digital blanking information encapsulation format in SAV, the auxiliary data, audio data, and other data are stripped, leaving only the image data of the digital valid line; in ST In the 2110-20 standard, the first packet of the video frame and all data packets of the same frame are obtained directly according to the encapsulation standard of the RTP payload header by judging the length, number of lines, and offset; then the original sampled pixel value tuple is restored by splicing adjacent bytes, and then the original tuple pixel data is processed in combination with the characteristics of computer display images to facilitate image display; finally, the Y, Cb, and Cr brightness and chromaticity of the corresponding pixels of the image are assigned until the image buffer is completely filled, and a complete image frame is obtained; the data processing process is repeated to continuously update the image in the image buffer.

Citation Information

Patent Citations

  • Buffer control method, relaying device and communication system

    CN101395871A

  • Method and apparatus for splicing

    US20060093045A1