Non-perceptual parsing and restoration method and system for RTP audio and video streams, and medium

By mirroring traffic data in the monitoring system, identifying and parsing signaling protocols and video encoding parameters, seamless RTP audio and video stream parsing is achieved, solving the video stream interference problem and ensuring system transparency and accurate format recognition.

WO2026066901A1PCT designated stage Publication Date: 2026-04-02SHANGHAI YVIEW TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

In existing technologies, video streams can easily interfere with monitoring systems when acquiring surveillance video data, especially in scenarios requiring highly concealed information acquisition, thus failing to meet usage requirements.

Method used

By monitoring system traffic data through the access module, identifying and filtering target data, parsing signaling protocols to obtain configuration information and video encoding parameters, dividing audio and video data, and restoring and synchronizing them, the system achieves seamless RTP audio and video stream parsing.

Benefits of technology

It enables accurate identification of video stream formats without relying on initial session encoding information, avoiding interference with the monitoring system, and is particularly suitable for highly covert information acquisition scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025117500_02042026_PF_FP_ABST
    Figure CN2025117500_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present invention are a non-perceptual parsing and restoration method and system for RTP audio and video streams, and a medium. The method comprises: accessing a monitoring system, and mirroring traffic data in the monitoring system to a system network card, so as to identify and filter out target data, the target data comprising signaling protocol data and video stream protocol data, and the monitoring system comprising a monitoring device and a storage device; parsing the signaling protocol data so as to correspondingly acquire configuration information, session control information, and a video encoding parameter set of a video stream; on the basis of the configuration information, parsing the video stream protocol data and extracting field information, and on the basis of the field information, dividing the video stream protocol data into video data and audio data; and restoring the audio data and the video data, and synthesizing and synchronizing the restored audio data and the restored video data, so as to obtain a target video file. The present invention achieves the accurate format identification of a video stream and the restoration of audio and video entity files, thereby avoiding any interference with an existing monitoring system.
Need to check novelty before this filing date? Find Prior Art

Description

Non-perception RTP audio and video stream analysis and restoration method, system and medium TECHNICAL FIELD

[0001] The application belongs to the technical field of monitoring, and relates to a monitoring data restoration method, in particular to a non-perception RTP audio and video stream analysis and restoration method, system and medium. BACKGROUND

[0002] A monitoring system is one of the most commonly used systems in a security system. The mainstream construction site monitoring is to realize video monitoring by using a handheld video communication device. From the earliest analog monitoring to the digital monitoring in recent years and to the current network video monitoring, the monitoring technology has changed greatly. Today, with the gradual unification of IP technology all over the world, it is necessary to re-understand the development history of the video monitoring system. From the technical point of view, the development of the video monitoring system is divided into the first generation of analog video monitoring system (CCTV), the second generation of digital video monitoring system (DVR) based on a PC and a multimedia card, and the third generation of network video monitoring system (IPVS) based on IP network.

[0003] In the network video monitoring system, an external device needs to obtain the data transmitted by the network monitoring in real time. However, in the prior art, due to the lack of accurate identification of the video stream, the video stream is easy to interfere with the monitoring system when obtaining the monitoring video data, especially in some information acquisition scenes that need to be highly concealed, the influence is greater, and the use requirements cannot be met. SUMMARY

[0004] The purpose of the present application is to provide a non-perception RTP audio and video stream analysis and restoration method, system and medium, which is used to solve the problems in the background.

[0005] In a first aspect, the present application provides a non-perception RTP audio and video stream analysis and restoration method, which comprises: connecting an access module to a monitoring system and mirroring the traffic data in the monitoring system to a network card connected with the access module to identify and filter out target data, the target data comprising signaling protocol data and video stream protocol data, the monitoring system comprising a monitoring device and a storage device; analyzing the signaling protocol data to correspondingly obtain configuration information, session control information and a video encoding parameter set of the video stream; analyzing the video stream protocol data according to the configuration information and extracting field information, and dividing the video stream protocol data into video data and audio data according to the field information; restoring the audio data, restoring the video data, and synthesizing and synchronizing the restored audio data and the video data to obtain a target video file.

[0006] In an implementation form of the first aspect, the identifying and filtering out the target data comprises: separating a data link layer in the traffic data to obtain IP network layer data packets; separating the IP network layer according to the IP network layer data packets to obtain a transport layer data after extracting source IP, destination IP and protocol number; separating the transport layer according to the transport layer data to obtain an application layer load after extracting source port and destination port of TCP protocol and UDP protocol.

[0007] In an implementation form of the first aspect, the parsing the signaling protocol data to correspondingly obtain configuration information, session control information and video encoding parameter set of the video stream comprises: parsing RTSP protocol and SIP protocol to obtain SDP information; extracting the configuration information, the session control information and the video encoding parameter set of the video stream according to the SDP information.

[0008] In an implementation form of the first aspect, the parsing the video stream protocol data according to the configuration information and extracting field information, and dividing the video stream protocol data into video data and audio data according to the field information comprises: parsing the video stream protocol data according to the configuration information to extract sequence number, timestamp and load type information; sorting the video stream protocol data according to the sequence number; dividing the sorted video stream protocol data according to the video encoding parameter set to obtain corresponding audio data and video data respectively.

[0009] In an implementation form of the first aspect, when there is no signaling information in the video stream protocol data, the sorted video stream protocol data is distinguished according to the load type information standard definition to obtain video data and audio data respectively.

[0010] In an implementation form of the first aspect, the restoring the audio data comprises: obtaining RTP load type information of the audio data, identifying audio format according to the RTP load type value; restoring the audio data into an audio stream according to a decoding mode corresponding to the identified audio format; wherein the audio data in a non-standard format is stopped from being parsed.

[0011] In an implementation form of the first aspect, the restoring the video data comprises: identifying a NAL unit with encoding parameters in the video stream data, and identifying whether it is a parameter set according to a first byte in the NAL unit; extracting the encoding parameters in the NAL unit after determining that the NAL unit is a parameter set; identifying a video encoding format of the video data according to the video encoding parameter set or the encoding parameters; sequentially adding a start code to the video encoding parameter set and splicing the video data in order to restore and decode a complete video raw stream.

[0012] In an implementation form of the first aspect, the synthesizing and synchronizing the audio data and the video data after the restoring comprises: in a streaming session, periodically sending, by a sender device, an RTCP control packet, the RTCP including an SR and an RR; parsing the RTCP control packet to extract an RTCP timestamp field; converting the RTCP timestamp field into an NTP timestamp field; adjusting a play time of the audio data and the video data by the NTP timestamp field to synthesize and synchronize the audio data and the video data.

[0013] In an implementation form of the first aspect, the restoring the video data further comprises: when it is determined that the video stream data lacks encoding parameters, using a plurality of general video encoding parameter combinations to restore and decode the video data until successful decoding or decoding failure.

[0014] In a second aspect, the present application provides a non-intrusive RTP audio and video stream analysis and restoration system, comprising: an access module, configured to access a monitoring system, mirror traffic data in the monitoring system to a system network card, identify and filter out target data, the target data including signaling protocol data and video stream protocol data, and the monitoring system including a monitoring device and a storage device; an analysis module, configured to analyze the signaling protocol data to correspondingly obtain configuration information, session control information and a video encoding parameter set of a video stream; a division module, configured to analyze the video stream protocol data according to the configuration information and extract field information, and divide the video stream protocol data into video data and audio data according to the field information; a restoration module, configured to restore the audio data, restore the video data, and synthesize and synchronize the restored audio data and the restored video data to obtain a target video file.

[0015] In a third aspect, the present application further provides a computer readable storage medium having a computer program stored thereon, the program being executed to implement the non-intrusive RTP audio and video stream analysis and restoration system.

[0016] As described above, the non-intrusive RTP audio and video stream analysis and restoration method, system and medium have the following beneficial effects:

[0017] Compared with the prior art, the application can automatically identify the video coding format (such as H264, H265, etc.) by accessing the network traffic mirror of the switch, extract the video stream in the target site intranet, and realize the accurate format identification of the video stream and the restoration of the audio and video entity files by deeply analyzing the RTP data without actively sending a SIP protocol request or other probing operations, and the non-perception feature ensures that the system is completely transparent to the monitoring party, avoids any interference with the existing monitoring system, and is particularly suitable for information acquisition scenarios that require high concealment. BRIEF DESCRIPTION OF DRAWINGS

[0018] FIG. 1 shows a flowchart of the non-perception RTP audio and video stream analysis and restoration method according to the embodiments of the application.

[0019] FIG. 2 shows a structural block diagram of the non-perception RTP audio and video stream analysis and restoration system according to the embodiments of the application. DETAILED DESCRIPTION

[0020] The embodiments of the application are described below by specific examples, and those skilled in the art can easily understand other advantages and effects of the application from the disclosure. The application can also be implemented or applied by different specific embodiments, and the details in the specification can be modified or changed based on different views and applications without departing from the spirit of the application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0021] It should be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the application, and only show the components related to the application, not the number, shape and size of the components when actually implemented. The actual implementation of each component may be randomly changed in type, number and proportion, and the component layout type may be more complex.

[0022] Referring to FIGS. 1 and 2, the embodiments of the application provide a non-perception RTP audio and video stream analysis and restoration method, system and medium, which can automatically identify the video coding format (such as H264, H265, etc.) without relying on the coding information transmitted in the initial session, extract the video stream in the target site intranet, and realize the accurate format identification of the video stream and the restoration of the audio and video entity files by deeply analyzing the RTP data without actively sending a SIP protocol request or other probing operations. Moreover, the non-perception feature ensures that the system is completely transparent to the monitoring party, avoids any interference with the existing monitoring system, and is particularly suitable for information acquisition scenarios that require high concealment.

[0023] RTP (Real-time Transport Protocol): Real-time Transport Protocol, used to transport audio and video data.

[0024] SIP (Session Initiation Protocol): Session Initiation Protocol, a signaling protocol used to initiate, maintain, and terminate multimedia communication sessions.

[0025] RTSP (Real-time Streaming Protocol): Real-time Streaming Protocol, a network control protocol used to control streaming media servers.

[0026] RTCP (Real-time Control Protocol): Real-time Control Protocol, provides quality feedback and synchronizes audio and video streams.

[0027] SDP (Session Description Protocol): Session Description Protocol, used to exchange audio and video information.

[0028] H.264: Digital video codec standard proposed by the Joint Video Team (JVT) of ITU-T Video Coding Experts Group and ISO / IEC Moving Picture Experts Group, has become one of the most commonly used formats for high-precision video encoding.

[0029] H.265: Digital video codec standard proposed by the Joint Video Team (JVT-VC), the next generation of H.264.

[0030] NAL (Network Abstraction Layer): Unit of video encoding data (such as H.264) transmitted over a network.

[0031] SPS (Sequence Parameter Set): Sequence Parameter Set, a key parameter set in H.264 and H.265 video encoding.

[0032] PPS (Picture Parameter Set): Picture Parameter Set, a key parameter set in H.264 and H.265 video encoding.

[0033] VPS (Video Parameter Set): Video Parameter Set, a key parameter set in H.265 video encoding.

[0034] I-frame (Intra-coded Frame): Intra-coded frame, usually contains complete one-frame image information.

[0035] Start Code Prefix: a prefix before each NAL, used to segment NAL units.

[0036] Five-tuple: the five parameters of source IP, destination IP, source port, destination port, and transport layer protocol.

[0037] NVR (Network Video Recorder): a network video recorder, a device used for managing and storing video monitoring.

[0038] UDP (User Datagram Protocol): a user datagram protocol, a datagram mode for providing packet switching computer communication in a group of interconnected computer networks.

[0039] TCP (Transmission Control Protocol): a transmission control protocol, a connection-oriented, reliable, and byte stream-based transport layer communication protocol.

[0040] The technical solutions in the embodiments of the present application will be described in detail below with reference to the drawings in the embodiments of the present application.

[0041] As shown in FIG. 1, in an embodiment, the present application provides a non-aware RTP audio and video stream analysis and restoration method, which comprises the following steps:

[0042] S101, an access module is accessed to a monitoring system, and traffic data in the monitoring system is mirrored to a system network card connected with the access module, so as to identify and filter out target data, the target data including signaling protocol data and video stream protocol data, and the monitoring system including monitoring devices and storage devices.

[0043] Specifically, first, a restoration system is accessed between monitoring devices and NVR devices in a monitoring system, so as to mirror traffic data in the monitoring devices to a system network card, so that the restoration system can identify and filter out target data. After the traffic data reaches the system network card, the restoration system analyzes the traffic data as follows.

[0044] In an embodiment, the identification and filtering of the target data include:

[0045] Separating the data link layer in the traffic data to obtain IP network layer data packets;

[0046] Separating the IP network layer to obtain the transport layer data according to the source IP, the destination IP, and the protocol number extracted from the IP network layer data packets;

[0047] The transport layer is separated to obtain the application layer load after the source port and the destination port of the TCP protocol and the UDP protocol are extracted from the transport layer data.

[0048] Specifically, in the process of restoring the system to analyze the traffic data, the data link layer is first separated to obtain the IP network layer data packet, the source IP, the destination IP, and the protocol number are extracted, and then the IP network layer is separated to obtain the transport layer data, and the source port and the destination port of the TCP and UDP protocols are extracted. After obtaining the five-tuple, the transport layer is separated to obtain the application layer load. According to the characteristics of the Real-Time Streaming Protocol (RTSP), the messages with the TCP protocol and the source or destination port of 554 are filtered out; according to the characteristics of the Session Initiation Protocol (SIP), the messages with the UDP protocol and the source or destination port of 5060 are filtered out. According to the characteristics of the Real-Time Control Protocol (RTCP), the messages with the UDP protocol, and the first byte of the UDP load being 0x80 or 0xA0 and the second byte being C8 are filtered out; according to the characteristics of the Real-Time Transport Protocol (RTP), the first byte of the UDP load being 0x80 or A0, and the 9th byte to the 12th byte of the UDP load being the same UDP stream from the first packet are filtered out.

[0049] S102, analyzing the signaling protocol data to correspondingly obtain configuration information, session control information, and a video encoding parameter set of the video stream.

[0050] In an embodiment, the analyzing the signaling protocol data to correspondingly obtain configuration information, session control information, and a video encoding parameter set of the video stream comprises:

[0051] analyzing the RTSP protocol and the SIP protocol to obtain SDP information;

[0052] extracting the configuration information, the session control information, and the video encoding parameter set of the video stream according to the SDP information.

[0053] Specifically, the system analyzes the RTSP protocol and the SIP protocol to analyze the SDP information, extracts the configuration information of the audio and video stream from the SDP information, such as the port of the RTP stream, the RTP load type and the sampling rate of the audio and video, and the video encoding parameter set such as SPS and PPS, decodes the parameter set by base64 to obtain the original hexadecimal value, and thus obtains the original video encoding parameter set.

[0054] S103, analyzing the video stream protocol data according to the configuration information and extracting field information, and dividing the video stream protocol data into video data and audio data according to the field information.

[0055] In an embodiment, the parsing the video stream protocol data according to the configuration information and extracting field information, and dividing the video stream protocol data into video data and audio data according to the field information, comprises:

[0056] parsing the video stream protocol data according to the configuration information to extract sequence number, timestamp and load type information;

[0057] sorting the video stream protocol data according to the sequence number;

[0058] dividing the sorted video stream protocol data according to the video encoding parameter set to obtain corresponding audio data and video data respectively.

[0059] After obtaining the configuration information, the video stream protocol data is further parsed according to the configuration information parsed from the signaling to extract fields including sequence number, timestamp, load type, etc. The RTP stream is correctly sorted according to the sequence number information to avoid the influence of out-of-order on the video stream. The same processing is also done for pure RTP stream without signaling information. Then the payload is extracted from the video stream protocol data, and the audio and video traffic is distinguished according to the video encoding parameter set parsed in the foregoing.

[0060] Among them, the load type of audio is generally less than 34; and the load type value of video is generally greater than or equal to 96, i.e. dynamic type. The audio and video data are processed respectively to obtain audio data and video data for subsequent decoding and restoration.

[0061] It should be noted that when the video stream protocol data does not have signaling information, the sorted video stream protocol data is distinguished according to the load type information standard definition to obtain video data and audio data respectively.

[0062] S104, restoring the audio data, restoring the video data, and synthesizing and synchronizing the restored audio data and video data to obtain a target video file.

[0063] In an embodiment, the restoring the audio data comprises:

[0064] obtaining the RTP load type information of the audio data, and identifying the audio format according to the RTP load type value;

[0065] restoring the audio data into an audio stream according to the decoding mode corresponding to the identified audio format;

[0066] Among them, for the audio data of non-standard format, the parsing is stopped.

[0067] In the embodiment, since the monitoring device generally uses standard audio formats such as PCMA, PCMU, G722, G729, etc., after the audio stream data is acquired, the restoration system identifies the audio format through the RTP payload type value, restores the audio stream according to the standard decoding process of each format, does not parse the audio system of a non-standard format, thereby obtaining the original audio data, and completes the restoration of the audio data.

[0068] In one embodiment, the restoration of the video data comprises:

[0069] identifying the NAL unit with the encoding parameter in the video stream data, and identifying whether it is a parameter set according to the first byte in the NAL unit;

[0070] after determining that the NAL unit is a parameter set, extracting the encoding parameter in the NAL unit;

[0071] identifying the video encoding format of the video data according to the video encoding parameter set or the encoding parameter;

[0072] adding a start code to the video encoding parameter set in sequence, and splicing the video data in sequence, to restore and decode the complete video raw stream.

[0073] In the embodiment, after the separated video data is acquired, the NAL unit with the encoding parameter is first identified, and whether it is a parameter set is identified through the first byte of the NAL unit. The judgment condition is:

[0074] A. reading the first byte of the NAL unit, taking the last 5 bits (from the 3rd bit to the 7th bit), which is 7 or 8 in decimal;

[0075] B. reading the first byte of the NAL unit, taking the middle 6 bits (from the 1st bit to the 6th bit), which is 32, 33 or 34 in decimal.

[0076] As long as one of the above conditions is met, that is, the type (nal_unit_type) of the NAL unit is 7 / 8 / 32 / 33 / 34, it is considered that the NAL unit is with the encoding parameter, and the parameter in the NAL unit is extracted as the encoding parameter.

[0077] Then, the video encoding format is identified according to the video encoding parameter set acquired from the signaling protocol data or the encoding parameter extracted from the video data.

[0078] If the video data does not have a corresponding signaling protocol, for example, the system accesses the traffic in the middle of the monitoring, at this time, the encoding format can be judged through the NAL unit type of the encoding parameter:

[0079] If the type nal_unit_type of the NAL unit in the video data is 7 or 8, it is determined that the encoding of the video is H.264;

[0080] If the type nal_unit_type of the NAL unit in the video data is 32, 33, or 34, it is determined that the encoding of the video stream is H.265.

[0081] If there is no corresponding signaling protocol and no NAL for transmitting encoding information in the video data, only pure video frame data is transmitted (for example, a monitoring device only sends encoding information at the beginning of traffic transmission), and the encoding format of the video needs to be identified by identifying the data characteristics of the I frame and other video frames.

[0082] Specifically, if the first byte of more than 90% of the NAL units in the video data is 0x7c, 0x5c, or 0x3c, it is determined that the encoding of the video stream is H.264.

[0083] If the first byte of more than 90% of the NAL units in the video data is 0x62 or 0x02, it is determined that the encoding of the video stream is H.265.

[0084] After determining the video encoding format of the video data, the extracted video encoding parameter information is added with a start code and parsed with the video data to splice and restore the complete video raw stream in sequence.

[0085] For H.264 encoding, the processing flow is as follows:

[0086] When the value of the 5 bits after the 0th byte of the NAL unit is 28 (0x1C), when the value of the 0th bit of the 1st byte is 1, it represents that the frame is the first frame, and the start code (0x00000001) is written first, then the value of the 3 bits before the 0th byte and the 5 bits after the 1st byte are taken to form a new byte, and the value of the byte is written, and then the remaining byte content starting from the 2nd byte is written. When the value of the 0th bit of the 1st byte is not 1, it represents that the frame is not the first frame, and the remaining byte content starting from the 2nd byte is directly written. When the value of the 5 bits after the 0th byte is not 28 (0x1C), but 1 / 5 / 6 / 7 / 8: the start code (0x00000001) is written first, and then the remaining byte content starting from the 0th byte is written.

[0087] For H.265 encoding, the processing flow is as follows:

[0088] When the value of the middle 6 bits of the 0th byte is 49 (0x31), when the value of the 0th bit of the 2nd byte is 1, it represents that the frame is the first frame, the start code (0x00000001) is written first, then the value of the last 6 bits of the 2nd byte is taken and left shifted by one bit, the value of the byte is written, then 0x01 is written, and then the remaining byte content from the 3rd byte is written. When the value of the 0th bit of the 2nd byte is not 1, it represents that the frame is not the first frame, and the remaining byte content from the 3rd byte can be directly written. When the value of the middle 6 bits of the 0th byte is not 49 (0x31), the start code (0x00000001) is written first, and then the remaining byte content from the 0th byte is written.

[0089] In the case of lacking video coding parameter set, the resolution, frame rate, chroma format, and color depth of the video stream are unknown. In practice, the monitoring device adopts default parameters, and the parameter value range is in a few fixed values, such as the resolution is generally 1920x1080, 1280x720, etc., and the frame rate is generally 24, 25, 50, etc. The system attempts to use multiple common video coding parameter combinations, combines with the video frame respectively, restores the video stream, and decodes the video stream until successful decoding or confirmation of failure to decode.

[0090] It should be noted that the correctness of some parameters does not determine whether the decoder can successfully decode, such as the frame rate only affects the frame interval and video duration, and incorrect resolution may cause the picture to present meaningless blurred images, so multiple video raw streams may be successfully decoded.

[0091] If the monitoring video adjusts the coding parameters in the middle, such as modifying the frame rate, resolution, or coding format, the modified data must have SPS and PPS parameter sets, which can be parsed according to the above process.

[0092] It should be noted that when it is determined that the video stream data lacks coding parameters, multiple common video coding parameter combinations are used to restore and decode the video data until successful decoding or failure to decode.

[0093] In an embodiment, the synthesis and synchronization of the restored audio data and video data include:

[0094] In a streaming media session, the sending end device periodically sends an RTCP control packet, and the RTCP includes SR and RR;

[0095] The RTCP control packet is parsed to extract the RTCP timestamp field;

[0096] The RTCP timestamp field is converted into an NTP timestamp field;

[0097] The play time of the audio data and the video data is adjusted through the NTP time stamp field to synchronize the audio data and the video data.

[0098] In the streaming session, the sender device periodically sends RTCP control packets, the control packets including a sender report (SR) and a receiver report (RR), and the system parses the RTCP protocols, extracts the time stamp fields provided by the RTCP, including an NTP time stamp and an RTP time stamp, and then converts the RTP time stamp into the NTP time stamp, adjusts the play time of the audio and video streams through the fields, and ensures that the restored video and audio are synchronized.

[0099] After the audio data and the video data are restored and synchronized, the restored audio stream and the video stream are encapsulated into an MP4 or other common media file format by using an ffmpeg tool for further processing or storage.

[0100] The protection scope of the non-aware RTP audio and video stream analysis and restoration method described in the embodiments of the present application is not limited to the execution order of the steps listed in the embodiments, and any scheme achieved by adding, replacing or replacing steps according to the principle of the present application is included in the protection scope of the present application.

[0101] The embodiments of the present application also provide a computer readable storage medium having a computer program stored thereon, the program being executed by an electronic device to implement the non-aware RTP audio and video stream analysis and restoration method described above.

[0102] Those skilled in the art can understand that all or part of the steps in the method of the above embodiments can be instructed by a program to complete the processor, and the program can be stored in a computer readable storage medium, the storage medium is a non-transitory medium, for example, random access memory, read only memory, flash memory, hard disk, solid state disk, magnetic tape, floppy disk, optical disc and any combination thereof. The above storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), optical medium (for example, digital video disc (digital video disc, DVD)) or semiconductor medium (for example, solid state disk (solid state disk, SSD)) etc.

[0103] The embodiment of the present application also provides a non-perception RTP audio and video stream analysis and restoration system which can realize the non-perception RTP audio and video stream analysis and restoration method of the present application, but the implementation device of the non-perception RTP audio and video stream analysis and restoration method of the present application includes but is not limited to the structure of the non-perception RTP audio and video stream analysis and restoration system listed in the embodiment, and any structure deformation and replacement of the prior art according to the principle of the present application is included in the protection scope of the present application.

[0104] As shown in Figure 2, in an embodiment, the present application also provides a non-perception RTP audio and video stream analysis and restoration system, which comprises:

[0105] The access module 201 is used for accessing a monitoring system and mirroring traffic data in the monitoring system to a system network card to identify and filter out target data, wherein the target data comprises signaling protocol data and video stream protocol data, and the monitoring system comprises a monitoring device and a storage device.

[0106] The analysis module 202 is used for analyzing the signaling protocol data to correspondingly acquire configuration information, session control information and video encoding parameter set of the video stream.

[0107] The division module 203 is used for analyzing the video stream protocol data according to the configuration information and extracting field information, and dividing the video stream protocol data into video data and audio data according to the field information.

[0108] The restoration module 204 is used for restoring the audio data, restoring the video data, and synthesizing and synchronizing the restored audio data and the video data to obtain a target video file.

[0109] It should be noted that the structure and principle of the access module 201, the analysis module 202, the division module 203 and the restoration module 204 correspond to the steps (step S101 to step S104) in the non-perception RTP audio and video stream analysis and restoration method, and the specific working principle can refer to the introduction of the non-perception RTP audio and video stream analysis and restoration method in the foregoing embodiment, and thus will not be described here.

[0110] In several embodiments provided by the present application, it should be understood that the disclosed system, device or method can be implemented in other manners. For example, the described device embodiment is merely illustrative. For example, the division of the modules / unit can be different, for example, a plurality of modules or units can be combined or can be integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or modules, and can be in electrical, mechanical or other forms.

[0111] The modules / unit described as separate components can or can not be physically separate, and the components shown as modules / unit can or can not be physical modules, i.e., can be located in one place or can be distributed to multiple network units. Some or all of the modules / unit can be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, the functional modules / unit in each embodiment of the present application can be integrated into a processing module, or each module / unit can be physically separate, or two or more modules / unit can be integrated into one module / unit.

[0112] Those of ordinary skill in the art should further appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0113] The above description of the flow or structure of each figure has its own emphasis, and the parts not described in detail in a certain flow or structure can refer to the related description of other flows or structures.

[0114] The above embodiments are only illustrative of the principles and effects of the present application, and are not intended to limit the present application. Any person skilled in the art can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those skilled in the art without departing from the spirit and technical idea disclosed by the present application should be covered by the claims of the present application.

Claims

1. A method for non-aware RTP audio / video stream parsing and restoring, characterized in that, The method comprises: Accessing a monitoring system with an access module, and mirroring traffic data in the monitoring system to a network card connected to the access module to identify and filter out target data, the target data including signaling protocol data and video stream protocol data, the monitoring system including monitoring equipment and storage equipment; Parsing the signaling protocol data to correspondingly obtain configuration information, session control information, and a video encoding parameter set of a video stream; According to the configuration information, parsing the video stream protocol data and extracting field information, and according to the field information, dividing the video stream protocol data into video data and audio data; Restoring the audio data, restoring the video data, and synthesizing and synchronizing the restored audio data and the video data to obtain a target video file.

2. The method of claim 1, wherein the method further comprises: The identifying and filtering out of the target data comprises: Separating a data link layer in the traffic data to obtain IP network layer data packets; According to the IP network layer data packets, extracting source IP, destination IP, and protocol number, and then separating the IP network layer to obtain a transport layer data; According to the transport layer data, extracting source ports and destination ports of TCP protocol and UDP protocol, and then separating the transport layer to obtain application layer load.

3. The method of claim 2, wherein the method further comprises: The parsing of the signaling protocol data to correspondingly obtain configuration information, session control information, and a video encoding parameter set of a video stream comprises: Parsing RTSP protocol and SIP protocol to obtain SDP information; According to the SDP information, extracting configuration information, session control information, and a video encoding parameter set of a video stream.

4. The method of claim 3, wherein the method further comprises: The parsing of the video stream protocol data according to the configuration information and the extraction of field information, and the division of the video stream protocol data into video data and audio data according to the field information comprise: According to the configuration information, parsing the video stream protocol data to extract sequence number, timestamp, and load type information; According to the sequence number, sorting the video stream protocol data; According to the video encoding parameter set, dividing the sorted video stream protocol data to obtain corresponding audio data and video data.

5. The method of claim 4, wherein the method further comprises: When there is no signaling information in the video stream protocol data, according to the load type information standard definition, distinguishing the sorted video stream protocol data to obtain video data and audio data.

6. The method of claim 1, wherein the method further comprises: The restoring of the audio data comprises: Obtaining RTP load type information of the audio data, and according to the RTP load type value, identifying an audio format; According to a decoding mode corresponding to the identified audio format, restoring the audio data into an audio stream; Wherein, for non-standard format audio data, stop parsing.

7. The method of claim 3, wherein the method further comprises: receiving a request for a media stream from a client; and sending a response to the request to the client, the response including a description of the media stream and a description of the media stream in a format that is different from the format of the media stream. The restoring of the video data comprises: Identifying NAL units with encoding parameters in the video stream data, and according to a first byte in the NAL units, identifying whether it is a parameter set; After determining that the NAL unit is a parameter set, extracting the encoding parameters in the NAL unit; According to the video encoding parameter set or the encoding parameters, identifying a video encoding format of the video data; The video coding parameter set is added with a start code in sequence, and the video data is spliced in sequence to restore and decode a complete video raw stream.

8. The method of claim 7, wherein the method further comprises: The restored audio data and the video data are synthesized and synchronized, comprising: In a streaming session, a sender device periodically sends an RTCP control packet, the RTCP including SR and RR; The RTCP control packet is parsed to extract an RTCP timestamp field; The RTCP timestamp field is converted into an NTP timestamp field; The play time of the audio data and the video data is adjusted by the NTP timestamp field to synthesize and synchronize the audio data and the video data.

9. The seamless RTP audio and video stream parsing and restoration method according to claim 7, characterized in that, The restoring of the video data further comprises: When it is determined that the video stream data lacks coding parameters, a plurality of general video coding parameter combinations are used to restore and decode the video data until successful decoding or decoding failure.

10. A non-aware RTP audio / video stream parsing and restoring system, characterized in that, Comprise: An access module, configured to access a monitoring system and mirror traffic data in the monitoring system to a network card connected with the access module to identify and filter out target data, the target data including signaling protocol data and video stream protocol data, the monitoring system including a monitoring device and a storage device; An analysis module, configured to analyze the signaling protocol data to correspondingly acquire configuration information, session control information and a video coding parameter set of a video stream; A division module, configured to analyze the video stream protocol data according to the configuration information and extract field information, and divide the video stream protocol data into video data and audio data according to the field information; A restoration module, configured to restore the audio data, restore the video data, and synthesize and synchronize the restored audio data and the video data to obtain a target video file.

11. A computer readable storage medium having stored thereon a computer program, characterized in that The program is executed to implement the non-aware RTP audio and video stream analysis and restoration system in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method and device for measuring lost step between network side audio and video streams based on RTSP

    CN103561260A

  • Video file restoration method and device, computer equipment and storage medium

    CN113438505A

  • Method and device for realizing conversion from any streaming media protocol to NDI

    CN113645485A

  • System and method for detecting and calibrating time abnormity of video and image acquisition equipment in real time

    CN116170104A

  • All-media fusion audio and video recording and video-on-demand system and processing method

    CN116614682A