Video synchronization method and video acquisition equipment

By packaging the video data of each camera under the same time stamp in a multi-mesh camera into target video data and transmitting it to the video receiver, the video processing error problem caused by independent asynchronous transmission of each camera in a multi-mesh camera is solved, and the synchronous transmission of video frames and private data is realized, ensuring the accuracy and stability of video processing.

CN120050489APending Publication Date: 2025-05-27HANGZHOU EZVIZ SOFTWARE CO LTD

Patent Information

Application Number
CN202510191920.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Each camera in a multi-eye camera independently transmits video frames and private data asynchronously, resulting in out-of-synchronous reception of the receiver and causing video processing errors.

Method used

By packaging the video data of each camera under the same time stamp into target video data and transmitting it to the video receiver, ensure that the video frame and private data arrive synchronously.

Benefits of technology

This avoids video processing errors caused by the video receiver due to the abnormal reception of video frames and private data of each camera, and ensures the accuracy and stability of video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050489A_ABST
    Figure CN120050489A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video synchronization method and video acquisition equipment. In the embodiment of the invention, the video data corresponding to each camera under the same timestamp is directly packaged into the target video data corresponding to the timestamp, and the target video data corresponding to the timestamp is transmitted to the video receiving end; therefore, the video frames and the private data corresponding to the cameras in the video acquisition equipment can be ensured to synchronously reach the video receiving end instead of independently and asynchronously transmitting the video frames and the private data corresponding to the cameras to the video receiving end in the same existing video acquisition equipment; and a video processing error caused by asynchronously receiving the video frames and the private data of the cameras at the video receiving end is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technologies, and in particular, to a video synchronization method and a video acquisition device. Background Art

[0002] A multi-camera refers to a camera that includes two or more cameras. Currently, in a real-time communication scenario, each camera in a multi-camera independently acquires and encodes video frames. Moreover, for each camera in the multi-camera, the video frame corresponding to this camera and private data are independently transmitted to a receiving end for corresponding video processing.

[0003] However, the method of independently and asynchronously transmitting video frames and private data by each camera in a multi-camera may cause video processing errors at the receiving end due to receiving the video frames and private data of each camera out of sync. Summary of the Invention

[0004] In view of this, this application provides a video synchronization method and a video acquisition device to avoid video processing errors caused by receiving the video frames and private data of each camera out of sync at the receiving end.

[0005] An embodiment of this application provides a video synchronization method. This method is applied to a video acquisition device, and the video acquisition device is equipped with N cameras, where N is greater than 1. The method includes:

[0006] Packaging the video data corresponding to each camera at the same timestamp into target video data corresponding to this timestamp; wherein, the video data corresponding to any one camera at this timestamp is obtained by associating the video frame of this camera at this timestamp and the Supplemental Enhancement Information (SEI) data packet of this camera at this timestamp; the SEI data packet is obtained based on the private data of this camera at this timestamp; the video data corresponding to any one camera at least carries a Synchronization Source Identifier (SSRC); the video data corresponding to different cameras carries different SSRCs;

[0007] Transmitting the target video data corresponding to this timestamp to a video receiving end, so that the video receiving end can obtain the private data of each camera at this timestamp and the video frame at this timestamp based on the SSRC carried by the target video data.

[0008] An embodiment of this application also provides a video acquisition device.

[0009] The video acquisition device is equipped with N cameras and a processor; N is greater than 1;

[0010] A camera for collecting original frames; the original frames are used to generate video frames of the camera;

[0011] A processor for obtaining the private data of each camera;

[0012] The processor is further configured to pack the video data corresponding to each camera at the same timestamp into target video data corresponding to the timestamp; wherein, the video data corresponding to any camera at the timestamp is obtained by associating the video frame of the camera at the timestamp and the SEI data packet of the camera at the timestamp; the SEI data packet is obtained based on the private data of the camera at the timestamp; the video data corresponding to any camera carries at least an SSRC; the video data corresponding to different cameras carries different SSRCs;

[0013] Transmit the target video data corresponding to the timestamp to a video receiving end, so that the video receiving end can obtain the private data of each camera at the timestamp and the video frames at the timestamp based on the SSRC carried by the target video data.

[0014] It can be seen from the above technical solutions that in the embodiments of the present application, the video data corresponding to each camera at the same timestamp is directly packed into the target video data corresponding to the timestamp, and the target video data corresponding to the timestamp is transmitted to the video receiving end, which can ensure that the video frames and private data corresponding to each camera in the video acquisition device reach the video receiving end synchronously, rather than independently and asynchronously transmitting the video frames and private data corresponding to each camera in the existing same video acquisition device to the video receiving end, avoiding video processing errors caused by receiving the video frames and private data of each camera out of sync at the video receiving end.

[0015] Furthermore, in the embodiments of the present application, by packing the private data of each camera at the same timestamp into SEI data packets, it can be associated with the video frames of each camera at the timestamp, thereby realizing the synchronous transmission of the private data and video frames of each camera at the same timestamp; in addition, the SSRCs carried by the video data corresponding to each camera in this embodiment are different for the video receiving end to distinguish each video data. Description of the Drawings

[0016] The drawings here are incorporated into the specification and form a part of this application, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0017] Figure 1 It is a schematic flowchart of the method provided by the embodiments of the present application.

[0018] Figure 2Schematic diagram of the implementation of video frame processing provided by an embodiment of the present application.

[0019] Figure 3 Schematic diagram of the implementation of another video frame processing provided by an embodiment of the present application.

[0020] Figure 4 Schematic diagram of the implementation of the video synchronization method provided by an embodiment of the present application.

[0021] Figure 5 Schematic diagram of the structure of the video acquisition device provided by an embodiment of the present application. Detailed implementation

[0022] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, and to make the above-mentioned objects, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0023] Refer to Figure 1 , Figure 1 which is a flowchart of the method provided by an embodiment of the present application. In this embodiment, the method is applied to a video acquisition device equipped with N cameras, where N is greater than 1. As an example, the video acquisition device in this embodiment may be, for example, a multi-camera, etc., and is not specifically limited here.

[0024] As Figure 1 shown, the process may include the following steps:

[0025] Step 101, pack the video data corresponding to each camera in the video acquisition device at the same timestamp into the target video data corresponding to the timestamp; among them, the video data corresponding to any camera at the timestamp is obtained by associating the video frame of the camera at the timestamp and the SEI data packet of the camera at the timestamp; the SEI data packet is obtained based on the private data of the camera at the timestamp.

[0026] In this embodiment, as an example, the video frame of any camera in the video acquisition device at a timestamp can be obtained based on the original frame collected by the camera at the timestamp.

[0027] For example, as an embodiment, obtaining the video frame of the camera at the timestamp based on the original frame collected by the camera at the timestamp can be specifically implemented as follows: First, encode the original frame collected by the camera, such as H.264 encoding. Then, package the encoded data in the RTP format to obtain the video frame of the camera at the timestamp. Here, the timestamp can be determined based on the acquisition time of the original frame, and this embodiment does not make specific limitations. RTP is the abbreviation of Real-time Transport Protocol.

[0028] In this embodiment, as an embodiment, in this step, obtaining the SEI data packet of the camera at the timestamp based on the private data of the camera at the timestamp can be specifically implemented as follows: Based on the SEI mechanism, convert the private data of the camera at the timestamp into SEI information, and package the SEI information in the RTP format to obtain the SEI data packet of the camera at the timestamp. Among them, SEI refers to a mechanism defined in the H.264 / AVC standard for transmitting supplementary information; the SEI information can be transmitted together with the video frame without a separate transmission channel; AVC is the abbreviation of Advanced Video Coding.

[0029] Based on the above description, it can be seen that by packaging the private data of any camera at a timestamp into an SEI data packet, it can be used to realize the association with the video frame of the camera at the timestamp, so as to further realize the synchronous transmission of the private data and the video frame of the camera at the timestamp. As for the specific content of the private data in this step, it will be described by examples below and will not be elaborated here.

[0030] In this embodiment, as an embodiment, the video data corresponding to any camera in the video acquisition device carries at least an SSRC. Among them, SSRC can refer to the identifier defined in RTP for uniquely identifying the data source; the data source here refers to the camera corresponding to the video data. Based on this, in this embodiment, the video data corresponding to different cameras carries different SSRCs to be used to distinguish each video data.

[0031] As for how to specifically package the video data corresponding to each camera at the same timestamp into the target video data corresponding to the timestamp and how to specifically obtain the video data corresponding to each camera at the timestamp, it will be described by examples below and will not be elaborated here for the time being.

[0032] Step 102: Transmit the target video data corresponding to the above timestamp to the video receiving end, so that the video receiving end can obtain the private data of each camera at this timestamp and the video frames at this timestamp based on the SSRC carried in the target video data.

[0033] In this embodiment, after obtaining the target video data corresponding to the above timestamp, the target video data can be transmitted to the video receiving end, which can ensure that the video frames and private data of each camera arrive at the video receiving end synchronously at this timestamp, thus avoiding video processing errors caused by the asynchronous reception of the video frames and private data of each camera at the video receiving end.

[0034] In this embodiment, the User Datagram Protocol (UDP) can ensure the smoothness and low latency of data transmission. Based on this, as an embodiment, the target video data corresponding to the above timestamp can be transmitted to the video receiving end based on the UDP protocol; this can ensure that the video frames and private data corresponding to each camera arrive at the video receiving end synchronously, and at the same time, can ensure the smoothness and low latency of the transmission of the video frames and private data corresponding to each camera.

[0035] Furthermore, in this embodiment, on the basis of the above UDP protocol, a Quality of Service (QoS) policy can be preconfigured for the transmission of the target video data to provide end-to-end quality of service guarantee, which can effectively improve the anti-weak network ability of data transmission. As for how to specifically configure the QoS policy, this embodiment does not make specific limitations. Among them, QoS refers to a network security mechanism that uses various basic technologies to provide better service capabilities for specified network communications, and can be used to solve problems such as network latency and congestion, and improve network service quality.

[0036] Thus far, the Figure 1 shown process is completed.

[0037] Through Figure 1 As can be seen from the shown process, in the embodiment of the present application, the video data corresponding to each camera at the same timestamp is directly packed into the target video data corresponding to this timestamp, and the target video data corresponding to this timestamp is transmitted to the video receiving end, which can ensure that the video frames and private data corresponding to each camera in the video acquisition device arrive at the video receiving end synchronously, rather than the existing independent asynchronous transmission of the video frames and private data corresponding to each camera in the same video acquisition device to the video receiving end, avoiding video processing errors caused by the asynchronous reception of the video frames and private data of each camera at the video receiving end.

[0038] Furthermore, in the embodiments of the present application, by packaging the private data of each camera at the same timestamp into an SEI data packet, it can be associated with the video frame of each camera at this timestamp, thereby realizing the synchronous transmission of the private data and the video frame of each camera at the same timestamp; in addition, the SSRCs carried by the video data corresponding to each camera in this embodiment are different, so as to be used by the video receiving end to distinguish each video data.

[0039] The private data in step 101 above will be described first below:

[0040] In this embodiment, as an example, obtaining the private data of any camera may be specifically implemented as follows: after the camera captures the original frame, the private data of the camera is obtained. Based on this, after obtaining the private data of the camera, the private data of the camera and the timestamp corresponding to the video frame of the camera can be set to the same timestamp.

[0041] In this embodiment, as an example, the private data of any camera at least includes the relevant information of the camera. Here, the relevant information of the camera is not specifically limited and can be flexibly set according to actual application requirements; for example, the relevant information of the camera may include the position information of the camera, and the human detection information obtained by performing human detection on the video frame of the camera, etc.

[0042] Among them, the above-mentioned position information of the camera may refer to the position information of the camera relative to the reference camera; for example, the position information of the camera may be that the camera is directly above the reference camera, or the camera is directly below the reference camera, and so on. The reference camera may be any camera in the multi-camera.

[0043] Next, how to determine the video data corresponding to any camera at a timestamp will be described:

[0044] In this embodiment, as an example, the above-mentioned determination of the video data corresponding to any camera at a timestamp may be specifically implemented as follows:

[0045] If the video frame of the camera at this timestamp is a key frame, where the key frame includes at least the following data packets: a Sequence Parameter Set (SPS) data packet, a Picture Parameter Set (PPS) data packet, and a Network Abstraction Layer Unit for I-frame (I-FU) data packet of at least one key frame, then: Insert the SEI data packet of the camera at this timestamp into the following position in the video frame: the position after the PPS data packet and before the last I-FU data packet, to obtain the video data corresponding to the camera at the timestamp.

[0046] If the video frame of the camera at this timestamp is a non-key frame, then insert the SEI data packet of the camera at this timestamp into the following position in the video frame: the position before the last data packet, to obtain the video data corresponding to the camera at this timestamp.

[0047] In this embodiment, as an example, to determine whether the video frame of any camera at a timestamp is a key frame, in a specific implementation, for example, it can be: Based on the frame identifier in the video frame of the camera at this timestamp, determine whether the video frame of the camera at this timestamp is a key frame.

[0048] For ease of understanding how to determine the video data corresponding to any camera at a timestamp as described above, the following is combined with Figure 2 and Figure 3 for an example description:

[0049] Exemplarily, as Figure 2 shown, assume that the video acquisition device is equipped with 2 cameras, which can be respectively denoted as the first-eye camera and the second-eye camera; among them, taking the first-eye camera as an example, assume that the first video frame of the first-eye camera at the first timestamp (i.e., timestamp 0) is a key frame (also known as an I-frame), which includes an SPS data packet, a PPS data packet, and 3 I-FU data packets. The SEI data packet obtained based on the private data of the first-eye camera at the first timestamp can be inserted into the position after the PPS data packet to obtain the video data corresponding to the camera at the first timestamp.

[0050] As Figure 3As shown, assume that the second video frame captured by the single-lens camera at the second timestamp (i.e., timestamp 1) is a predicted frame (also known as a P-frame), which includes 2 Network Abstraction Layer Units for P-frame (P-FU) data packets. The SEI data packet obtained based on the private data of the single-lens camera at the second timestamp can be inserted before the first P-FU data packet to obtain the video data corresponding to the camera at the second timestamp. Here, NAL is the abbreviation of Network Abstract Layer.

[0051] Thus, the description of how to determine the video data corresponding to any camera at a timestamp is completed.

[0052] Next, a description will be given on how to pack the video data corresponding to each camera at the same timestamp into the target video data corresponding to this timestamp:

[0053] In this embodiment, as an example, each data packet included in the video data corresponding to the same camera carries the same SSRC; each data packet included in the video data corresponding to different cameras carries different SSRCs to distinguish each video data.

[0054] In this embodiment, as an example, in the target video data, the last video data carries a marker bit; the marker bit is a first specified value used to indicate the end; where if other video data in the target video data also carries a marker bit, then this marker bit is set to a second specified value used to indicate non-end, so as to prevent the receiving end from missing video data when parsing the target video data and ensure the normal parsing of each video data in the target video data.

[0055] Based on this, as an example, to pack the video data corresponding to each camera in the video capture device at the same timestamp into the target video data corresponding to this timestamp, the specific implementation can be, for example: after obtaining the video data corresponding to each camera at the same timestamp, first arrange the video data corresponding to each camera at this timestamp in a specified order to obtain a video data sequence; here, the specified order can be, for example, the order determined based on the identifier of the camera. Specifically, it can be, for example, the single-lens camera, the dual-lens camera,..., the N-lens camera. This embodiment does not make specific limitations on this.

[0056] After that, adjust the above video data sequence. Specifically, each data packet included in the above video data sequence corresponds to a serial number, and the data packets included in the above video data sequence are organized together in the order of the serial numbers, and the difference between adjacent serial numbers is a preset value; the preset value here can be flexibly set according to actual application requirements. For example, it can be 1, and this embodiment does not make specific limitations.

[0057] In addition, set the marker bit carried by the last video data in the above video data sequence to a first specified value to indicate the end; among them, if the marker bit is carried by other video data except the last video data in the video data sequence, set the marker bit carried by the other video data to a second specified value to indicate non-end.

[0058] After completing the adjustment of the above video data sequence, pack the video data sequence to obtain the target video data corresponding to the time stamp.

[0059] In this embodiment, as an example, the first data packet included in the target video data has a corresponding serial number; among them, if the time stamp is the first time stamp in the video synchronization process, it means that the video frame at this time stamp is the first video frame of the camera. In this case, the serial number of the first data packet included in the target video data is the initial value; the initial value here can be 1, for example, and this embodiment does not make specific limitations.

[0060] If the time stamp is not the first time stamp in the video synchronization process, it means that the video frame captured at this time stamp is not the first video frame of the camera. In this case, the serial number of the first data packet included in the target video data is the sum of the serial number of the last data packet included in the target video data corresponding to the previous time stamp and the preset value.

[0061] To facilitate understanding of how to pack the video data corresponding to each camera at the same time stamp into the target video data corresponding to the time stamp, the following is combined with Figure 2 and Figure 3 for an example description:

[0062] Exemplarily, as Figure 2 shown, assume that the data packets included in the first video frame of the first camera at the first time stamp can be respectively recorded as: the first SPS data packet, the first PPS data packet, and 3 first I-FU data packets; the data packets included in the first video frame of the second camera at the first time stamp can be respectively recorded as: the second SPS data packet, the second PPS data packet, and 2 second I-FU data packets; the SEI data packets of the first camera and the second camera at the first time stamp can be respectively recorded as the first SEI data packet and the second SEI data packet.

[0063] Based on the above description, first, arrange the video data corresponding to each camera at the first timestamp in the order of the single-lens camera and the dual-lens camera to obtain the first video data sequence. At this time, the sequence structure of the first video data sequence is: the first SPS data packet with serial number 1, the first PPS data packet with serial number 2, the first SEI data packet with serial number 1, 3 first I-FU data packets with serial numbers 3, 4, and 5 respectively, the second SPS data packet with serial number 1, the second PPS data packet with serial number 2, the second SEI data packet with serial number 1, and 2 second I-FU data packets with serial numbers 3 and 4 respectively; among them, each data packet included in the video data corresponding to the same camera carries a different SSRC. The SSRC carried by each data packet included in the video data corresponding to the single-lens camera is 1111, and the SSRC carried by each data packet included in the video data corresponding to the dual-lens camera is 2222.

[0064] After that, adjust the serial numbers of each data packet included in the above first video data sequence so that the difference between adjacent serial numbers is 1. And since the first timestamp is the first timestamp in the video synchronization process, the serial number of the first data packet in the first video data sequence is the initial value such as 1. At this time, the sequence structure of the first video data sequence is: the first SPS data packet with serial number 1, the PPS data packet with serial number 2, the first SEI data packet with serial number 3, 3 first I-FU data packets with serial numbers 4, 5, and 6 respectively, the second SPS data packet with serial number 7, the second PPS data packet with serial number 8, the second SEI data packet with serial number 9, and 2 second I-FU data packets with serial numbers 10 and 11 respectively.

[0065] As for the video data corresponding to the single-lens camera and the dual-lens camera at the second timestamp, as Figure 3 shown, assume that each data packet included in the second video frame of the single-lens camera at the second timestamp can be respectively recorded as: 2 first P-FU data packets; each data packet included in the second video frame of the dual-lens camera at the second timestamp can be respectively recorded as: the second P-FU data packet; the SEI data packets of the single-lens camera and the dual-lens camera at the second timestamp can be respectively recorded as: the third SEI data packet and the fourth SEI data packet; among them, the second timestamp is the next timestamp of the first timestamp.

[0066] Based on this, first, arrange the video data corresponding to each camera at the second timestamp in the order of the single-lens camera and the dual-lens camera to obtain the second video data sequence.

[0067] After that, the sequence numbers of each data packet included in the above second video data sequence are adjusted so that the difference between adjacent sequence numbers is 1; and since the second timestamp is not the first timestamp in the video synchronization process, therefore, in order to ensure the temporal coherence of the video data under different timestamps, the sequence number of the first data packet in this second video data sequence is the sum of the sequence number of the last data packet included in the target video data corresponding to the previous timestamp (i.e., 11) and a preset value (such as 1), which is 12. At this time, the sequence structure of this second video data sequence is: the third SEI data packet with sequence number 12, 2 first P-FU data packets with sequence numbers 13 and 14 respectively, the fourth SEI data packet with sequence number 15, and the second P-FU data packet with sequence number 16.

[0068] So far, the description of how to pack the video data corresponding to each camera in the video acquisition device at the same timestamp into the target video data corresponding to this timestamp is completed.

[0069] For the convenience of understanding the specific implementation process of the above video synchronization method, the following combines Figure 4 , and describes it by way of specific embodiments.

[0070] See Figure 4 The schematic diagram of the implementation of the video synchronization method shown. As Figure 4 shown, the sending end may refer to the end where the video acquisition device such as a multi-camera is located. The multi-camera includes N cameras, namely, a 1-camera, a 2-camera,..., an N-camera, and N is greater than 1. This method may include the following steps:

[0071] First, for each camera, after the sending end obtains the original frame collected by this camera, it will obtain the video frame of this camera at a timestamp based on the original frame collected by this camera, and will also obtain the private data of this camera, and obtain the SEI data packet of this camera at this timestamp based on the private data of this camera.

[0072] Among them, the sending end refers to the end where the video acquisition device is located. The above timestamp may be a timestamp determined based on the acquisition moment of the original frame; in this embodiment, the timestamps corresponding to the video frames of each camera obtained based on the original frames of each camera obtained at the same moment are the same.

[0073] After that, the sending end associates the video frame of each camera at the above timestamp and the SEI data packet to obtain the video data corresponding to each camera at the above timestamp.

[0074] After the sending end obtains the video data corresponding to each camera at the above-mentioned timestamp, it will package the video data corresponding to each camera at the above-mentioned timestamp into the target video data corresponding to the above-mentioned timestamp based on the RTP format; and based on the UDP protocol and the pre-configured QoS policy, transmit the target video data corresponding to the above-mentioned timestamp to the receiving end.

[0075] After the receiving end receives the target video data corresponding to the above-mentioned timestamp, it will parse the target video data based on the SSRC carried by the target video data to obtain the SEI data packets and video frames of each camera at the above-mentioned timestamp; and convert the SEI data packets of each camera at the above-mentioned timestamp into the private data of each camera at the above-mentioned timestamp.

[0076] After the receiving end obtains the private data and video frames of each camera at the above-mentioned timestamp, it will perform video frame synthesis and rendering processing based on the private data and video frames of each camera at the above-mentioned timestamp.

[0077] So far, the description of the method provided in the embodiments of the present application is completed. Next, the video acquisition device provided in the embodiments of the present application will be described:

[0078] See Figure 5 , Figure 5 which is a schematic structural diagram of a video acquisition device provided in the embodiments of the present application. As Figure 5 shown, the video acquisition device 500 is equipped with N cameras 501 and a processor 502; N is greater than 1;

[0079] The camera 501 is used to collect the original frame; the original frame is used to generate the video frame of the camera;

[0080] The processor 502 is used to obtain the private data of each camera;

[0081] The processor 502 is further used to package the video data corresponding to each camera at the same timestamp into the target video data corresponding to the timestamp; wherein, the video data corresponding to any camera at the timestamp is obtained by associating the video frame of the camera at the timestamp and the SEI data packet of the camera at the timestamp; the SEI data packet is obtained based on the private data of the camera at the timestamp; the video data corresponding to any camera carries at least the SSRC; the video data corresponding to different cameras carries different SSRCs;

[0082] Transmit the target video data corresponding to the timestamp to the video receiving end, so that the video receiving end can obtain the private data of each camera at the timestamp and the video frame at the timestamp based on the SSRC carried by the target video data.

[0083] As an embodiment, the video data corresponding to any camera at this timestamp is determined through the following steps:

[0084] If the video frame of the camera at this timestamp is a key frame, and the key frame includes at least the following data packets: SPS data packet, PPS data packet, and at least one I-FU data packet, then: insert the SEI data packet into the following position in the video frame: the position after the PPS data packet and before the last I-FU data packet, to obtain the video data corresponding to the camera at this timestamp;

[0085] If the video frame of the camera at this timestamp is a non-key frame, then insert the SEI data packet into the following position in the video frame: the position before the last data packet, to obtain the video data corresponding to the camera at this timestamp.

[0086] As an embodiment, each data packet included in the video data corresponding to the same camera carries the same SSRC;

[0087] Each data packet included in the video data corresponding to different cameras carries different SSRCs;

[0088] In the target video data, the data packets in the video data corresponding to all cameras are organized together in the order of the sequence numbers, and the difference between adjacent sequence numbers is a preset value.

[0089] As an embodiment, the first data packet included in the target video data has a corresponding sequence number;

[0090] Among them, if the timestamp is the first timestamp in the video synchronization process, the sequence number of the first data packet is the initial value; if the timestamp is not the first timestamp in the video synchronization process, the sequence number of the first data packet is the sum of the sequence number of the last data packet included in the target video data corresponding to the previous timestamp and the preset value.

[0091] As an embodiment, in the target video data, the last video data carries a marker bit; the marker bit is the first specified value, which is used to indicate the end;

[0092] If other video data in the target video data also carries a marker bit, then this marker bit is set to the second specified value, which is used to indicate non-end.

[0093] Thus far, the Figure 5 structural description of the shown video acquisition device is completed.

[0094] For the implementation processes of the functions and roles of each component in the above video acquisition device, please refer to the implementation processes of the corresponding steps in the above method for details, and will not be elaborated here.

[0095] For the embodiments of the video acquisition device, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. Those of ordinary skill in the art can understand and implement them without creative efforts.

[0096] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the scope of protection of the present application.

Claims

1. A video synchronization method, characterized in that: The method is applied to a video acquisition device, wherein the video acquisition device is equipped with N cameras, where N is greater than 1; the method comprises: Packing the video data corresponding to each camera at the same timestamp into the target video data corresponding to the timestamp; wherein the video data corresponding to any camera at the timestamp is obtained by associating the video frame of the camera at the timestamp and the media supplementary enhancement information SEI data packet of the camera at the timestamp; the SEI data packet is obtained based on the private data of the camera at the timestamp; the video data corresponding to any camera at least carries a synchronization source identifier SSRC; the video data corresponding to different cameras carry different SSRCs; The target video data corresponding to the timestamp is transmitted to the video receiving end, so that the video receiving end obtains the private data of each camera at the timestamp and the video frame at the timestamp based on the SSRC carried by the target video data.

2. The method according to claim 1, characterized in that The video data corresponding to any camera at the timestamp is determined by the following steps: If the video frame of the camera at the timestamp is a key frame, and the key frame includes at least the following data packets: a sequence parameter set SPS data packet, a picture parameter set PPS data packet, and a network abstraction layer unit I-FU data packet of at least one key frame, then: The SEI data packet is inserted into the following position in the video frame: after the PPS data packet and before the last I-FU data packet, so as to obtain the video data corresponding to the camera at the timestamp.

3. The method according to claim 1, characterized in that: The video data corresponding to any camera at the timestamp is determined by the following steps: If the video frame of the camera at the timestamp is a non-key frame, the SEI data packet is inserted into the following position in the video frame: the position before the last data packet, so as to obtain the video data corresponding to the camera at the timestamp.

4. The method according to claim 2 or 3, characterized in that: Each data packet contained in the video data corresponding to the same camera carries the same SSRC; Each data packet contained in the video data corresponding to different cameras carries a different SSRC; In the target video data, data packets in the video data corresponding to all cameras are organized together in sequence, and the difference between adjacent sequence numbers is a preset value.

5. The method according to claim 1, characterized in that: The first data packet included in the target video data has a corresponding sequence number; Wherein, if the timestamp is the first timestamp in the video synchronization process, the sequence number of the first data packet is an initial value; If the timestamp is not the first timestamp in the video synchronization process, the sequence number of the first data packet is the sum of the sequence number of the last data packet included in the target video data corresponding to the previous timestamp and a preset value.

6. The method according to claim 1, characterized in that In the target video data, the last video data carries a marker bit; the marker bit is a first specified value, which is used to indicate the end; If other video data in the target video data also carries a marker bit, the marker bit is set to a second specified value to indicate non-end.

7. A video acquisition device, characterized in that: The video acquisition device is equipped with N cameras and a processor; N is greater than 1; The camera is used to collect original frames; the original frames are used to generate video frames of the camera; The processor is used to obtain private data of each camera; The processor is further used to package the video data corresponding to each camera at the same timestamp into the target video data corresponding to the timestamp; wherein the video data corresponding to any camera at the timestamp is obtained by associating the video frame of the camera at the timestamp and the media supplementary enhancement information SEI data packet of the camera at the timestamp; the SEI data packet is obtained based on the private data of the camera at the timestamp; the video data corresponding to any camera at least carries a synchronization source identifier SSRC; and the video data corresponding to different cameras carry different SSRCs; The target video data corresponding to the timestamp is transmitted to the video receiving end, so that the video receiving end obtains the private data of each camera at the timestamp and the video frame at the timestamp based on the SSRC carried by the target video data.

8. The video acquisition device according to claim 7, characterized in that: The video data corresponding to any camera at the timestamp is determined by the following steps: If the video frame of the camera at the timestamp is a key frame, and the key frame includes at least the following data packets: a sequence parameter set SPS data packet, a picture parameter set PPS data packet, and a network abstraction layer unit I-FU data packet of at least one key frame, then: insert the SEI data packet into the following position in the video frame: after the PPS data packet and before the last I-FU data packet, so as to obtain the video data corresponding to the camera at the timestamp; If the video frame of any camera at the timestamp is a non-key frame, the SEI data packet is inserted into the following position in the video frame: the position before the last data packet, so as to obtain the video data corresponding to the camera at the timestamp.

9. The video acquisition device according to claim 8, characterized in that: Each data packet contained in the video data corresponding to the same camera carries the same SSRC; Each data packet contained in the video data corresponding to different cameras carries a different SSRC; In the target video data, data packets in the video data corresponding to all cameras are organized together in sequence, and the difference between adjacent sequence numbers is a preset value.

10. The video acquisition device according to claim 7, characterized in that: The first data packet included in the target video data has a corresponding sequence number; Wherein, if the timestamp is the first timestamp in the video synchronization process, the sequence number of the first data packet is an initial value; if the timestamp is not the first timestamp in the video synchronization process, the sequence number of the first data packet is the sum of the sequence number of the last data packet contained in the target video data corresponding to the previous timestamp and a preset value; and / or, In the target video data, the last video data carries a marker bit; the marker bit is a first specified value, which is used to indicate the end; If other video data in the target video data also carries a marker bit, the marker bit is set to a second specified value to indicate non-end.

Citation Information

Patent Citations

  • Video frame synchronization method, system and device and readable storage medium

    CN113163222A

  • Video structured information storage method and system

    CN113360707A

  • TS stream time synchronization information insertion method, apparatus and device, and readable storage medium

    CN115866300A

  • Method and apparatus for video frame marking

    EP1995965A1

  • Immersive viewport dependent multiparty video communication

    WO2021074005A1

Cited By

  • Private frame framing method and device of code stream and computer equipment

    CN121567869A