Multi-camera spatio-temporal joint synchronous acquisition processing method, device and storage medium

By generating a video data packet queue and calculating the absolute UTC timestamp, the synchronization time point is determined, which solves the problems of high cost and complex wiring in multi-camera synchronization technology. It realizes simple and efficient spatiotemporal joint synchronous acquisition of multiple cameras, and improves system stability and synchronization accuracy.

CN119583732BActive Publication Date: 2026-04-14ORANGE LION SPORTS (ZHEJIANG) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ORANGE LION SPORTS (ZHEJIANG) CO LTD
Filing Date
2024-12-06
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing multi-camera synchronization technology solutions are costly, have complex wiring, and are not very adaptable. In particular, when multiple points are far apart, signal attenuation is severe, and engineering wiring is messy and disorderly.

Method used

By receiving data packets from each camera, a video data packet queue is generated, and the UTC absolute timestamp and display timestamp are calculated to determine the video synchronization time point. This enables spatiotemporal joint synchronous acquisition of multiple cameras, supports synchronous processing of audio data packets, and optimizes synchronization by using the synchronization vector and error exponent calculation on the host side, thereby reducing the need for external electronic components.

Benefits of technology

It enables spatiotemporal joint synchronous acquisition by multiple cameras, reduces hardware costs, simplifies wiring, improves system stability and reliability, and supports spatiotemporal collaboration and precise synchronization of multiple cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119583732B_ABST
    Figure CN119583732B_ABST
Patent Text Reader

Abstract

The application provides a multi-camera space-time joint synchronous acquisition processing method and device and a storage medium. The method comprises the following steps: receiving data packets uploaded by each camera and generating corresponding video data packet queues; acquiring the UTC absolute timestamp of the first frame video data packet of each queue and the PTS of each video data packet; calculating the incremental offset of the PTS of each video data packet relative to the PTS of the first frame video data packet; calculating the UTC absolute timestamp of each video data packet according to the UTC absolute timestamp of the first frame video data packet and the incremental offset corresponding to each video data packet; determining the video receiving time point with the optimal relative synchronization degree as the video synchronization time point according to the UTC absolute timestamp of the video data packet of each camera received at different video receiving time points, and generating a multi-camera based synchronous multi-view video sequence according to the video data packet corresponding to the video synchronization time point, so as to realize the space-time joint synchronous acquisition of multi-camera video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and information communication technology, and in particular to a multi-camera spatiotemporal joint synchronous acquisition and processing method, device and storage medium. Background Technology

[0002] Multi-camera systems targeting different physical locations provide fundamental capabilities for the construction of digital sports venues. Deploying multi-point, multi-camera devices at various physical locations within the sports venue helps capture the field conditions from all angles, especially supporting the synchronization of video capture by multiple cameras, which is crucial for capturing fast-paced ball sports. Multi-camera synchronization systems based on physical multiple points differ from multi-camera output stereo vision acquisition systems with shared substrates or supports, which are similar to single-point systems. The issues involved are more complex and therefore have significant research value.

[0003] The existing technologies for achieving multi-camera synchronization mainly include the following methods: For example, the main body uses multiple image acquisition cards connected in series to trigger the synchronization control between the host and multiple cameras, and a high-precision timer module is used to generate a high-precision clock signal for communication synchronization between devices; another example is to use a hardware synchronizer motherboard and a hardware synchronizer slave board to control and communicate to acquire trigger signals, driving the camera acquisition host to output multi-view video sequences.

[0004] Existing methods for achieving multi-camera synchronization have at least the following shortcomings: 1. The technical solutions employ a large number of electronic component modules, resulting in high overall costs; 2. For multi-camera systems with large physical spatial distances, the existing serial connection method for image acquisition cards, as well as the connection between the motherboard and slave board of the hardware synchronizer, is not very engineering-friendly for such situations. Long-distance wiring leads to severe signal attenuation, poor adaptability, and the space appears rather cluttered and disorderly, making the solution unsuitable for practical implementation. Summary of the Invention

[0005] In view of the above problems, the present invention is proposed to provide a multi-camera spatiotemporal joint synchronous acquisition and processing method, device and storage medium that solves or at least partially solves the above technical problems.

[0006] One aspect of the present invention provides a multi-camera spatiotemporal joint synchronous acquisition and processing method, the method comprising:

[0007] Receive data packets uploaded by each camera, including video data packets, and generate a video data packet queue corresponding to each camera;

[0008] Get the UTC absolute timestamp and display timestamp (PTS) of the first video data packet in each video data packet queue, as well as the PTS of each subsequent video data packet in the queue except for the first video data packet;

[0009] Calculate the incremental offset of the PTS of each subsequent video data packet to the PTS of the first frame video data packet, and calculate the UTC absolute timestamp of the corresponding subsequent video data packet based on the UTC absolute timestamp of the first frame video data packet and the incremental offset of each subsequent video data packet.

[0010] The relative synchronization of video data acquired by each camera is determined based on the UTC absolute timestamp of the video data packets received by each camera at different video reception times, and the video reception time with the best relative synchronization is taken as the video synchronization time point.

[0011] The video data packets received from each camera at the video synchronization time point are used as the first frame video data packet to be output by the corresponding camera to generate a synchronized multi-view video sequence based on multiple cameras.

[0012] Furthermore, the data packet also includes audio data packets;

[0013] The method further includes:

[0014] Generate audio data packet queues corresponding to each camera;

[0015] Obtain the PTS of each audio data packet in the audio data packet queue, calculate the incremental offset of the PTS of each audio data packet to the PTS of the first frame video data packet, and calculate the UTC absolute timestamp of each video data packet based on the UTC absolute timestamp of the first frame video data packet and the incremental offset corresponding to each audio data packet.

[0016] Accordingly, the step of generating a multi-camera synchronized multi-view video sequence by using the video data packets received from each camera at the video synchronization time point as the first frame video data packet to be output by the corresponding camera includes:

[0017] Obtain the target audio data packet that is synchronized with the video data packet received from each camera at the video synchronization time point. For the video data packet queue and audio data packet queue corresponding to each camera, generate a synchronized multi-view video sequence based on multiple cameras, with the video data packet received at the video synchronization time point and the target audio data packet synchronized with it as the starting positions.

[0018] Furthermore, after calculating the UTC absolute timestamp of the corresponding subsequent video data packet based on the UTC absolute timestamp of the first frame video data packet and the incrementing offset corresponding to each subsequent video data packet, the method further includes:

[0019] The absolute UTC timestamp and PTS of each data packet are bound to the corresponding data packet, and the host UTC timestamp when the host receives each data packet is updated to the video data packet queue as the video reception time of the corresponding data packet.

[0020] Furthermore, the step of determining the relative synchronization degree of video data acquired by each camera based on the UTC absolute timestamps of the video data packets received by each camera at different video reception times, and taking the video reception time with the optimal relative synchronization degree as the video synchronization time point, includes:

[0021] Construct a synchronization vector T = [t1, t2, ..., t3] using the UTC absolute timestamps of the video data packets uploaded by each camera received by the host at the same video reception time. n ], where t i It is the absolute UTC timestamp of the video data packet captured by the i-th camera;

[0022] The relative synchronization error index α of each camera's video data at each video reception time is calculated based on the synchronization vector T at different video reception times. The formula A for calculating α is:

[0023]

[0024] Where t is the video reception time corresponding to the synchronization vector T;

[0025] Find the video reception time corresponding to the minimum value of the relative synchronization error index, and use the current video reception time as the video synchronization time point.

[0026] Furthermore, before calculating the relative synchronization error exponent α of the video data acquired by each camera at each video reception time based on the synchronization vector T at different video reception times, the method further includes:

[0027] The asynchronous deviation index β of the video data acquired by each camera at each video reception time is calculated based on the synchronization vector T corresponding to different video reception times. The formula B for calculating β is:

[0028] B = max i |t i -t|

[0029] When the asynchronous deviation index β corresponding to a certain video reception time is detected to be less than the preset system nominal error, the relative synchronization error index α of the video data acquired by each camera at each video reception time is calculated for the synchronization vector T corresponding to different subsequent video reception times, starting from the current video reception time.

[0030] Furthermore, the acquisition of the target audio data packet synchronized with the video data packets received from each camera at the video synchronization time point includes:

[0031] Search the video data packet queue corresponding to each camera to obtain the host UTC timestamp and PTS of the target video data packet corresponding to the current camera. The target video data packet is the video data packet of the corresponding camera received at the video synchronization time point.

[0032] Search the audio data packet queue of each camera to obtain the target audio data packet whose host UTC timestamp is equal to or greater than the host UTC timestamp of the target video data packet corresponding to the current camera.

[0033] Furthermore, the method also includes:

[0034] After searching the audio data packet queue of each camera to obtain the target audio data packet whose host UTC timestamp is equal to or greater than the host UTC timestamp of the target video data packet corresponding to the current camera, it is determined whether the PTS of the target video data packet corresponding to each camera is corrected and rewritten when it is output as the first frame of video data.

[0035] If the target video data packet is corrected and rewritten when it is output as the first frame of video data, then the PTS of the target audio data packet is updated synchronously according to the offset time difference between the target audio data packet and the target video data packet.

[0036] For each camera's corresponding video data packet queue and audio data packet queue, the PTS of each data packet is synchronously updated one by one towards the tail of its respective queue based on the PTS of the target video data packet and the target audio data packet.

[0037] Furthermore, receiving data packets uploaded by each camera includes:

[0038] Before the host sends acquisition commands to each camera, video preloading processing is performed on each camera to receive data packets uploaded by each camera; and / or

[0039] Send acquisition commands to each camera to receive data packets uploaded by each camera.

[0040] In another aspect, the present invention provides a multi-camera spatiotemporal joint synchronous acquisition and processing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the multi-camera spatiotemporal joint synchronous acquisition and processing method described above.

[0041] In another aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multi-camera spatiotemporal joint synchronous acquisition and processing method described above.

[0042] The multi-camera spatiotemporal joint synchronous acquisition and processing method, device, and storage medium provided in this invention can achieve multi-camera spatiotemporal joint synchronous acquisition without the need for other external electronic devices. This not only saves on the overall cost of hardware equipment but also simplifies the wiring of multi-camera projects based on multiple physical points, making it highly operable. Moreover, this invention fully embodies the spatiotemporal collaboration and joint synchronization between the host and multiple cameras, supporting the updating and determination of video synchronization time points by detecting and calculating the relative synchronicity of video data acquired by each camera. This effectively ensures the accuracy of multi-camera spatiotemporal joint synchronization and improves system stability and reliability.

[0043] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0044] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. In the drawings:

[0045] Figure 1 This is a flowchart of a multi-camera spatiotemporal joint synchronous acquisition and processing method according to an embodiment of the present invention;

[0046] Figure 2 This is a flowchart of a multi-camera spatiotemporal joint synchronous acquisition and processing method according to another embodiment of the present invention. Detailed Implementation

[0047] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0048] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the meaning consistent with their meaning in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined.

[0049] The image sensor used in the camera involved in this invention is not limited to charge-coupled devices (CCD), complementary metal-oxide-semiconductor (CMOS), linear scanning sensors, infrared sensors, thermal imagers, 3D depth sensors, X-ray sensors, and ultraviolet sensors.

[0050] This invention relates to the connection between a camera and a host, not limited to interfaces such as USB, MIPI, HDMI, GigE, Ethernet, Wi-Fi, SDI, VGA, AV, Thunderbolt, and FireWire (IEEE 1394).

[0051] The multi-camera spatiotemporal joint synchronous acquisition and processing method provided by this invention is applicable to multi-camera synchronous acquisition systems. The real-time system includes at least two or more cameras and a host terminal. The host terminal sends acquisition commands to each camera. Upon receiving the acquisition commands, the cameras begin acquiring live video data and uploading it to the host. The host acquires the video data packets from the cameras and performs spatiotemporal joint synchronization and alignment processing to generate an aligned multi-viewpoint video sequence. The multi-camera spatiotemporal joint synchronous acquisition and processing method provided by the embodiments of this invention will be described in detail below.

[0052] like Figure 1 As shown in the figure, the multi-camera spatiotemporal joint synchronous acquisition and processing method proposed in this embodiment of the invention includes the following steps:

[0053] S11. Receive data packets uploaded by each camera, the data packets including video data packets, and generate a video data packet queue corresponding to each camera.

[0054] In this embodiment, after receiving the video streams transmitted from each camera, the host generates a data packet queue and obtains the parameter characteristics of the video stream, including video resolution, frame rate, encoder type, and time base.

[0055] S12. Obtain the UTC absolute timestamp and display timestamp (PTS) of the first video data packet in each video data packet queue, as well as the PTS of each subsequent video data packet in the queue except for the first video data packet.

[0056] S13. Calculate the incremental offset of the PTS of each subsequent video data packet to the PTS of the first frame video data packet, and calculate the UTC absolute timestamp of the corresponding subsequent video data packet based on the UTC absolute timestamp of the first frame video data packet and the incremental offset of each subsequent video data packet.

[0057] In this embodiment, the host calculates and generates the UTC absolute timestamp (utc_timestamp) and the presentation time stamp (PTS) of the first frame video data packet based on the received video data packets from the camera, as well as the PTS corresponding to each video data packet. For each video data packet, based on the UTC absolute timestamp of the first frame video, the PTS of the first frame video data packet, and the PTS corresponding to each data packet, the host calculates the incremental offset of the PTS of each data packet relative to the PTS of the first frame video. The host then calculates the UTC absolute timestamp of each data packet by superimposing this offset with the UTC absolute timestamp of the first frame video.

[0058] Furthermore, the method also includes binding the UTC absolute timestamp and PTS of each data packet with the corresponding data packet. Moreover, the host records the local host UTC timestamp (local_utc_timestamp) of each acquired data packet and updates the video data packet queue with the host UTC timestamp at the time the host receives each data packet, using it as the video reception time of the corresponding data packet.

[0059] The specific logic for calculating the absolute UTC timestamp of video data packets is as follows:

[0060] Obtain the absolute UTC timestamp of the first frame video data packet:

[0061] first_frame_utc=first_video_packet.utc_timestamp;

[0062] Obtain the PTS of the first frame video data packet:

[0063] first_frame_pts=first_video_packet.pts;

[0064] Iterate through all video data packets and calculate the incrementing offset of the current data packet's PTS relative to the PTS of the first frame of video:

[0065] delta_video_pts=packet.pts-first_frame_pts;

[0066] Wherein, video_time_unit is the unit of video clock frequency, and its value is equivalent to video time base, with a typical value of 1 / 90000;

[0067] Calculate the absolute UTC timestamp of the current data packet, packet_utc:

[0068] packet_utc = first_frame_utc + delta_video_pts * video_time_unit, where video_time_unit is a conversion factor, such as converting from PTS units to UTC units.

[0069] Bind the calculated UTC absolute timestamp and PTS to the data packet:

[0070] packet.utc_timestamp=packet_utc

[0071] Update to the video data packet queue:

[0072] video_packets_queue.append(packet).

[0073] S14. Determine the relative synchronization degree of the video data acquired by each camera based on the UTC absolute timestamp of the video data packets received by each camera at different video reception times, and take the video reception time with the best relative synchronization degree as the video synchronization time point.

[0074] S15. Generate a multi-camera synchronized multi-view video sequence by using the video data packets received from each camera at the video synchronization time point as the first frame video data packet to be output by the corresponding camera.

[0075] The multi-camera spatiotemporal joint synchronous acquisition and processing method provided in this invention can achieve multi-camera spatiotemporal joint synchronous acquisition without the need for additional external electronic devices. This not only saves on the overall cost of hardware equipment but also simplifies the wiring of multi-camera projects based on multiple physical points, making it highly operable. Furthermore, this invention fully embodies the spatiotemporal collaboration and joint synchronization between the host and multiple cameras, supporting the updating and determination of video synchronization time points by detecting and calculating the relative synchronicity of video data acquired by each camera. This effectively ensures the accuracy of multi-camera spatiotemporal joint synchronization and improves system stability and reliability.

[0076] In this embodiment of the invention, receiving video data packets uploaded by each camera can be achieved by preloading the video data packets of each camera before the host sends the acquisition command to each camera; alternatively, the acquisition command can be directly sent to each camera to receive the video data packets uploaded by each camera.

[0077] Since the first frame of video data packets uploaded by cameras is often large after compression by the video encoder, this invention employs preloading processing before officially acquiring video data to avoid network congestion when receiving the first frame, thereby improving the first frame experience. Specifically, before the host sends acquisition commands to each camera, the host acquires the video data packets from each camera through a video preloading process, pre-fetching the video header data of each camera to pre-determine the video synchronization time point. The duration 'd' of the preloaded video is a floating-point number greater than 0, in seconds. During the preloading phase, the host receives a certain amount of video data packets from each camera, forming a corresponding video data packet queue. To reduce computational overhead during the preloading phase, decoding of the data packets is not performed first; only the packet header information is parsed, including at least the packet type (PacketType), packet index, and presentation time stamp (PTS), and this information is bound to the corresponding data packet and updated in the video data packet queue of the preloading phase.

[0078] Furthermore, this invention pre-synchronizes the Universal Time Coordinated (UTC) clocks of each camera with a system reference through a synchronized time synchronization method, such as using Precision Time Protocol (PTP) or Network Time Protocol (NTP) to synchronize the local system clocks of each camera. NTP provides millisecond accuracy, while PTP can provide accuracy up to submicron levels. More preferably, this invention synchronizes the host and each camera with the same UTC server to coordinate the UTC clocks. In this way, the host and each camera have a local clock synchronized with the global clock. Thus, data communication between the host and the cameras is based on a global clock, thereby ensuring that the session is in sync.

[0079] In this embodiment of the invention, step S14, which determines the relative synchronization degree of video data acquired by each camera based on the UTC absolute timestamps of the video data packets received by each camera at different video reception times, and takes the video reception time with the optimal relative synchronization degree as the video synchronization time point, includes:

[0080] Construct a synchronization vector T = [t1, t2, ..., t3] using the UTC absolute timestamps of the video data packets uploaded by each camera received by the host at the same video reception time. n ], where t i It is the absolute UTC timestamp of the video data packet captured by the i-th camera;

[0081] The relative synchronization error index α of each camera's video data at each video reception time is calculated based on the synchronization vector T at different video reception times. The formula A for calculating α is:

[0082]

[0083] Where t is the video reception time corresponding to the synchronization vector T;

[0084] Find the video reception time corresponding to the minimum value of the relative synchronization error index, and use the current video reception time as the video synchronization time point.

[0085] Furthermore, before calculating the relative synchronization error index α of the video data acquired by each camera at each video reception time based on the synchronization vector T at different video reception times, the asynchronous deviation index β of the video data acquired by each camera at each video reception time can be calculated based on the synchronization vector T corresponding to different video reception times. The formula B for calculating β is:

[0086] B = max i |t i -t|

[0087] When the asynchronous deviation index β corresponding to a certain video reception time is detected to be less than the preset system nominal error, the relative synchronization error index α of the video data acquired by each camera at each video reception time is calculated for the synchronization vector T corresponding to different subsequent video reception times, starting from the current video reception time.

[0088] In this embodiment, multi-camera relative synchronization is used to measure the temporal alignment of frames captured by all cameras in a multi-camera system. Assume there are n cameras, each capturing a frame at time t, and each frame has a UTC absolute timestamp. We can define a vector T = [t1, t2, ..., t...]. n ], where t i This is the absolute UTC timestamp of the frame captured by the i-th camera. Ideally, if all cameras are perfectly synchronized, then all elements of this vector should be equal, i.e., t1 = t2 = ... = t n However, in practical applications, this ideal situation is difficult to achieve due to hardware limitations, network latency, and processing delays. This invention uses the optimal synchronicity equation A to measure the relative synchronicity of the camera under optimal conditions, characterizing the Euclidean distance between perception vectors T and t. This invention also uses the worst-case desynchronization equation B to measure the worst-case desynchronization of the camera, characterizing the maximum deviation of a sample element of vector T from its true value.

[0089]

[0090] B = maxi |t i -t|

[0091] As shown in the equation above, A, B, t i t is the utc_timestamp of the frame captured by the i-th camera, and t is the local UTC timestamp corresponding to the data packet received by the host from the camera.

[0092] The worst-case nominal error of desynchronization allowed by the multi-camera system is defined as σ. Optionally, when the camera video stream frame rate is 120Hz, σ is 0.005 seconds, or 5 milliseconds precision; when the camera video stream frame rate is 60Hz, σ is 0.010 seconds, or 10 milliseconds precision; and when the camera video stream frame rate is 30Hz, σ is 0.020 seconds, or 20 milliseconds precision.

[0093] In practical applications, the need for relative synchronization error index detection can be determined through conditional judgment. For example, multi-camera systems require spatiotemporal joint synchronization processing during initialization, and spatiotemporal joint synchronization processing is also required if desynchronization occurs during acquisition. Specific conditions can be set such that relative synchronization error index detection is performed during system initialization or when the β value calculated by the worst-case desynchronization index module of the multi-camera system does not reach the system's nominal error. The desynchronization deviation index β is calculated based on equation B. When β < σ, the relative synchronization error index α is calculated based on equation A. The specific calculation of the relative synchronization error index involves traversing the video data packet queue from newest to oldest (i.e., from the tail of the queue to the head of the queue) to obtain the calculated value α of equation A at a certain moment (i.e., under the condition of a similarity value t), which minimizes the spatial correlation synchronization error of the joint multi-camera system, i.e., α. min This corresponds to the optimal synchronization point for each camera. Thus, this synchronization point is both the alignment point of the physical objective world time for each camera and the first frame of video processed by the camera synchronization logic.

[0094] Furthermore, the present invention can also define a period for normalized relative synchronization error index detection, which is used to start normalized relative synchronization error index detection calculation and feedback at one period. This can effectively eliminate the impact of camera hardware processing and network transmission jitter, and further improve the robustness and reliability of multi-camera system synchronization performance.

[0095] This invention supports real-time detection, calculation, and feedback closed-loop of the optimal synchronization index and worst asynchrony index of multiple cameras, and proposes an anti-shake processing method. It features high system stability and reliability, effectively ensuring the accuracy of spatiotemporal joint synchronization of multiple cameras and improving system stability and reliability.

[0096] For multi-camera systems that also include audio options, the present invention provides another embodiment for detailed description. (Refer to...) Figure 2 As shown in the figure, the multi-camera spatiotemporal joint synchronous acquisition and processing method proposed in this embodiment of the invention includes the following steps:

[0097] S21. Receive video data packets and audio data packets uploaded by each camera, and generate video data packet queues and audio data packet queues corresponding to each camera.

[0098] After receiving video and audio data packets from each camera, the host generates video and audio data packet queues respectively, and obtains the parameter characteristics of the audio and video streams. The video parameter characteristics include video resolution, frame rate, encoder type and time base, while the audio parameter characteristics include audio sampling rate, time base, encoder type, number of channels and sample format.

[0099] S22. Obtain the UTC absolute timestamp and display timestamp (PTS) of the first video data packet in each video data packet queue, as well as the PTS of each subsequent video data packet and audio data packet in the queue, excluding the first video data packet.

[0100] S23. Calculate the incremental offset of the PTS of each subsequent video data packet and audio data packet corresponding to the PTS of the first frame video data packet, and calculate the corresponding UTC absolute timestamp of the subsequent video data packet and audio data packet based on the UTC absolute timestamp of the first frame video data packet and the incremental offset of each subsequent video data packet and audio data packet.

[0101] In this embodiment, the host calculates and generates the UTC absolute timestamp (utc_timestamp) and the presentation time stamp (PTS) of the first frame video data packet based on the received video data packets from the camera, as well as the PTS corresponding to each video and audio data packet. For each video or audio data packet, based on the UTC absolute timestamp of the first frame video, the PTS of the first frame video data packet, and the PTS corresponding to each data packet, the host calculates the incremental offset of the PTS of each data packet relative to the PTS of the first frame video. The UTC absolute timestamp of each data packet is then calculated by superimposing this incremental offset with the UTC absolute timestamp of the first frame video.

[0102] Furthermore, the method also includes binding the UTC absolute timestamp and PTS of each data and audio data packet with the corresponding data packet. Moreover, the host records the local host UTC timestamp (local_utc_timestamp) of each acquired data packet and updates the video data packet queue with the host UTC timestamp at the time the host receives each data packet, using it as the video reception time of the corresponding data packet.

[0103] The specific logic for calculating the absolute UTC timestamps of video and audio data packets is as follows:

[0104] Obtain the absolute UTC timestamp of the first frame video data packet:

[0105] first_frame_utc=first_video_packet.utc_timestamp;

[0106] Obtain the PTS of the first frame video data packet:

[0107] first_frame_pts=first_video_packet.pts;

[0108] Iterate through all video data packets and calculate the incrementing offset of the current data packet's PTS relative to the PTS of the first frame of video:

[0109] delta_video_pts=packet.pts-first_frame_pts;

[0110] Wherein, video_time_unit is the unit of video clock frequency, and its value is equivalent to video time base, with a typical value of 1 / 90000;

[0111] Calculate the absolute UTC timestamp of the current data packet, packet_utc:

[0112] packet_utc = first_frame_utc + delta_video_pts * video_time_unit, where video_time_unit is a conversion factor, such as converting from PTS units to UTC units.

[0113] Bind the calculated UTC absolute timestamp and PTS to the data packet:

[0114] packet.utc_timestamp=packet_utc

[0115] Update to the video data packet queue:

[0116] video_packets_queue.append(packet)

[0117] Iterate through all audio data packets and calculate the elapsed time offset of the current data packet's PTS relative to the PTS of the first video frame, in seconds:

[0118] since_1st_packet=packet.pts*audio_time_unit-first_frame_pts*video_time_unit;

[0119] Among them, the audio clock frequency audio_time_unit is equivalent to the audio time base. Its typical value is the reciprocal of the audio sampling rate. For example, if the sampling rate is 16000 samples / second, then it is 1 / 16000.

[0120] Calculate the absolute UTC timestamp of the current data packet:

[0121] packet_utc=first_frame_utc+since_1st_packet;

[0122] Bind the calculated UTC absolute timestamp and PTS to the data packet:

[0123] packet.utc_timestamp=packet_utc;

[0124] Update to the audio packet queue: audio_packets_queue.append(packet).

[0125] S24. Determine the relative synchronization degree of the video data acquired by each camera based on the UTC absolute timestamp of the video data packets received by each camera at different video reception times, and take the video reception time with the best relative synchronization degree as the video synchronization time point.

[0126] S25. Obtain the target audio data packet that is synchronized with the video data packet received from each camera at the video synchronization time point. For the video data packet queue and audio data packet queue corresponding to each camera, generate a multi-camera synchronized multi-view video sequence with the video data packet received at the video synchronization time point and the target audio data packet synchronized with it as the starting position.

[0127] In this embodiment of the invention, step S24, which determines the relative synchronization degree of video data acquired by each camera based on the UTC absolute timestamps of the video data packets received from each camera at different video reception times, and takes the video reception time with the optimal relative synchronization degree as the video synchronization time point, includes:

[0128] Construct a synchronization vector T = [t1, t2, ..., t3] using the UTC absolute timestamps of the video data packets uploaded by each camera received by the host at the same video reception time. n ], where t i It is the absolute UTC timestamp of the video data packet captured by the i-th camera;

[0129] The relative synchronization error index α of each camera's video data at each video reception time is calculated based on the synchronization vector T at different video reception times. The formula A for calculating α is:

[0130]

[0131] Where t is the video reception time corresponding to the synchronization vector T;

[0132] Find the video reception time corresponding to the minimum value of the relative synchronization error index, and use the current video reception time as the video synchronization time point.

[0133] Furthermore, before calculating the relative synchronization error exponent α of the video data acquired by each camera at each video reception time based on the synchronization vector T at different video reception times, the method further includes:

[0134] The asynchronous deviation index β of the video data acquired by each camera at each video reception time is calculated based on the synchronization vector T corresponding to different video reception times. The formula B for calculating β is:

[0135] B = max i |t i -t|

[0136] When the asynchronous deviation index β corresponding to a certain video reception time is detected to be less than the preset system nominal error, the relative synchronization error index α of the video data acquired by each camera at each subsequent video reception time is calculated, starting from the current video reception time and applied to the synchronization vector T corresponding to different video reception times. Otherwise, the generation of a multi-camera synchronized multi-view video sequence can continue, using the target video data packet corresponding to the historical video synchronization time point as the first frame video data packet to be output.

[0137] In this embodiment of the invention, step S25, obtaining the target audio data packet synchronized with the video data packets received from each camera at the video synchronization time point, specifically includes: searching the video data packet queue corresponding to each camera to obtain the host UTC timestamp and PTS of the target video data packet corresponding to the current camera, wherein the target video data packet is the video data packet of the corresponding camera received at the video synchronization time point; and searching the audio data packet queue of each camera to obtain the target audio data packet in the audio data packet queue whose host UTC timestamp is equal to or greater than the host UTC timestamp of the target video data packet corresponding to the current camera.

[0138] To achieve synchronization in the camera's temporal domain, this invention employs an audio-visual synchronization technique for processing the audio and video streams output by a single camera based on the video synchronization time point, i.e., the optimal synchronization point. To achieve optimal audio-visual synchronization, it is necessary to calculate the audio data packet corresponding to the video synchronization time point. The specific implementation process is as follows: The host UTC timestamp (local_utc_timestamp) and pts of the target video data packet corresponding to the video synchronization time point are obtained from the video data packet queue. Then, in the audio data packet queue of the same camera, an audio data packet whose host UTC timestamp (local_utc_timestamp) is greater than or equal to the host UTC timestamp (local_utc_timestamp) of the target video data packet is searched, and this audio data packet is used as the target audio data packet.

[0139] To achieve optimal audio-visual synchronization, further steps are taken: after searching the audio data packet queue of each camera to obtain target audio data packets whose host UTC timestamp is equal to or greater than the host UTC timestamp of the target video data packet corresponding to the current camera, it is determined whether the PTS of the target video data packets corresponding to each camera is corrected and rewritten when output as the first frame of video data. If the PTS of the target video data packets is corrected and rewritten when output as the first frame of video data, the PTS of the target audio data packets is synchronously updated according to the offset time difference between the target audio data packets and the target video data packets. Then, for the video data packet queue and audio data packet queue corresponding to each camera, the PTS of each data packet is synchronously updated one by one towards the tail of its respective queue according to the PTS of the target video data packets and the target audio data packets. If the PTS of the target video data packets is not corrected and rewritten when output as the first frame of video data, the PTS of the target audio data packets calculated from the search is used.

[0140] Calculate the time difference between the target audio data packet and the target video data packet:

[0141] audio_elapsed_since_video_sync=audio_sync.pts*audio_time_unit–video_sync.pts*video_time_unit;

[0142] With the video baseline PTS set to 0, calculate and update the audio synchronization point PTS to achieve optimal audio-visual synchronization:

[0143] audio.sync.pts=round(audio_elapsed_since_video_sync / audio_time_unit);

[0144] Where: audio_sync is the target audio data packet, audio_sync.pts is the PTS of the target audio data packet, video_sync.pts is the PTS of the video packet at the optimal synchronization point, audio_elapsed_since_video_sync is the natural time elapsed value of the audio synchronization point based on the video synchronization point, in seconds; audio_time_unit is the audio clock frequency, and video_time_unit is the video clock frequency.

[0145] Accordingly, a multi-camera synchronized multi-view video sequence is generated, starting with the video data packets received at the video synchronization time point and the target audio data packets synchronized with them. This includes processing the audio and video data packets in the queue into corresponding audio and video frames according to the timeline corresponding to their respective PTSs, starting with the video data packets received at the video synchronization time point and the target audio data packets synchronized with them, to generate a multi-camera synchronized multi-view video sequence. This invention updates the PTS corresponding to each data packet sequentially towards the tail of its respective buffer queue based on the optimal video and audio synchronization point positions. Then, based on the multiple cameras and the timeline, it restores the audio and video data packets of each camera queue into corresponding audio and video frames, thereby generating a multi-camera synchronized multi-view video sequence.

[0146] In one embodiment of the present invention, the camera is a network camera, and the transmitted data is a byte stream data packet compressed by an audio or video encoder. In this case, the host performs packet assembly processing and decoding processing of the data packets according to the corresponding protocol standard and the relevant audio or video compression standard. Preferably, for video streams transmitted by the camera, such as ITU-T H.264, ITU-TH.265, and AOMedia Video 1 standard protocols, the host performs hardware video decoding processing according to its own hardware capabilities to reduce processing latency and improve performance. For example, Nvidia video decoding capabilities can be used.

[0147] The multi-camera spatiotemporal joint synchronous acquisition and processing method provided by this invention, if the multi-camera acquisition system is configured not to generate an audio stream, this invention only outputs the sequence of pure video portions from multiple viewpoints; it also includes cases where the generated multi-viewpoint video sequence is simply a pass-through bypass of the original individual camera stream portions, in which case this invention outputs a composite audio and video output multi-viewpoint video sequence; it also includes cases where the generated multi-viewpoint video needs to be rewritten or transcoded, such as when correction processing or overlaying of logos or other modifications are required on the original video, in which case corresponding video processing is performed on the original video before recombining and outputting a multi-viewpoint video sequence.

[0148] Preferably, in this embodiment of the invention, when the host performs video and audio processing, such as camera distortion correction, audio or video resampling, and video format conversion, GPU CUDA / OpenCL technology is preferentially used to release or reduce the computing pressure and bottleneck of the host CPU. Simultaneously, the host can support video processing capabilities at ultra-high frame rates of 60Hz, 120Hz, or higher for multiple cameras. Therefore, the system has the ability to expand the number of cameras according to the host's performance configuration capabilities. This is beneficial for adapting the technical solution of this invention to sites of different sizes by selecting and assembling different numbers of cameras.

[0149] The multi-camera spatiotemporal joint synchronous acquisition and processing method proposed in this invention only includes multiple cameras and a host unit in the hardware system, without involving unnecessary cost expenditures. It can realize multi-camera spatiotemporal joint synchronous acquisition without setting up other external electronic devices. It can not only save hardware equipment costs, but also make the wiring of multi-camera based on physical multiple points simpler and more operable.

[0150] Furthermore, this invention fully embodies the characteristics of spatiotemporal collaboration and joint synchronization between the host and multiple cameras, achieving high accuracy and conforming to system specifications. It supports real-time detection, calculation, and feedback closed-loop of optimal and worst-case asynchrony indices related to multiple cameras, proposes an anti-jitter processing method, and features high system stability and reliability. It adaptively performs multi-camera synchronous acquisition and processing at low, high, or ultra-high frame rates, and has the ability to expand the number of multiple cameras according to the host's performance configuration capabilities. This is beneficial for the invention to be implemented in sites of different spatial sizes.

[0151] for Figure 2 Regarding the illustrated embodiment, since it is related to Figure 1 The embodiments shown are basically similar, so the description is relatively simple. For relevant details, please refer to [link / reference]. Figure 1 The description of the embodiments shown is only a partial one and has the corresponding technical effects.

[0152] For the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0153] Another embodiment of the present invention provides a multi-camera spatiotemporal joint synchronous acquisition and processing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the methods provided in the above-described method embodiments. For example... Figure 1 Steps S11 to S15 shown, or, Figure 2 Steps S21 to S25 shown

[0154] Another embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods provided in the above-described method embodiments.

[0155] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as ROM, RAM, magnetic disk, or optical disk.

[0156] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0157] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, any of the claimed embodiments can be used in any combination.

[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-camera spatiotemporal joint synchronous acquisition and processing method, characterized in that, The method includes: Receive data packets uploaded by each camera, including video data packets, and generate a video data packet queue corresponding to each camera; Get the UTC absolute timestamp and display timestamp (PTS) of the first video data packet in each video data packet queue, as well as the PTS of each subsequent video data packet in the queue except for the first video data packet; Calculate the incremental offset of the PTS of each subsequent video data packet to the PTS of the first frame video data packet, and calculate the UTC absolute timestamp of the corresponding subsequent video data packet based on the UTC absolute timestamp of the first frame video data packet and the incremental offset of each subsequent video data packet. The relative synchronization of video data acquired by each camera is determined based on the UTC absolute timestamp of the video data packets received by each camera at different video reception times, and the video reception time with the best relative synchronization is taken as the video synchronization time point. The video data packets received from each camera at the video synchronization time point are used as the first frame video data packet to be output by the corresponding camera to generate a synchronized multi-view video sequence based on multiple cameras.

2. The method according to claim 1, characterized in that, The data packet also includes audio data packets; The method further includes: Generate audio data packet queues corresponding to each camera; Get the PTS of each audio data packet in the audio data packet queue, calculate the incremental offset of the PTS of each audio data packet to the PTS of the first frame video data packet, and calculate the UTC absolute timestamp of each audio data packet based on the UTC absolute timestamp of the first frame video data packet and the incremental offset of each audio data packet. The step of generating a multi-camera synchronized multi-view video sequence by using the video data packets received from each camera at the video synchronization time point as the first frame video data packet to be output by the corresponding camera includes: Obtain the target audio data packet that is synchronized with the video data packet received from each camera at the video synchronization time point. For the video data packet queue and audio data packet queue corresponding to each camera, generate a synchronized multi-view video sequence based on multiple cameras, with the video data packet received at the video synchronization time point and the target audio data packet synchronized with it as the starting positions.

3. The method according to claim 1 or 2, characterized in that, After calculating the UTC absolute timestamp of the corresponding subsequent video data packet based on the UTC absolute timestamp of the first frame video data packet and the incrementing offset corresponding to each subsequent video data packet, the method further includes: The absolute UTC timestamp and PTS of each data packet are bound to the corresponding data packet, and the host UTC timestamp when the host receives each data packet is updated to the video data packet queue as the video reception time of the corresponding data packet.

4. The method according to claim 1 or 2, characterized in that, The step of determining the relative synchronization degree of video data acquired by each camera based on the UTC absolute timestamps of the video data packets received from each camera at different video reception times, and taking the video reception time with the optimal relative synchronization degree as the video synchronization time point, includes: A synchronization vector T = [t1, t2, ..., t] is constructed by using the UTC absolute timestamps of the video data packets uploaded by each camera received by the host at the same video reception time. n ], where t i It is the absolute UTC timestamp of the video data packet captured by the i-th camera; The relative synchronization error index α of each camera's video data at each video reception time is calculated based on the synchronization vector T at different video reception times. The formula A for calculating α is: , Where t is the video reception time corresponding to the synchronization vector T; Find the video reception time corresponding to the minimum value of the relative synchronization error index, and use the current video reception time as the video synchronization time point.

5. The method according to claim 4, characterized in that, Before calculating the relative synchronization error exponent α of the video data acquired by each camera at each video reception time based on the synchronization vector T at different video reception times, the method further includes: The asynchronous deviation index β of the video data acquired by each camera at each video reception time is calculated based on the synchronization vector T corresponding to different video reception times. The formula B for calculating β is: , When the asynchronous deviation index β corresponding to a certain video reception time is detected to be less than the preset system nominal error, the relative synchronization error index α of the video data acquired by each camera at each video reception time is calculated for the synchronization vector T corresponding to different subsequent video reception times, starting from the current video reception time.

6. The method according to claim 2, characterized in that, The target audio data packet synchronized with the video data packets received from each camera at the video synchronization time point includes: Search the video data packet queue corresponding to each camera to obtain the host UTC timestamp and PTS of the target video data packet corresponding to the current camera. The target video data packet is the video data packet of the corresponding camera received at the video synchronization time point. Search the audio data packet queue of each camera to obtain the target audio data packet whose host UTC timestamp is equal to or greater than the host UTC timestamp of the target video data packet corresponding to the current camera.

7. The method according to claim 6, characterized in that, The method further includes: After searching the audio data packet queue of each camera to obtain the target audio data packet whose host UTC timestamp is equal to or greater than the host UTC timestamp of the target video data packet corresponding to the current camera, it is determined whether the PTS of the target video data packet corresponding to each camera is corrected and rewritten when it is output as the first frame of video data. If the target video data packet is corrected and rewritten when it is output as the first frame of video data, then the PTS of the target audio data packet is updated synchronously according to the offset time difference between the target audio data packet and the target video data packet. For each camera's corresponding video data packet queue and audio data packet queue, the PTS of each data packet is synchronously updated one by one towards the tail of its respective queue based on the PTS of the target video data packet and the target audio data packet.

8. The method according to claim 1 or 2, characterized in that, The receiving of data packets uploaded by each camera includes: Before the host sends acquisition commands to each camera, video preloading processing is performed on each camera to receive data packets uploaded by each camera; and / or Send acquisition commands to each camera to receive data packets uploaded by each camera.

9. A multi-camera spatiotemporal joint synchronous acquisition and processing device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Synchronous output method and system for multi-channel collected data and RGBD camera

    CN115484407A

  • Video data processing method and device, equipment and storage medium

    CN115695883A