Video synchronization processing method and device, electronic equipment and storage medium

CN115988172BActive Publication Date: 2026-08-21ZHEJIANG UNIVIEW TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111191020.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-13
Publication Date
2026-08-21
Estimated Expiration
2041-10-13

AI Technical Summary

Technical Problem

为了目标对象框叠加效果更好,会采取视频帧与算法检测帧同步的方法,每帧运算同步方法的目标叠加准确性强,目标对象框随目标运动流畅,但硬件实现不友好,需要算法检测帧率与视频帧率同步,因此实用性差,目标对象框随目标运动不流畅,有严重顿挫感

Benefits of technology

[0018]本发明实施例中提供了一种视频同步处理方法,确定视频帧的当前检测帧信息与下一检测帧信息;当前检测帧与下一检测帧为间隔至少一个视频帧对不同视频帧进行目标检测所得的两个连续检测帧;依据所述当前检测帧信息与下一检测帧信息,推算得到从当前检测帧与下一检测帧期间未进行目标检测的视频帧的目标检测帧信息,再向视频帧中添加与视频帧匹配的检测帧信息。采用本申请方案,针对每个视频帧均叠加了对应的检测帧信息,保证每一个视频帧均有其同步的检测帧信息,这样保持检测帧信息对应的目标对象框跟踪视频帧中目标对象的流畅性;同时,考虑到算法检测速率要低于视频输出速率,在目标检测时很难针对所有视频帧均进行一次目标检测,而是选择通过动量估计运算期间未进行目标检测的其他视频帧的检测帧信息,实现为每个视频帧叠加检测帧信息,既保持了目标对象框跟踪目标的流畅性,又可以提高方案在不同性能的硬件平台上的适应性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115988172B_ABST
    Figure CN115988172B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a video synchronization processing method and device, electronic equipment and storage medium. The method comprises: determining current detection frame information and next detection frame information of a video frame; the current detection frame and the next detection frame are two continuous detection frames obtained by performing target detection on different video frames with an interval of at least one video frame; according to the current detection frame information and the next detection frame information, target detection frame information of a video frame during which target detection is not performed is calculated; matched detection frame information is added to the video frame, so that a target object box indicated by the matched detection frame information is synchronously displayed when the video frame is displayed. According to the application, corresponding detection frame information is added to each video frame, so that each video frame has its synchronous detection frame information, and the smoothness of the target object box corresponding to the detection frame information in tracking the target object in the video frame is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of intelligent monitoring technology, and in particular to a video synchronization processing method, apparatus, electronic device and storage medium. Background Technology

[0002] Face detection and human / machine detection are widely used applications of network cameras (IPCs), and are effective technologies for quickly detecting and extracting target objects in surveillance footage.

[0003] In practical applications, most monitoring scenarios involve complex environments where vehicles, non-vehicles, and humans coexist. To highlight the location of a target object in the monitored image, a bounding box is overlaid on the target object. To achieve a better overlay effect, a method of synchronizing video frames with algorithm-detected frames is used. This method, which performs frame-by-frame synchronization, offers high accuracy in target overlay and allows the bounding box to move smoothly with the target. However, it is not hardware-friendly, requiring the algorithm's detection frame rate to be synchronized with the video frame rate. Therefore, its practicality is poor, and the bounding box does not move smoothly with the target, resulting in a noticeable stutter. Summary of the Invention

[0004] This invention provides a video synchronization processing method, apparatus, electronic device, and storage medium to achieve smooth target object frame tracking and adaptability to hardware platforms with different performance levels.

[0005] In a first aspect, embodiments of the present invention provide a video synchronization processing method, the method comprising:

[0006] Determine the current detection frame information and the next detection frame information of the video frame; the current detection frame and the next detection frame are two consecutive detection frames obtained by performing target detection on different video frames at least one video frame interval;

[0007] Based on the current detection frame information and the next detection frame information, the target detection frame information of the video frames that were not detected between the current detection frame and the next detection frame is calculated.

[0008] Add matching detection frame information to the video frame so that the target object box indicated by the matching detection frame information of the video frame is displayed synchronously when the video frame is displayed.

[0009] Secondly, this invention also provides a video synchronization processing apparatus, which includes:

[0010] The detection frame determination module is used to determine the current detection frame information and the next detection frame information of the video frame; the current detection frame and the next detection frame are two consecutive detection frames obtained by performing target detection on different video frames at an interval of at least one video frame.

[0011] The target estimation module is used to estimate the target detection frame information of video frames that have not been detected between the current detection frame and the next detection frame based on the current detection frame information and the next detection frame information.

[0012] The detection frame matching module is used to add matching detection frame information to video frames, so as to synchronously display the target object box indicated by the matching detection frame information when displaying video frames.

[0013] Thirdly, this invention also provides an electronic device, comprising:

[0014] One or more processing devices;

[0015] Storage device for storing one or more programs;

[0016] When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the video synchronization processing method provided in any embodiment of the present invention.

[0017] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processing device, implements the video synchronization processing method provided in any embodiment of the present invention.

[0018] This invention provides a video synchronization processing method, which determines the current detection frame information and the next detection frame information of a video frame. The current detection frame and the next detection frame are two consecutive detection frames obtained by performing target detection on different video frames at least one video frame interval. Based on the current detection frame information and the next detection frame information, the target detection frame information of video frames that have not undergone target detection between the current detection frame and the next detection frame is calculated, and then the detection frame information matching the video frame is added to the video frame. Using this application's scheme, corresponding detection frame information is superimposed on each video frame, ensuring that each video frame has its own synchronized detection frame information. This maintains the smoothness of the target object box tracking the target object in the video frame. Simultaneously, considering that the algorithm's detection rate is lower than the video output rate, it is difficult to perform target detection on all video frames. Instead, the detection frame information of other video frames that have not undergone target detection during the momentum estimation operation is selected, thus superimposing detection frame information on each video frame. This maintains the smoothness of the target object box tracking the target and improves the adaptability of the scheme to different hardware platforms.

[0019] The above description of the invention is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0020] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0021] Figure 1 This is a flowchart of a video synchronization processing method provided in an embodiment of the present invention;

[0022] Figure 2 This is a schematic diagram illustrating the motion relationship between a target object and its frame under intermittent display, provided in an embodiment of the present invention.

[0023] Figure 3 This is a flowchart of another video synchronization processing method provided in an embodiment of the present invention;

[0024] Figure 4 This is a schematic diagram of video frame and detection frame synchronization control provided in an embodiment of the present invention;

[0025] Figure 5 This is a schematic diagram of the buffer area at time t2 during video synchronization processing provided in an embodiment of the present invention;

[0026] Figure 6 This is a schematic diagram of a buffer area that completes T2-T7 interpolation at time t4 during video synchronization processing, provided in an embodiment of the present invention.

[0027] Figure 7 This is a schematic diagram of the buffer area after frame synchronization is completed at time t4 during video synchronization processing, provided in an embodiment of the present invention.

[0028] Figure 8 This is a flowchart of another video synchronization processing method provided in this embodiment of the invention;

[0029] Figure 9 This is a schematic diagram illustrating the calculation of target movement offset during video synchronization processing, provided in an embodiment of the present invention.

[0030] Figure 10 This is a schematic diagram of cubic spline-natural boundary interpolation during video synchronization processing provided in an embodiment of the present invention;

[0031] Figure 11 This is a schematic diagram of cubic spline-non-node boundary interpolation during video synchronization processing provided in an embodiment of the present invention;

[0032] Figure 12 This is a structural block diagram of a video synchronization processing device provided in an embodiment of the present invention;

[0033] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0034] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention and not the entire structure.

[0035] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations (or steps) may be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations may be rearranged. The process may be terminated when its operation is completed, but may also have additional steps not included in the figures. The process may correspond to a method, function, procedure, subroutine, subroutine, etc.

[0036] The video synchronization processing method, apparatus, electronic device, and storage medium provided in this application will be described in detail below through various embodiments and their optional solutions.

[0037] Figure 1 This is a flowchart of a video synchronization processing method provided in an embodiment of the present invention. This embodiment of the present invention is applicable to situations where a target object frame is superimposed on the target object in the image to highlight its location. This method can be executed by a video synchronization processing device, which can be implemented in software and / or hardware and integrated into any electronic device with network communication capabilities. Figure 1 As shown, the video synchronization processing method provided in this application embodiment may include the following steps:

[0038] S110. Determine the current detection frame information and the next detection frame information of the video frame.

[0039] The current detection frame and the next detection frame are two consecutive detection frames obtained by performing target detection on different video frames at least one video frame apart.

[0040] When the target detection rate is synchronized with the video frame rate, for example, when the hardware performance is very high and there are sufficient computing resources to ensure that the target detection rate is synchronized with the video frame rate, a frame-by-frame synchronization scheme can be used to make the target object bounding box indicated by the target detection frame information move with the target object flow field in the video frame. However, in many cases, for some embedded devices, the hardware performance is not sufficient, and there may not be enough computing resources to ensure that the target detection rate is synchronized with the video frame rate, resulting in the target detection rate being lower than the video frame rate.

[0041] Therefore, when performing object detection on a series of video frames, instead of performing intelligent object detection on all video frames, object detection is performed on different video frames before and after the interval of at least one video frame. This reduces the computational load of object detection and avoids requiring more hardware performance due to a large number of object detections. Furthermore, in selecting detection frames, two consecutive detection frames are chosen: the current detection frame information and the next detection frame information. This allows the information of the video frames between the current and next detection frames to be estimated based on the information of the current detection frame and the information of the next detection frame generated after a delay.

[0042] Here, a video frame can represent a single frame of video data from a surveillance camera used for video playback, while a detection frame can represent the set of coordinates of a target object within the surveillance camera's view, calculated once based on a single video frame by intelligent target detection. For example, intelligent target detection methods can include, but are not limited to, face detection and vehicle / non-motorized vehicle (MOV) detection, where MOV detection includes the detection of motor vehicles, non-motorized vehicles, and pedestrians using intelligent AI algorithms.

[0043] S120. Based on the current detection frame information and the next detection frame information, calculate the target detection frame information of the video frames that were not detected between the current detection frame and the next detection frame.

[0044] Since the aforementioned approach does not perform object detection for each video frame, it reduces the computational burden of object detection to some extent. However, see... Figure 2 When the video frames between two consecutive detection frames are not subject to target detection, and the target object bounding box is used to track the target object in the video frame, the target object bounding box will not move smoothly with the target object because there is no matching target detection frame information for these video frames. This will result in a severe sense of stuttering, especially on low-performance hardware, where the target object bounding box lag will be more obvious.

[0045] Detection frame information indicates the target detection results of the target objects identified through target detection in the video frame within the captured image. The target detection results can be highlighted and described in the form of target object bounding boxes, highlighting the position of the target object within the captured image. Detection frame information may include target object bounding box position information used to annotate the target object. The current detection frame and the next detection frame are two consecutive detection frames, and the corresponding video frames are very close in time. Therefore, the target objects in the two corresponding video frames have a certain offset relationship in position.

[0046] Therefore, the target object box position of each video frame during the detection period between the current detection frame and the next detection frame can be obtained by calculating the offset between the target object boxes indicated by the two detection frames. In this way, for some video frames that are missed due to limited target detection, the target detection frame information can be inferred by the movement and offset of the target object in two consecutive detection frames to supplement the target object boxes of the video frames that have not been detected.

[0047] S130. Add matching target detection frame information to the video frame to synchronously display the target object box indicated by the matching detection frame information of the video frame when displaying the video frame.

[0048] Each video frame in the video stream has a matching frame sequence identifier (ID). The detection frame information obtained by performing object detection on the video frame has the same frame sequence identifier (ID) as the video frame, thus achieving a matching association between the target object bounding box positions in the video frame and the detection frame information. Simultaneously, after calculating the target detection frame information for each video frame between the current detection frame and the next detection frame, the frame sequence identifier of the video frame can be assigned to the corresponding calculated target detection frame information.

[0049] By combining the target object bounding boxes obtained from target detection of different video frames in the current detection frame information and the next detection frame information, as well as the target object bounding boxes included in the target detection frame information of each video frame between the current and next detection frames, the target object bounding boxes in the detection frame information matched by each video frame can be determined. This allows for the addition of the target object bounding box data, including the detected and inferred target object bounding boxes, to the matched video frames based on the association and matching between the video frames and the detection frame information. Subsequently, while playing the video frames, the target object bounding boxes indicated by the detection frame information for that frame will be displayed on the video frame screen, effectively highlighting the location of target objects in complex scenarios involving mixed traffic of motor vehicles, non-motor vehicles, and pedestrians.

[0050] In one optional embodiment, adding matching detection frame information to a video frame may include the following steps:

[0051] The target detection frame information matched by the video frame is encoded into the video frame according to the preset encoding method, so as to decode and display the video frame from the video frame and synchronously overlay the target detection frame information included in the video frame.

[0052] The process from the current detection frame to the next detection frame is considered as a detection cycle. After obtaining the detection frame information that matches each video frame within a detection cycle, the system synchronizes with the video frames. By using the frame sequence identifier ID recorded in the video frame and the frame sequence identifier ID corresponding to the detection frame information, the system matches the video frames and adds detection frame information. The successfully matched detection frame information is encoded into the video frame.

[0053] The encoding method used to encode the detection frame information into video frames can be common metadata encoding methods such as SEI or RTP metadata. The video stream transmitted over the network can be encoded using the RTP encoding format. The video receiving and decoding end on the network extracts the video frame data from the video stream, and while playing the video frames, it overlays the target information corresponding to the frames onto the screen. This process is a conventional video decoding process, and this article does not impose any limitations on it.

[0054] In another optional embodiment, adding matching detection frame information to the video frame may include the following steps:

[0055] Before the video frame is sent, the detection frame information matching the video frame is superimposed on the video frame so that the video frame with the superimposed detection frame information can be directly displayed after the video frame is sent.

[0056] The target object bounding box indicated by the detection frame information matched to each video frame can also be superimposed onto the video frame image within the network camera IPC before the video stream is sent. Subsequently, the decoding end does not need to perform the target object bounding box superimposition process; after acquiring the video frame data from the video stream, directly playing the video frame will simultaneously display the target object bounding box information of the corresponding video frame that was superimposed before the stream was sent.

[0057] This invention provides a video synchronization processing method that overlays corresponding detection frame information onto each video frame, ensuring that each video frame has its own synchronized detection frame information. This maintains the smoothness of the target object bounding box corresponding to the detection frame information tracking the target object in the video frame. Simultaneously, considering that the algorithm's detection rate is lower than the video output rate, it is difficult to perform target detection on all video frames. Instead, the detection frame information of other video frames that were not detected during momentum estimation is selected. This allows the target object bounding box indicated in the detection frame information to be overlaid on each video frame, highlighting the target object's position during display. This maintains the smoothness of the target object bounding box tracking the target and improves the adaptability of the solution to hardware platforms with different performance levels.

[0058] Figure 3This is a flowchart of another video synchronization processing method provided in this embodiment of the invention. This embodiment further optimizes the aforementioned embodiments, and can be combined with various optional solutions from one or more of the above embodiments. For example... Figure 3 As shown, the video synchronization processing method provided in this application embodiment may include the following steps:

[0059] S310. When buffering the first video frame split from the source video stream, perform target detection once for the second video frame split from the source video stream at least one video frame interval.

[0060] Among them, the first video frame and the second video frame belong to the same source video frame, and both come from the video frames in the source video stream. In addition, the frame sequence identifier of the video frame is the same as the frame sequence identifier of the video frame in the source video stream.

[0061] See Figure 4 The source video stream can be split into two video frames, namely the first video frame and the second video frame. The first video frame is buffered, and while the first video frame is buffered, the second video frame is input into the intelligent object detection to generate detection frame information. Since the detection rate of the object detection algorithm is lower than the video output rate, it is impossible to perform object detection on all video frames in the second video frame. In order to keep the buffering of the first video frame synchronized with the object detection and buffering of the second video frame, object detection is performed on the second video frame at an interval of at least one video frame. Thus, between two object detections, some video frames in the second video frame will have been object detected, while others will not.

[0062] In one optional embodiment, performing target detection on the second video frame derived from the source video stream at least once every other video frame may include the following steps A1-A2:

[0063] Step A1: After performing target detection on one video frame in the second video frame to obtain current detection frame information, select the next video frame from the second video frame after an interval of at least one video frame.

[0064] Among them, at least one video frame between two consecutive detection frames includes a first value of video frames; the first value is determined by rounding the ratio of the video frame rate to the detection rate of the target detection.

[0065] See Figure 4 and Figure 5While performing object detection on a video frame from the second video stream, the first video stream is also being continuously buffered. Since the detection rate of the object detection algorithm is lower than the video output rate, it's possible that multiple video frames from the first video stream need to be buffered before object detection can be completed on a single video frame. If, at this point, object detection is then performed on the already buffered video frames from the first video stream that are being synchronized with the video frames from the second video stream (for example, performing object detection on F1 to obtain result T1, and then sequentially performing object detection on F2), it will inevitably consume a significant amount of time, causing subsequent video frames to be unable to be detected quickly, resulting in a buildup of detections.

[0066] Therefore, the detection of video frame buffers and detection frames can be synchronized. By setting the relationship between the time interval between two consecutive detection frames and the video frame interval, the timing of the algorithm's detection startup can be controlled. Let the video frame rate be Ffps, the target detection rate of the target detection algorithm be X times / s, and the target detection rate of the target detection algorithm can be dynamically changed in real time. The output interval between two consecutive detection frames is... With a certain number of video frames (ceil is a decimal rounding method), we can then determine how many video frames from the first video stream will be buffered before completing a target detection operation on the second video stream.

[0067] Step A2: Start and perform target detection on the next video frame selected from the second video frame, and output the next detection frame information for the target detection of the next video frame.

[0068] See Figure 4 After performing a target detection on one video frame in the second video frame to obtain a current detection frame information, the next video frame in the second video frame is selected by maintaining an integer multiple relationship between the time interval between two consecutive detections and the video frame interval. A target detection is then performed on the next video frame selected from the second video frame. At the same time as the target detection, video frames from the same source as the next video frame selected from the second video frame are cached.

[0069] See Figure 4 and Figure 5 Let the frame rate of the video be F fps, the target detection rate of the target detection algorithm be X times / s, and the interval between each video frame generation be... The time required for each object detection algorithm to perform object detection To make the synchronization control diagram more intuitive, the diagram is set to F = 25fps, X = 4 times / s, and the interval between two detection frame outputs is... There are 10 video frames. At time t1, the same source video frames of the scene corresponding to video frame F1 are simultaneously input into the target detection algorithm for target detection. By time t2, the detection frame T1 is output. F1 and T1 are a pair of corresponding and matched video frames and detection frames.

[0070] See Figure 4 and Figure 5 The time interval between two consecutive detection frames is kept as an integer multiple of the video frame interval. This ensures that the buffering of the first video frame is synchronized with the target detection of the second video frame. When target detection is initiated on a video frame from the second video frame, a corresponding video frame from the first video frame is being buffered, guaranteeing more accurate calculation of the target motion offset. Therefore, at time t3, the next frame of the video frame is selected, and the image corresponding to video frame F8 is simultaneously input into the target detection algorithm for target detection. At time t4, detection frame T8 is output. F8 and T8 are a pair of corresponding matching video frames and detection frames.

[0071] In one optional embodiment, the buffer capacity required for video frames is greater than or equal to the capacity of a second set of video frames, for example, for the first video frame stream. The buffer capacity required for detection frames is greater than the capacity of a first set of detection frames, for example, for the detection frames of the second video frame stream. The first value is determined by rounding down the ratio of the video frame rate to the detection rate of the target detection, and the second value is equal to the first value by a preset multiple.

[0072] See Figure 4 and Figure 6 To complete the target object bounding box calculation, it is usually necessary to obtain two consecutive detection frame information. To obtain two consecutive detection frame information, it is usually necessary to cache the number of video frames that correspond to the two consecutive detection frames. These cached video frames need to be retained, so the video frame cache capacity needs to be greater than or equal to the capacity required for this process.

[0073] For example, suppose the video frame rate is F fps, the target detection rate of the target detection algorithm is X times / s, and the output interval between two consecutive detection frames is... The video frames (ceil is a decimal rounding method) are used. Since the detection frame does not need to retain the video frame size corresponding to the interval between two consecutive detection frames, therefore... Let it be the first value, and then... The value that is twice the value of is called the second value.

[0074] S320. Perform frame matching between the target detection information obtained from the second video frame and the cached first video frame to obtain the detection frame information that matches some video frames in the first video frame and cache it.

[0075] The detection frame generated by performing target detection on the second video frame split from the source video stream also uses the frame sequence identifier ID of the corresponding second video frame undergoing target detection; that is, the frame sequence identifier ID of the detection frame is the same as that of the second video frame undergoing target detection. Furthermore, the first and second video frames split from the source video stream belong to the same source video frame, and both source video frames are labeled using the frame sequence identifier ID of the video frame in the source video stream.

[0076] Based on the frame sequence identifier IDs matched from the first video frame and the frame sequence identifier IDs matched from the detection frame information, the frame sequence identifier IDs corresponding to the detection frame information obtained from target detection of the second video frame can be matched with the frame sequence identifiers corresponding to the cached first video frame. In this way, the portion of the first video frame that matches the frame sequence identifier of the detection frame information of the second video frame can be matched from the first video frame.

[0077] S330: Obtain two consecutive detection frame information from the buffer, and use them as the current detection frame information and the next detection frame information for the video stream, respectively.

[0078] The current detection frame and the next detection frame are two consecutive detection frames obtained by performing target detection on different video frames at least one video frame apart.

[0079] See Figure 7 After frame synchronization is completed at time t4, F1(T1)-F7(T7) are sent out of the buffer, and the last frame of the previous cycle, F8(T8), remains in the buffer as the start frame of the next cycle, and then the cycle operation is repeated.

[0080] S340. Based on the current detection frame information and the next detection frame information, calculate the target detection frame information of the video frame that was not detected between the current detection frame and the next detection frame.

[0081] S350. Add matching detection frame information to the video frame to synchronously display the target object box indicated by the matching detection frame information when displaying the video frame.

[0082] This invention provides a video synchronization processing method. The time interval between two consecutive detection frames is maintained as an integer multiple of the video frame interval, achieving synchronization between target detection and video frame buffering. Furthermore, corresponding detection frame information is superimposed on each video frame to ensure that each video frame has its own synchronized detection frame information. This maintains the smoothness of the target object bounding box tracking the target object in the video frame. Simultaneously, considering that the algorithm's detection rate is lower than the video output rate, it is difficult to perform target detection on all video frames. Instead, the detection frame information of other video frames that were not detected during momentum estimation is selected. This allows the target object bounding box indicated in the detection frame information to be superimposed on each video frame to highlight the target object's position during display. This maintains the smoothness of the target object bounding box tracking the target and improves the adaptability of the solution to hardware platforms with different performance levels.

[0083] Figure 8 This is a flowchart of another video synchronization processing method provided in this embodiment of the invention. This embodiment further optimizes the aforementioned embodiments, and can be combined with various optional solutions from one or more of the above embodiments. For example... Figure 8 As shown, the video synchronization processing method provided in this application embodiment may include the following steps:

[0084] S810. Determine the current detection frame information and the next detection frame information of the video frame.

[0085] The current detection frame and the next detection frame are two consecutive detection frames obtained by performing target detection on different video frames at least one video frame interval.

[0086] S820. Based on the current detection frame information and the next detection frame information, determine the relative offset between the target object indicated by the current detection frame information and the target object indicated by the next detection frame information.

[0087] See Figure 4 and Figure 9 At time t4, based on the two consecutive detection frames T1 and T8, and using the target object bounding box coordinates (X1, Y1) and (X2, Y2) in detection frame T1 and (X3, Y3) and (X4, Y4) in detection frame T8 (detected by video frame F1), the coordinates are known values ​​in the frame. Since the interval between detection frames T1 and T8 is relatively short, it can be assumed that the target object moves at a constant speed during these frames. Therefore, the relative offset of the target object during its movement can be determined based on its position in the two detection frames.

[0088] S830. Based on the relative offset between target objects, interpolate to estimate the position of the target object in the video frame that was not detected between the current detection frame and the next detection frame, so as to obtain the target detection frame information of the video frame that was not detected between the current detection frame and the next detection frame.

[0089] After determining the relative offset of the target object based on the target object position in the two detection frame information, the target object position in each video frame that has not undergone target detection during the period from the current detection frame to the next detection frame can be estimated by interpolation. In this way, the target object bounding box position that needs to be marked in the video frames that have not undergone target detection can be estimated according to the estimated target object position, and the target detection frame information of the video frames that have not undergone target detection can be obtained.

[0090] In one optional embodiment, interpolation based on the relative offset between target objects may include the following steps:

[0091] First-order linear interpolation is performed based on the linear motion offset between the target objects in the two detection frames, which are indicated by the relative offset between the target objects.

[0092] See Figure 9 First-order linear interpolation: Given the coordinates of the top-left and bottom-right corners of the target object bounding box in T1 (X1, Y1) and (X2, Y2), and the coordinates of the top-left and bottom-right corners of the target object bounding box in T8 (X3, Y3) and (X4, Y4), calculate the coordinates of T2 and T3 as follows:

[0093]

[0094]

[0095] The coordinate values ​​for T2-T7 are obtained by analogy. Other target information is filled using the content of T1, and the calculation method for other cycle periods is the same. In actual calculations, the content in the formula is replaced according to the actual video frame rate and the algorithm's detection rate.

[0096] Tn=(X i Y i ), (X′ i ,Y′ i ), (n = 1, 2, 3, ..., n)

[0097] Where (X) i Y i (X′) is the coordinate of the top-left point of Tn. i ,Y′ i Let ) be the coordinates of the lower right point of Tn. Based on the above, the following formula can be derived:

[0098]

[0099] In one optional embodiment, interpolation based on the relative offset between target objects may include the following steps:

[0100] Cubic spline interpolation is performed based on the relative offset between the target objects to smooth the interpolation trajectory points.

[0101] See Figure 10 Cubic spline interpolation can smooth the interpolated trajectory points, making them more consistent with the actual movement trajectory of the monitored target. Cubic splines segment the trajectory. Assume there is a set of target trajectory points within the trajectory interval [a, b]: (X1, Y1), (X2, Y2), (X3, Y3)...(X... n Y n ), a≤X1 <X2<···X n ≤b, let each segmented interval [x1, x2, ..., x3] be [x1, x3, ..., x4]. 1+1 S(x) is a cubic equation. i )=y i (i = 1, 2, ..., n).

[0102] See also Figure 10 Cubic splines have three types of boundary conditions: natural boundary, fixed boundary, and non-node boundary; the second derivative of a natural boundary is 0 at a specified endpoint, i.e., S″(x1)=S″(x1). n When the boundary condition is 0, the continuity of the boundary connection points is poor. Since the process of superimposing the target box in the video is real-time and continuous, the algorithm needs to superimpose the target as soon as it exits the detection frame. It needs to complete the interpolation of an interval and synchronize. Each interpolation is performed at the edge. Therefore, using natural boundaries as conditions will result in poor trajectory smoothness.

[0103] See Figure 10 and Figure 11 To achieve better trajectory smoothness, non-node boundaries can be used as boundary conditions, such that the third derivative of the first point equals the third derivative of the second point, and the third derivative of the last point equals the third derivative of the penultimate point; that is, S″′1(x1)=S″′2(x2), S″′ n-2 (x n-1 )=S″′ n-1 (x n ).

[0104] The following can be obtained by solving the problem using commonly used cubic splines:

[0105] In each subinterval x i ≤x≤x i+1 There exists an equation in S: i (x)=a i +bi (xx i )+c i (xx i ) 2 +d i (xx i ) 3 ;

[0106] The step size is represented by h1, h i =x i+1 -x i Second derivative m i =S″ i (x i );

[0107] a i =y i

[0108]

[0109]

[0110]

[0111] When solving m i It is the only unknown value, and it can be solved using the common method Gaussian elimination, which will not be described in detail in this article.

[0112] In summary, the starting point coordinates T1 = (X1, Y1), and the ending point coordinates T... n =(X n Y n Substituting the target position into the detection frame, we get the formula:

[0113] Tn=(X i Y i ), (X′ i ,Y′ i ), (n = 1, 2, 3, ..., n)

[0114] Equations in each subinterval:

[0115] Y i =a i +b i (X i -x i )+c i (X i -x i ) 2 +d i (X i -x i ) 3

[0116] Solving the matrix equation yields the result; the last interval is then... Divide the matrix into equal parts, substitute the X-axis coordinates into the formula to obtain the corresponding Y-axis coordinates, and the interpolation will be completed.

[0117] S840. Add matching detection frame information to the video frame so that the target object box indicated by the matching detection frame information of the video frame is displayed synchronously when the video frame is displayed.

[0118] See Figure 6 By inserting detection frames corresponding to video frames into the target bounding box buffer, the target bounding box is superimposed on each frame, maintaining the smoothness of target bounding box tracking and improving the adaptability of the solution to hardware platforms with different performance levels, thus balancing the effect and efficiency of target superposition.

[0119] This invention provides a video synchronization processing method that overlays corresponding detection frame information onto each video frame, ensuring that each video frame has its own synchronized detection frame information. This maintains the smoothness of the target object bounding box corresponding to the detection frame information tracking the target object in the video frame. Simultaneously, considering that the algorithm's detection rate is lower than the video output rate, it is difficult to perform target detection on all video frames. Instead, the detection frame information of other video frames that were not detected during momentum estimation is selected. This allows the target object bounding box indicated in the detection frame information to be overlaid on each video frame, highlighting the target object's position during display. This maintains the smoothness of the target object bounding box tracking the target and improves the adaptability of the solution to hardware platforms with different performance levels.

[0120] Figure 12 This is a structural block diagram of a video synchronization processing device provided in an embodiment of the present invention. This embodiment of the present invention is applicable to situations where a target object frame is superimposed on the target object in the image to highlight its position. This device can be implemented in software and / or hardware and integrated into any electronic device with network communication capabilities. Figure 12 As shown, the video synchronization processing apparatus provided in this embodiment may include the following: a detection frame determination module 1210, a target estimation module 1220, and a detection frame matching module 1230. Wherein:

[0121] The detection frame determination module 1210 is used to determine the current detection frame information and the next detection frame information of the video frame; the current detection frame and the next detection frame are two consecutive detection frames obtained by performing target detection on different video frames at an interval of at least one video frame.

[0122] The target estimation module 1220 is used to estimate the target detection frame information of video frames that have not been detected between the current detection frame and the next detection frame based on the current detection frame information and the next detection frame information;

[0123] The detection frame matching module 1230 is used to add matching detection frame information to the video frame so as to synchronously display the target object box indicated by the matching detection frame information of the video frame when displaying the video frame.

[0124] Based on the above embodiments, optionally, the detection frame determination module 1210 includes:

[0125] While buffering the first video frame derived from the source video stream, target detection is performed on the second video frame derived from the source video stream at least once every other video frame.

[0126] The detection frame information obtained by target detection of the second video frame is matched with the cached first video frame to obtain the detection frame information that matches some video frames in the first video frame and cached; the first video frame and the second video frame belong to the same source video frame;

[0127] Retrieve two consecutive detection frame information from the cache, and use them as the current detection frame information and the next detection frame information for the video stream, respectively.

[0128] Based on the above embodiments, optionally, target detection is performed once every at least one video frame interval on the second video frame split from the source video stream, including:

[0129] After performing target detection on one video frame in the second video frame and obtaining current detection frame information, the next video frame is selected from the second video frame after an interval of at least one video frame.

[0130] Start and perform target detection on the next video frame selected from the second video frame, and output the next detection frame information for the target detection of the next video frame;

[0131] Wherein, at least one video frame between two consecutive detection frames includes a first value of video frames; the first value is determined by rounding the ratio of the video frame rate to the detection rate of the target detection.

[0132] Based on the above embodiments, optionally, the buffer capacity required for video frames is greater than or equal to the second value of video frame capacity, and the buffer capacity required for detection frames is greater than the first value of detection frame capacity; the first value is determined by rounding the ratio of video frame rate to target detection rate, and the second value is equal to the first value by a preset multiple.

[0133] Based on the above embodiments, optionally, the target estimation module 1220 includes:

[0134] Based on the current detection frame information and the next detection frame information, determine the relative offset between the target object indicated by the current detection frame information and the target object indicated by the next detection frame information;

[0135] Interpolation is performed based on the relative offset between the target objects to estimate the position of the target object in the video frame from the current detection frame to the next detection frame that has not been detected, so as to obtain the target detection frame information.

[0136] Based on the above embodiments, optionally, interpolation is performed based on the relative offset between target objects, including:

[0137] First-order linear interpolation is performed based on the linear motion offset between the target objects in two detection frames, indicated by the relative offset between the target objects; or...

[0138] Cubic spline interpolation is performed based on the relative offset between the target objects to smooth the interpolation trajectory points.

[0139] Based on the above embodiments, optionally, the frame matching detection module 1230 includes:

[0140] The detection frame information matching the video frame is encoded into the video frame according to a preset encoding method, so as to decode and display the video frame and synchronously display the detection frame information included in the video frame; or,

[0141] Before the video frame is sent, the detection frame information matching the video frame is superimposed on the video frame so that the video frame with the superimposed detection frame information can be directly displayed after the video frame is sent.

[0142] The video synchronization processing device provided in the embodiments of the present invention can execute the video synchronization processing method provided in any of the embodiments of the present invention, and has the corresponding functions and beneficial effects of executing the video synchronization processing method. For details, please refer to the relevant operations of the video synchronization processing method in the foregoing embodiments.

[0143] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. For example... Figure 13 The structure shown in this embodiment of the invention includes an electronic device comprising one or more processors 1310 and a storage device 1320; the processors 1310 in this electronic device may be one or more. Figure 13 Taking a processor 1310 as an example; storage device 1320 is used to store one or more programs; the one or more programs are executed by the one or more processors 1310, so that the one or more processors 1310 implement the video synchronization processing method as described in any one embodiment of the present invention.

[0144] The electronic device may also include an input device 1330 and an output device 1340.

[0145] The processor 1310, storage device 1320, input device 1330, and output device 1340 in this electronic device can be connected via a bus or other means. Figure 13 Taking the example of a connection between China and Israel via a bus.

[0146] The storage device 1320 in this electronic device serves as a computer-readable storage medium, which can be used to store one or more programs. These programs can be software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the video synchronization processing method provided in this embodiment of the invention. The processor 1310 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the storage device 1320, thereby implementing the video synchronization processing method described in the above embodiment.

[0147] Storage device 1320 may include a stored program area and a stored data area, wherein the stored program area may store the operating system and applications required for at least one function; the stored data area may store data created based on the use of the electronic device, etc. Furthermore, storage device 1320 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, storage device 1320 may further include memory remotely located relative to processor 1310, and this remote memory may be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0148] Input device 1330 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the electronic device. Output device 1340 may include display devices such as a display screen.

[0149] Furthermore, when one or more programs included in the aforementioned electronic device are executed by one or more processors 1310, the programs perform the following operations:

[0150] Determine the current detection frame information and the next detection frame information of the video frame; the current detection frame and the next detection frame are two consecutive detection frames obtained by performing target detection on different video frames at least one video frame interval;

[0151] Based on the current detection frame information and the next detection frame information, the target detection frame information of the video frames that were not detected between the current detection frame and the next detection frame is calculated.

[0152] Add matching detection frame information to the video frame so that the target object box indicated by the matching detection frame information of the video frame is displayed synchronously when the video frame is displayed.

[0153] Of course, those skilled in the art will understand that when one or more programs included in the above-mentioned electronic device are executed by one or more processors 1310, the programs can also perform related operations in the video synchronization processing method provided in any embodiment of the present invention.

[0154] This invention provides a computer-readable medium storing a computer program thereon, which, when executed by a processor, performs a video synchronization processing method, the method comprising:

[0155] Determine the current detection frame information and the next detection frame information of the video frame; the current detection frame and the next detection frame are two consecutive detection frames obtained by performing target detection on different video frames at least one video frame interval;

[0156] Based on the current detection frame information and the next detection frame information, the target detection frame information of the video frames that were not detected between the current detection frame and the next detection frame is calculated.

[0157] Add matching detection frame information to the video frame so that the target object box indicated by the matching detection frame information of the video frame is displayed synchronously when the video frame is displayed.

[0158] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, a portable CD-ROM, an optical storage device, a magnetic storage device, or any suitable combination thereof. A computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0159] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit programs for use by or in connection with an instruction execution system, apparatus, or device.

[0160] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, radio frequency (RF), etc., or any suitable combination thereof.

[0161] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0162] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0163] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

Claims

1. A video synchronization processing method, characterized in that, The method includes: Determine the current detection frame information and the next detection frame information of the video frame; the current detection frame and the next detection frame are two consecutive detection frames obtained by performing target detection on different video frames at least one video frame interval; Based on the current detection frame information and the next detection frame information, the target detection frame information of the video frames that were not detected between the current detection frame and the next detection frame is calculated. Add matching detection frame information to the video frame so that the target object box indicated by the matching detection frame information of the video frame is displayed synchronously when the video frame is displayed. The determination of the current detection frame information and the next detection frame information of the video frame includes: While buffering the first video frame derived from the source video stream, target detection is performed on the second video frame derived from the source video stream at least once every other video frame. The detection frame information obtained by target detection of the second video frame is matched with the cached first video frame to obtain the detection frame information that matches some video frames in the first video frame and cached; the first video frame and the second video frame belong to the same source video frame; Two consecutive detection frame information are obtained from the cache and used as the current detection frame information and the next detection frame information of the video stream, respectively; wherein, at least one video frame between the two consecutive detection frames includes a first value of video frames; the first value is determined by rounding the ratio of the video frame rate to the detection rate of the target detection.

2. The method according to claim 1, characterized in that, For the second video frame segmented from the source video stream, target detection is performed once every at least one video frame interval, including: After performing target detection on one video frame in the second video frame and obtaining current detection frame information, the next video frame is selected from the second video frame after an interval of at least one video frame. Initiate and perform target detection on the next video frame selected from the second video frame, and output the next detection frame information for the target detection of the next video frame.

3. The method according to any one of claims 1-2, characterized in that, The buffer capacity required for video frames is greater than or equal to the second value of video frame capacity, and the buffer capacity required for detection frames is greater than the first value of detection frame capacity; the first value is determined by rounding the ratio of video frame rate to target detection rate, and the second value is equal to the first value by a preset multiple.

4. The method according to claim 1, characterized in that, Based on the current detection frame information and the next detection frame information, the target detection frame information of video frames that did not undergo target detection between the current detection frame and the next detection frame is calculated, including: Based on the current detection frame information and the next detection frame information, determine the relative offset between the target object indicated by the current detection frame information and the target object indicated by the next detection frame information; Interpolation is performed based on the relative offset between the target objects to estimate the position of the target object in the video frame from the current detection frame to the next detection frame that has not been detected, so as to obtain the target detection frame information.

5. The method according to claim 4, characterized in that, Interpolation is performed based on the relative offsets between target objects, including: First-order linear interpolation is performed based on the linear motion offset between the target objects in two detection frames, indicated by the relative offset between the target objects; or... Cubic spline interpolation is performed based on the relative offset between the target objects to smooth the interpolation trajectory points.

6. The method according to claim 1, characterized in that, Add matching detection frame information to the video frames, including: The detection frame information matching the video frame is encoded into the video frame according to a preset encoding method, so as to decode and display the video frame and synchronously display the detection frame information included in the video frame; or, Before the video frame is sent, the detection frame information matching the video frame is superimposed on the video frame so that the video frame with the superimposed detection frame information can be directly displayed after the video frame is sent.

7. A video synchronization processing device, characterized in that, The device includes: The detection frame determination module is used to determine the current detection frame information and the next detection frame information of the video frame; the current detection frame and the next detection frame are two consecutive detection frames obtained by performing target detection on different video frames at an interval of at least one video frame. The target estimation module is used to estimate the target detection frame information of video frames that have not been detected between the current detection frame and the next detection frame based on the current detection frame information and the next detection frame information. The detection frame matching module is used to add matching detection frame information to the video frame so that the target object box indicated by the matching detection frame information of the video frame can be displayed synchronously when the video frame is displayed. The detection frame determination module includes: While buffering the first video frame derived from the source video stream, target detection is performed on the second video frame derived from the source video stream at least once every other video frame. The detection frame information obtained by target detection of the second video frame is matched with the cached first video frame to obtain the detection frame information that matches some video frames in the first video frame and cached; the first video frame and the second video frame belong to the same source video frame; Two consecutive detection frame information are obtained from the cache and used as the current detection frame information and the next detection frame information of the video stream, respectively; wherein, at least one video frame between the two consecutive detection frames includes a first value of video frames; the first value is determined by rounding the ratio of the video frame rate to the detection rate of the target detection.

8. An electronic device, characterized in that, include: One or more processing devices; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the video synchronization processing method according to any one of claims 1-6.

9. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processing device, it implements the video synchronization processing method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method and device for detecting objects in video and computer storage medium

    CN108256506A

  • Target frame and video frame synchronous display method, system and device and medium

    CN113296723A