Multi-channel video frame-by-frame synchronous playing system and method and storage medium

By employing a time synchronization algorithm and an intelligent pre-loading cache cleanup mechanism in the in-vehicle multi-channel video playback system, frame-by-frame accurate synchronization and smooth playback of dual-channel video in the vehicle are achieved. This solves the problems of insufficient synchronization accuracy and low cache scheduling efficiency in existing technologies, and improves the reliability of event tracing and playback experience.

CN121940577APending Publication Date: 2026-04-28CHONGQING RUIMING INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING RUIMING INFORMATION TECH CO LTD
Filing Date
2026-01-30
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing in-vehicle multi-channel video playback systems have shortcomings in frame-by-frame synchronization accuracy and cache scheduling efficiency, resulting in insufficient picture synchronization accuracy, frame-by-frame playback stuttering, and low reliability of event tracing. In particular, in reverse frame-by-frame scenarios, untimely loading of I-frames leads to decoding failures.

Method used

A multi-channel video frame-by-frame synchronous playback system is adopted. The time synchronization algorithm of the playback control layer accurately locates the synchronization time point of multiple channels. Combined with the intelligent preloading and cache clearing mechanism of the data processing layer, it can achieve frame-by-frame accurate synchronization and smooth playback of dual-channel video in the vehicle.

Benefits of technology

It improves the ability to accurately synchronize video frames one by one in vehicle monitoring, ensuring that video frames from different cameras are accurately synchronized in a frame-by-frame retrospective scenario, thereby improving the accuracy of event tracing and the stability of the playback experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940577A_ABST
    Figure CN121940577A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicle-mounted video processing, and discloses a multi-channel video frame-by-frame synchronous playing system and method and a storage medium, and the system comprises a vehicle-mounted terminal which is used for collecting a video stream and transmitting the video stream to a data processing layer; the playing control layer is used for receiving and processing a playing instruction triggered by a user, determining a synchronization time point of multiple paths of video streams and coordinating acquisition and rendering scheduling of frame data; the data processing layer is used for identifying different types of video frames in combination with a GOP structure and executing an intelligent preloading and cache cleaning algorithm on the frame data of the multiple paths of video streams; according to a scheduling instruction of the play control layer, frame data with time synchronization alignment is generated; the user display layer is used for receiving and synchronously rendering and displaying the frame data from the data processing layer; the time synchronization precision and the playing fluency of the frame-by-frame playing of the vehicle-mounted multi-channel video are effectively improved, and the cache management efficiency of the frame data is optimized at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of vehicle-mounted video processing technology, and in particular to a multi-channel video frame-by-frame synchronous playback system, method, and storage medium. Background Technology

[0002] As the core support for vehicle safety tracing and incident evidence collection, the vehicle-mounted multi-channel video surveillance system's ability to play multi-channel video frame by frame and accurately align time is a key technical requirement in scenarios such as accident retrospection and behavior analysis.

[0003] Existing technical solutions have several shortcomings: Firstly, most systems can only achieve macroscopic timeline alignment of multi-channel video, making it difficult to meet high-precision time synchronization at the frame-by-frame level. When there are video frame drops, frame rate fluctuations, or storage delays, the timestamp discrepancies between different channels can lead to image misalignment, making it impossible to accurately restore event details. Secondly, traditional cache management strategies only clean up data based on fixed time windows and do not incorporate preloading logic based on video encoding dependency characteristics, resulting in frequent stuttering and frame drops during frame-by-frame playback. This is especially pronounced in reverse frame-by-frame scenarios, where decoding failures caused by untimely loading of I-frames are more significant. Furthermore, while some existing technologies introduce time synchronization algorithms, they can only handle alignment requirements at normal frame rates and lack adaptive compensation capabilities for complex situations common in in-vehicle scenarios, such as abnormal frame intervals and accelerated playback, failing to guarantee accurate traceability of timestamps and actual recording times.

[0004] Therefore, there is an urgent need for a multi-channel video frame-by-frame synchronous playback system, method, and storage medium to achieve accurate time alignment of multiple channels and adapt to video playback systems in complex vehicle scenarios. This would overcome the shortcomings of existing technologies in terms of frame-by-frame synchronization accuracy, cache scheduling efficiency, and adaptability to abnormal scenarios, and meet the high reliability requirements for tracing vehicle monitoring events. Summary of the Invention

[0005] In view of this, the present invention aims to propose a multi-channel video frame-by-frame synchronous playback system, method and storage medium to solve the problems of insufficient picture synchronization accuracy, frame-by-frame playback stuttering and low reliability of event tracing caused by the lack of multi-channel frame-by-frame time alignment mechanism and the failure of cache scheduling to adapt to encoding dependency features in existing in-vehicle multi-channel video playback systems, thereby achieving accurate frame-by-frame synchronization and smooth playback of dual-channel in-vehicle video.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: A multi-channel video frame-by-frame synchronous playback system includes: The vehicle-mounted terminal includes a front-view camera and an in-vehicle camera; used to capture video streams and transmit them to the data processing layer. The playback control layer is used to receive and process frame-by-frame playback commands triggered by the user, determine the synchronization time points of multiple video streams based on the time synchronization algorithm, and coordinate the data processing layer to perform frame data acquisition and rendering scheduling. The data processing layer is communicatively connected to both the vehicle-mounted terminal and the playback control layer. It is used to identify video frames by combining the GOP structure, and to perform intelligent preloading and cache clearing algorithms on the frame data of multiple video streams based on the encoding dependency characteristics of each type of video frame. According to the scheduling instructions of the playback control layer, it generates time-synchronized frame data of the two channels of the front-view camera and the in-vehicle camera, and feeds it back to the playback control layer. The user display layer is communicatively connected to the data processing layer; it is used to receive and synchronously render and display frame data from the data processing layer.

[0007] The beneficial effects of this solution are as follows: In existing technologies, in-vehicle multi-channel video playback systems generally suffer from insufficient precision in achieving macroscopic timeline alignment and frame-by-frame synchronization, and the cache scheduling is not adapted to encoding dependency characteristics, resulting in frame-by-frame playback stuttering and unclear details in event tracing, which restricts the reliability of in-vehicle monitoring in scenarios such as accident backtracking. This system accurately locates the synchronization time points of multiple channels through the time synchronization algorithm of the playback control layer, and combines it with the intelligent preloading and cache clearing mechanism based on encoding dependency characteristics in the data processing layer, achieving accurate frame-by-frame synchronization and smooth playback of in-vehicle dual-channel video, effectively improving the accuracy of event tracing and the stability of the playback experience.

[0008] Furthermore, the playback control layer includes: The state manager is used to monitor the operational status of the system and manage the operation sequence through a token mechanism to prevent data conflicts. The frame-by-frame controller is used to receive frame-by-frame playback instructions triggered by the user and convert the instructions into frame-by-frame control signals for processing by the time synchronizer. A time synchronizer is used to align the time points of the video stream collected by the vehicle terminal during frame-by-frame playback based on a preset time synchronization algorithm.

[0009] Beneficial effects: The token mechanism of the state manager can effectively avoid data conflicts when multiple operations are concurrent, ensuring the stability of system operation. At the same time, the frame-by-frame controller can accurately respond to and convert the user's frame-by-frame playback instructions. Combined with the preset time synchronization algorithm of the time synchronizer, the time alignment accuracy of multi-channel video frame-by-frame playback can be further improved, ensuring that video frames from different cameras can be accurately synchronized in frame-by-frame backtracking scenarios, providing more reliable technical support for event tracing.

[0010] Furthermore, the time synchronizer executes the time synchronization algorithm process including: Obtain the timestamps of the current video frames to be processed from the front-view camera and the in-vehicle camera; Based on the frame-by-frame playback direction, determine the synchronization time point for time alignment between the two video channels; if it is a forward frame-by-frame playback, select the maximum value of the two timestamps as the synchronization time point; if it is a reverse frame-by-frame playback, select the minimum value of the two timestamps as the synchronization time point. For each video channel, find the video frame whose timestamp is closest to the synchronization time point, and complete the basic time alignment of the video frames of the two channels. Frame time difference detection is performed on the two video frames after basic time alignment, and synchronization error compensation is performed based on the detected time deviation value to achieve high-precision alignment.

[0011] Beneficial effects: The time synchronization algorithm obtains the timestamps of video frames from both cameras and dynamically determines the synchronization time point based on the forward or reverse playback direction. It can accurately locate the most matching video frame in each channel. Combined with frame time difference detection and synchronization error compensation, it can achieve high-precision frame-by-frame alignment of multi-channel video frames. This not only solves the limitation of traditional solutions that can only achieve macro-time alignment, but also avoids image misalignment caused by changes in playback direction, effectively improving the detailed restoration capability of vehicle monitoring event tracing.

[0012] Furthermore, the data processing layer includes: A multi-channel cache manager is used to cache video frames from the front-view camera and the in-vehicle camera into a frame sequence format based on the GOP structure, and to execute intelligent preloading algorithm and cache cleanup algorithm in combination with encoding dependency features. The GOP data processor is used to retrieve GOP frame sequence data from the multi-channel buffer manager, identify the frame type and detect GOP boundaries based on the feature identification information of the frame encoding header, and decode the single-frame video data corresponding to the target synchronization time point; the frame types include I-frames, P-frames and B-frames. The frame data renderer is used to process the time-synchronized video frames from the decoded front-view camera and in-vehicle camera channels according to the scheduling instructions of the playback control layer, generate frame data for direct rendering by the user display layer, and feed it back to the playback control layer.

[0013] Beneficial effects: By using a multi-channel cache manager to cache frame sequences based on the GOP structure and combining encoding dependency features to perform intelligent preloading and cache cleanup, key I-frames can be obtained in advance and expired caches can be released efficiently, avoiding stuttering and decoding delays during frame-by-frame playback. At the same time, the GOP data processor can accurately identify frame types and GOP boundaries, quickly decode the video frames corresponding to the target synchronization time points, and then, after format processing by the frame data renderer, directly output frame data adapted to the display layer, further ensuring the synchronous rendering efficiency and smoothness of multi-channel video frames, providing solid underlying support for high-precision time-aligned video playback.

[0014] Furthermore, the process of the multi-channel cache manager executing the intelligent preloading algorithm includes: Detect the current amount of cached data; if it is less than a preset threshold, start the preloading process. If the playback is in a forward frame-by-frame manner and the current playback position is at the end boundary of the current GOP, then the I-frame data and the complete frame sequence of the next GOP are preloaded. If the playback is in reverse frame-by-frame and the current playback position is at the beginning boundary of the current GOP, then the I-frame data of the previous GOP is preloaded, and then the remaining frame data of the GOP is loaded.

[0015] Beneficial effects: By monitoring the amount of cached data in real time and dynamically triggering preloading, and combining the forward and reverse playback directions with the current GOP position for precise frame sequence preloading, the complete frame sequence of the next GOP is obtained in advance during forward playback, and the key I-frames of the previous GOP are loaded first during reverse playback before supplementing the remaining frame data. This not only avoids playback stuttering caused by insufficient cache, but also solves the decoding failure problem caused by untimely loading of I-frames during reverse frame-by-frame playback, effectively improving the smoothness and response speed of multi-channel video frame-by-frame playback.

[0016] Furthermore, the process of the multi-channel cache manager executing the cache cleanup algorithm includes: Traverse the buffered frame sequences of the front-view camera and the in-vehicle camera; If the timestamp of the frame data in the cached frame sequence is earlier than the difference between the current playback time and the preset cache window duration, then the frame data is determined to be expired frame data; and the expired frame data is cleaned up in units of GOP.

[0017] Beneficial effects: By traversing the cached frame sequence of dual cameras, expired frame data is accurately identified based on timestamps combined with preset cache window duration. Cache cleanup is performed in units of Groups of Pictures (GOPs), which avoids the inefficiency of single-frame cleanup and ensures that only valid frame data required for current playback is retained in the cache. This effectively optimizes the utilization efficiency of cache resources, prevents redundant cache accumulation from affecting the retrieval and decoding speed of frame data, and works in conjunction with the intelligent preloading algorithm to make cache management more targeted, further ensuring the smoothness of multi-channel video frame-by-frame playback.

[0018] Furthermore, the time synchronizer also executes an adaptive frame time calculation algorithm to accurately calculate the target frame time when the video frame interval is irregular. This algorithm includes: Obtain the timestamp and frame rate information of the current video frame to be processed on the vehicle-mounted device; The theoretical frame interval of the video frame to be processed is calculated based on the frame rate information. The actual frame interval mode of the video stream at the vehicle terminal is detected. If the actual frame interval mode is an abnormal mode, the frame extraction interval is used as a candidate frame interval. If the actual frame interval mode is a normal mode, the theoretical frame interval is used as a candidate frame interval. Based on the candidate frame interval, the actual frame interval is determined in combination with the playback speed requirements of the video stream; if the playback speed is greater than or equal to 2, the actual frame interval is recalculated based on the GOP length; if the playback speed is equal to 1, the current actual frame interval remains unchanged. When moving forward frame by frame to the last frame of the current second, if the decimal part of the timestamp calculated based on the actual frame interval exceeds a preset threshold, the time will be forcibly rounded up to the next second, and the precise timestamp of the target frame will be output.

[0019] Beneficial effects: By obtaining video frame timestamps and frame rate information, the theoretical frame interval is calculated, and the actual frame interval pattern of the video stream is detected to dynamically determine the candidate frame interval. It can also adapt and adjust the actual frame interval according to the needs of speed playback. At the same time, threshold verification and carry calibration are performed on the decimal part of the timestamp. This effectively solves the problem of target frame time calculation deviation caused by irregular video frame intervals and poor adaptability to speed playback in vehicle scenarios. It accurately outputs the precise timestamp of the target frame, makes up for the adaptation shortcomings of conventional time synchronization algorithms in abnormal frame interval scenarios, and further improves the accuracy of time alignment and scene adaptability when playing multi-channel video frame by frame, ensuring the accuracy of frame time calculation in various playback scenarios.

[0020] Furthermore, a method for synchronous playback of multiple video streams frame by frame, the method comprising: S1. Collect multiple video streams through the vehicle's front-view camera and in-vehicle camera; S2. Receive the frame-by-frame operation command triggered by the user, and calculate the synchronization time point of each video stream through the time synchronizer. S3. Based on the synchronization time point, obtain the target frame data from the corresponding video stream, and perform decoding and rendering processing based on the GOP structure; S4. The user display layer synchronously renders and displays the processed multiple video frames.

[0021] Beneficial effects: By standardizing the steps, the system achieves the acquisition of dual-channel video streams in vehicles, the calculation of synchronous time points, the decoding and rendering of frame data based on the GOP structure, and the synchronous display. The process is simple and easy to implement, and can efficiently achieve frame-by-frame time alignment and smooth playback of dual-channel video. It is suitable for actual application scenarios of vehicle monitoring, and relies on the aforementioned system architecture to ensure execution accuracy, thereby improving the practicality of frame-by-frame backtracking and event tracing of vehicle monitoring videos.

[0022] Furthermore, a computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the multi-channel video frame-by-frame synchronous playback method. Attached Figure Description

[0023] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein: Figure 1 This is an exemplary system architecture diagram of a multi-channel video frame-by-frame synchronous playback system; Figure 2 This is an exemplary architecture diagram of a multi-channel cache manager; Figure 3 This is an exemplary flowchart of the time synchronization algorithm used by the time synchronizer; Figure 4 This is an exemplary flowchart of a method for synchronous playback of multiple video streams frame by frame. Detailed Implementation

[0024] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0025] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0026] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0027] The following detailed explanation illustrates the specific implementation methods: Example 1: Figure 1 This is an exemplary system architecture diagram of a multi-channel video frame-by-frame synchronous playback system, such as... Figure 1 As shown, a multi-channel video frame-by-frame synchronous playback system (hereinafter referred to as the playback system) includes: The vehicle-mounted component includes a front-view camera and an in-vehicle camera; it is used to capture video streams and transmit them to the data processing layer.

[0028] A vehicle-mounted terminal refers to a collection of hardware devices installed inside or outside a vehicle to collect video data about the vehicle's surroundings or interior. It typically includes one or more cameras, such as a front-view camera and an in-vehicle camera, to capture video information from different perspectives. These cameras transmit the raw video streams to subsequent data processing layers.

[0029] As one implementation, the vehicle-mounted terminal can use standard digital video interfaces such as MIPI CSI-2 or LVDS to transmit video data to the data processing layer; or it can use network protocols such as RTSP or RTMP for streaming. In another implementation, the vehicle-mounted terminal can store the video data on local storage media, and the data processing layer can then retrieve the video data via file reading.

[0030] Video streams are used to represent moving images. In this embodiment, the video stream is captured and encoded in real time by a camera for storage and transmission.

[0031] The playback control layer is used to receive and process frame-by-frame playback commands triggered by the user, determine the synchronization time points of multiple video streams based on the time synchronization algorithm, and coordinate the data processing layer to perform frame data acquisition and rendering scheduling.

[0032] The playback control layer refers to the module in this playback system responsible for receiving the user's frame-by-frame playback instructions, coordinating the playback process, and managing time synchronization.

[0033] A frame-by-frame playback command is a command triggered by the user through the user interface, which requests the playback system to play the video forward or backward frame by frame.

[0034] In this embodiment, the playback control layer receives frame-by-frame playback commands triggered by the user on the operation interface through its built-in frame-by-frame controller, and converts the received frame-by-frame playback commands into standardized frame-by-frame control signals adapted to the time synchronizer, thereby realizing effective reception and format conversion of commands.

[0035] The state manager monitors the system's operational status and manages operation sequences through a token mechanism to prevent data conflicts.

[0036] In this embodiment, the playback system is initially in standby mode, waiting for the user to initiate a frame-by-frame playback operation. When the user triggers the frame-by-frame playback command, the system enters the frame-by-frame startup phase, generates a unique operation token, and then sets the waiting and enable flags to true to indicate that it is currently in a state of waiting for data and is allowed to process, and sends a data request for the target frame to the data processing layer.

[0037] After receiving frame data, the data processing layer performs token verification. If the token does not match, the data request is deemed invalid and the data is discarded to avoid data conflicts caused by concurrent operations. If the token matches, the frame data is decoded and formatted to complete frame rendering. After the data processing process is completed, the system enters a waiting phase, setting `waiting` to `false` and resetting relevant status flags to prepare for the next user operation.

[0038] The frame-by-frame controller receives user-triggered frame-by-frame playback commands and converts these commands into frame-by-frame control signals for processing by the time synchronizer.

[0039] A time synchronizer is used to align the time points of the video stream captured by the vehicle terminal during frame-by-frame playback based on a preset time synchronization algorithm.

[0040] A time synchronization algorithm refers to the logic and methods used to calculate and determine the frame alignment of multiple video streams at the synchronization time point. Its goal is to eliminate time deviations between different video channels and ensure that the images from each channel are precisely synchronized during frame-by-frame playback.

[0041] Synchronization time point refers to a unified reference time point established on the timeline for all video channels to achieve precise frame-by-frame alignment of multiple video streams.

[0042] In this embodiment, further, as Figure 3 As shown, the time synchronization algorithm includes: Obtain the timestamps of the current video frames to be processed from the front-view camera and the in-vehicle camera.

[0043] In this embodiment, when the vehicle-mounted terminal acquires video streams, the front-view camera and the in-vehicle camera synchronously record the acquisition timestamp of each frame by a hardware encoder during the encoding and generation of each frame of video data. This timestamp is based on a unified system clock, ensuring the consistency of the time base of the two video frames. When the playback control layer initiates a frame-by-frame playback command, the data processing layer directly retrieves the timestamp information of the corresponding frame from the cached frame sequence metadata, or obtains the embedded timestamp field by parsing the encoded header of the video frame, thereby obtaining the timestamp of the current video frame to be processed in both video channels.

[0044] Based on the frame-by-frame playback direction, determine the synchronization point for time alignment between the two video channels; if it is a forward frame-by-frame playback, select the maximum value of the two timestamps as the synchronization point; if it is a reverse frame-by-frame playback, select the minimum value of the two timestamps as the synchronization point.

[0045] For each video channel, find the video frame whose timestamp is closest to the synchronization time point, and complete the basic time alignment of the video frames of the two channels.

[0046] In this embodiment, when the front-view camera and the in-vehicle camera acquire and transmit video streams, they will add a timestamp to each frame of video data based on a unified system clock, and generate a frame time index table containing all frame timestamps in chronological order. This index table is equivalent to a time directory of frames, which can support quick location of the video frame corresponding to a certain time point.

[0047] As an example, when the playback control layer determines the synchronization time point to be 10:00:02 based on the frame-by-frame playback direction, the playback system will initiate searches in the frame time index tables of the front-view camera and the in-vehicle camera respectively. It will find the frame with timestamp 10:00:02.122 in the index table of the front-view camera and the frame with timestamp 10:00:02.124 in the index table of the in-vehicle camera. The timestamps of these two frames are closest to the synchronization time point of 10:00:02, thereby completing the basic time alignment of the two video frames and providing the initial aligned frame data for subsequent synchronization error compensation.

[0048] Frame time difference detection is performed on the two video frames after basic time alignment, and synchronization error compensation is performed based on the detected time deviation value to achieve high-precision alignment.

[0049] In this embodiment, after basic time alignment is completed, the system first calculates the difference between the timestamps of the two target frames to detect their time deviation. Taking the timestamp of the front-view camera frame (10:00:02.122) and the in-vehicle camera frame (10:00:02.124) after basic alignment as an example, the calculated time deviation between the two frames is 0.002 seconds. If the deviation is less than a preset threshold, the current alignment result is directly retained; if the deviation exceeds the preset threshold, a synchronization error compensation mechanism is triggered to eliminate the deviation by adjusting the rendering sequence of the frames. For example, the in-vehicle camera frame with a slightly later timestamp is rendered 0.002 seconds earlier, or the front-view camera frame with a slightly earlier timestamp is delayed by 0.002 seconds, so that the two frames are rendered and output at the same time, thereby achieving high-precision alignment of the dual-channel video frames.

[0050] In this embodiment, by acquiring the timestamps of the video frames from the dual cameras respectively and dynamically determining the synchronization time point according to the forward or reverse playback direction, the most matching video frame in each channel can be accurately located. Combined with frame time difference detection and synchronization error compensation, high-precision frame-by-frame alignment of multi-channel video frames can be achieved. This not only solves the limitation of traditional solutions that can only achieve macro-time alignment, but also avoids image misalignment caused by changes in playback direction, effectively improving the detailed restoration capability of vehicle monitoring event tracing.

[0051] The data processing layer communicates with both the vehicle-mounted terminal and the playback control layer. It is used to identify different types of video frames by combining the GOP structure, and to perform intelligent preloading and cache clearing algorithms on the frame data of multiple video streams based on the encoding dependency characteristics of each type of video frame. According to the scheduling instructions of the playback control layer, it generates time-synchronized frame data of the two channels of the front-view camera and the in-vehicle camera, and feeds it back to the playback control layer.

[0052] The data processing layer refers to the core functional module in the playback system responsible for receiving, processing, and managing video data.

[0053] Encoding dependency features refer to the overall encoding structure features of a GOP starting with an I-frame, as well as the decoding dependency features between I-frames, P-frames, and B-frames within a GOP; including the encoding dependency of P-frames on forward I-frames or P-frames, the encoding dependency of B-frames on forward and backward I-frames or P-frames, and also covering the encoding association attributes between adjacent GOPs with I-frames as the boundary.

[0054] For example, the forward encoding dependency of P-frames. In a GOP of a forward-looking camera, when the I-frame is the first frame and the third frame is a P-frame, the P-frame only records pixel changes and motion trajectory information compared with the forward reference frame, i.e., the first I-frame or the previous P-frame in the same GOP. It does not have complete image data. During decoding, its forward reference frame must be retrieved and parsed first in order to reconstruct the complete video image of the third frame. If the reference frame is missing, the P-frame cannot be decoded and displayed independently.

[0055] For example, there is the bidirectional coding dependency of B-frames. In a certain GOP of an in-vehicle camera, when the 5th frame is a B-frame, this B-frame will simultaneously record the pixel difference information with the forward reference frame 3 (P-frame) and the backward reference frame 7 (I-frame). During decoding, the complete data of the two reference frames must be retrieved simultaneously to complete the image reconstruction of this B-frame.

[0056] A Group of Pictures (GOP) is a basic unit in video coding, consisting of one I-frame and several P-frames and B-frames. The I-frame is the keyframe; the P-frame is the prediction frame; and the B-frame is the bidirectional prediction frame.

[0057] Intelligent preloading algorithm refers to a strategy to optimize cache management. Based on information such as video playback direction, current playback position, and GOP structure, it predictively loads video frame data that may be played later into the cache in advance to reduce playback latency and stuttering.

[0058] A cache cleanup algorithm is a strategy for managing cache space by identifying and removing expired frame data that is no longer needed, thereby freeing up storage resources.

[0059] Furthermore, the data processing layer includes: The multi-channel cache manager is used to cache video frames from the front-view camera and the in-vehicle camera video streams into a frame sequence format based on the GOP structure, and to execute intelligent preloading algorithm and cache cleanup algorithm in combination with encoding dependency features.

[0060] The GOP data processor is used to retrieve GOP frame sequence data from the multi-channel buffer manager, identify frame types and detect GOP boundaries based on the feature identification information of the frame encoding header, and decode the single-frame video data corresponding to the target synchronization time point; the frame types include I-frames, P-frames and B-frames.

[0061] Feature identification information refers to the standardized marker fields embedded in the frame encoding header. These include the frame type feature code, GOP boundary identifier, and frame sequence number, and can be directly obtained by parsing the encoding header.

[0062] In this embodiment, the playback system extracts frame type feature codes and distinguishes between I, P, and B frames according to preset code values. It also detects boundaries using boundary markers and sequence numbers. The start marker and start number identify the GOP start boundary, and the end marker and end number identify the end boundary. If no boundary marker is found, the new I frame and its reset sequence number are used to determine the next GOP start boundary and the previous frame the previous GOP end boundary.

[0063] The frame data renderer is used to process the time-synchronized video frames from the decoded front-view camera and in-vehicle camera channels according to the scheduling instructions of the playback control layer, generate frame data for direct rendering by the user display layer, and feed it back to the playback control layer.

[0064] In this embodiment, Figure 2 This is an exemplary architecture diagram of a multi-channel cache manager, such as... Figure 2 As shown, the entire architecture is divided into a cache management strategy layer on the left and dual-camera cache instances on the right. The cache management strategy layer, with GOP structure optimization at its core, drives two key modules: intelligent preloading algorithm and cache cleanup algorithm. Each camera cache instance contains a frame data list (frameList), cache state management, and a renderer reference (renderer).

[0065] Furthermore, the intelligent preloading algorithm includes: The current amount of cached data is checked, and if it is less than a preset threshold, the preloading process is started.

[0066] If the playback is in a forward, frame-by-frame manner and the current playback position is at the end boundary of the current GOP, then the I-frame data and the complete frame sequence of the next GOP are preloaded.

[0067] If the playback is in reverse frame-by-frame and the current playback position is at the beginning boundary of the current GOP, then the I-frame data of the previous GOP is preloaded, and then the remaining frame data of the GOP is loaded.

[0068] The end boundary of a GOP refers to the position of the last frame of a GOP, which is the dividing point between the frame sequence of the current GOP and the frame sequence of the next GOP; the start boundary of a GOP refers to the starting position of a GOP, which is the dividing point between the current GOP and the previous GOP.

[0069] In this embodiment, by real-time detection of cached data volume and dynamic triggering of preloading, and by combining the forward and reverse playback directions with the boundary position of the current GOP, precise frame sequence preloading is performed. When playing forward to the end boundary of the current GOP, the I-frame and complete frame sequence of the next GOP are obtained in advance. When playing backward to the start boundary of the current GOP, the I-frame of the previous GOP is loaded first, and then the remaining frame data is supplemented. This not only avoids playback stuttering caused by insufficient cache, but also solves the decoding failure problem caused by untimely loading of I-frames when playing backward frame by frame. This effectively improves the smoothness and response speed of multi-channel video frame by frame playback. At the same time, preloading in units of GOP also makes the utilization of cached resources more targeted.

[0070] Furthermore, cache cleanup algorithms include: Iterate through the cached frame sequences of the front-view camera and the in-vehicle camera.

[0071] If the timestamp of the frame data in the cached frame sequence is earlier than the difference between the current playback time and the preset cache window duration, the frame data is determined to be expired frame data; and expired frame data is cleaned up in units of GOP.

[0072] The preset buffer window duration refers to a fixed time threshold set by the playback system for buffer management. It is set based on the actual application scenario and hardware performance of the vehicle monitoring system, for example, 30 seconds to 5 minutes.

[0073] In this embodiment, the cache cleanup algorithm avoids the inefficiency of single-frame cleanup, reduces cache redundancy, optimizes cache resource utilization efficiency, prevents invalid data from affecting frame data retrieval and decoding speed, and works in conjunction with the intelligent preloading algorithm to form a closed-loop cache management, ensuring the smoothness and response efficiency of multi-channel video playback frame by frame.

[0074] The user display layer communicates with the data processing layer; it is used to receive and synchronously render and display frame data from the data processing layer.

[0075] The user display layer may include a display screen and related rendering software, used to present synchronized video footage to the user.

[0076] In this embodiment, the user display layer receives frame data processed by the data processing layer through a preset data communication interface. This interface adopts an asynchronous data transmission method, which can receive dual-channel video frame decoding data after time synchronization and error compensation in real time, and perform format verification on the received frame data to ensure data integrity. During the rendering stage, the user display layer calls the rendering engine adapted to the vehicle display terminal to map the frame data of the front-view camera and the in-vehicle camera to the corresponding display area of ​​the display interface, and triggers the synchronous rendering of dual-channel frame data based on a unified rendering clock.

[0077] In this embodiment, the playback system accurately locates the synchronization time points of multiple channels through the time synchronization algorithm of the playback control layer, and combines the intelligent preloading and cache clearing mechanism based on the encoding dependency characteristics of the data processing layer to achieve frame-by-frame accurate synchronization and smooth playback of dual-channel video in the vehicle, effectively improving the accuracy of event tracing and the stability of the playback experience.

[0078] Example 2: Building upon Example 1, to accurately calculate the target frame time when video frame intervals are irregular, the time synchronizer further executes an adaptive frame time calculation algorithm, which includes: The system acquires the timestamp and frame rate information of the current video frame to be processed on the vehicle terminal; calculates the theoretical frame interval of the video frame to be processed based on the frame rate information; detects the actual frame interval mode of the video stream on the vehicle terminal; if the actual frame interval mode is an abnormal mode, the frame extraction interval is used as a candidate frame interval; if the actual frame interval mode is a normal mode, the theoretical frame interval is used as a candidate frame interval; based on the candidate frame interval, the actual frame interval is determined in combination with the playback speed requirements of the video stream; if the playback speed is greater than or equal to 2, the actual frame interval is recalculated based on the GOP length; if the playback speed is equal to 1, the current actual frame interval remains unchanged; when moving forward frame by frame to the last frame of the current second, if the decimal part of the timestamp calculated based on the actual frame interval exceeds a preset threshold, the time is forcibly rounded up to the next second, and the accurate timestamp of the target frame is output.

[0079] In this embodiment, the playback system can calibrate the timestamp using the vehicle-mounted unified system clock, embed a specified field in the frame encoding header, and parse it to read the timestamp of the current video frame to be processed. The frame rate is an inherent encoding parameter of the video stream, pre-stored in the video stream encapsulation header or frame encoding header parameter field, or it can be obtained by calculating and verifying the difference between consecutive frame timestamps.

[0080] In this embodiment, the theoretical frame interval is the standard time difference between two adjacent video frames. It is calculated using the frame rate as the base and taking its reciprocal. That is: Theoretical frame interval = 1 / frame rate; the unit is seconds per frame.

[0081] In this embodiment, the actual frame interval mode is determined based on the difference in actual timestamps between adjacent frames of the video stream, including abnormal mode and normal mode.

[0082] In this context, "normal mode" refers to a situation where the difference between the actual timestamps of adjacent frames is consistent with or deviates from the theoretical frame interval within a preset reasonable threshold. "Abnormal mode" refers to a situation where the deviation between the actual timestamp difference of adjacent frames and the theoretical frame interval exceeds the preset reasonable threshold. The preset reasonable threshold is a threshold set based on industry standards for vehicle-mounted video surveillance data transmission and encoding / decoding, and is determined by domain personnel based on experience.

[0083] In this embodiment, the playback operation of the in-vehicle video stream may require speed adjustment. Based on the candidate frame interval, if the playback multiplier is greater than or equal to 2, the actual frame interval is recalculated based on the GOP length. The calculation logic is: actual frame interval = GOP length / 1000. If the playback multiplier is equal to 1, the current actual frame interval remains unchanged.

[0084] In this embodiment, due to the slight deviation in frame interval during the encoding process of the vehicle video stream and the cumulative error of timestamp after playback at double speed, there are boundary conditions for second-level time nodes when the vehicle video stream is played back frame by frame. In order to avoid the target frame time positioning deviation caused by the accumulation of the decimal part of the timestamp, it is necessary to ensure that the frame timestamp corresponds accurately with the actual playback time.

[0085] Therefore, during forward frame-by-frame playback, it is monitored in real time whether the current frame is the last frame of the current second. For the last frame of that second, its final timestamp is recalculated based on the actual frame interval. The decimal part of the recalculated timestamp is extracted and compared with a preset threshold. If the decimal part exceeds the preset threshold, it means that the actual playback time of the frame is close to the next second. At this time, the integer part of the timestamp is forcibly incremented by 1 and the decimal part is set to 0 to complete the time carry-over process. If the decimal part does not exceed the preset threshold, the original timestamp after recalculation is retained. The preset threshold is set by those skilled in the art to be 0.8~0.9 seconds based on the time accuracy requirements of in-vehicle video.

[0086] In this embodiment, this processing method can effectively avoid the problem of timestamp ambiguity at the second-level boundary, ensure the accuracy of multi-channel video frame time synchronization, and meet the high-precision requirements for frame time positioning in vehicle monitoring scenarios.

[0087] Example 3: Figure 4 This is an exemplary flowchart of a method for synchronous frame-by-frame playback of multiple video streams. Figure 4 As shown, a method for synchronous playback of multiple video streams frame by frame includes: Step S1: Collect multiple video streams through the vehicle's front-view camera and in-vehicle camera.

[0088] Step S2: Receive the frame-by-frame operation command triggered by the user, and calculate the synchronization time point of each video stream through the time synchronizer.

[0089] Step S3: Based on the synchronization time point, obtain the target frame data from the corresponding video stream and perform decoding and rendering processing based on the GOP structure.

[0090] Step S4: The user display layer synchronously renders and displays the processed multi-channel video frames.

[0091] In this embodiment, the acquisition of dual-channel video streams in the vehicle, the calculation of synchronization time points, the decoding and rendering of frame data based on the GOP structure, and the synchronous display are achieved through standardized steps. The process is simple and easy to implement, and can efficiently achieve frame-by-frame time alignment and smooth playback of dual-channel video. It is suitable for actual vehicle monitoring application scenarios. Relying on the aforementioned system architecture, the execution accuracy is guaranteed, and the operability of frame-by-frame backtracking and event tracing of vehicle monitoring video is improved.

[0092] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method for synchronous playback of multiple video streams frame by frame.

[0093] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.

[0094] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this specification are not intended to limit the order of the processes and methods described herein. Although various examples have been discussed in the foregoing disclosure of some embodiments of the invention that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the spirit and scope of the embodiments described herein. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely using software solutions, such as installing the described system on existing servers or mobile devices.

[0095] Similarly, it should be noted that, in order to simplify the description disclosed herein and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of embodiments in this specification may sometimes combine multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of a single embodiment disclosed above.

[0096] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values ​​are set as precisely as feasible.

[0097] For each patent, patent application, patent application publication, and other material such as articles, books, specifications, publications, and documents referenced in this specification, the entire contents of which are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this specification, as well as documents that limit the broadest scope of the claims in this specification (currently or subsequently appended to this specification). It should be noted that in the event of any inconsistency or conflict between the descriptions, definitions, and / or terminology used in the supplementary materials to this specification and the content of this specification, the descriptions, definitions, and / or terminology used in this specification shall prevail.

[0098] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.

Claims

1. A multi-channel video frame-by-frame synchronous playback system, characterized in that, include: Vehicle-mounted devices include front-view cameras and in-vehicle cameras; Used to capture video streams and transmit them to the data processing layer; The playback control layer is used to receive and process frame-by-frame playback commands triggered by the user, determine the synchronization time points of multiple video streams based on the time synchronization algorithm, and coordinate the data processing layer to perform frame data acquisition and rendering scheduling. The data processing layer is communicatively connected to both the vehicle-mounted terminal and the playback control layer. It is used to identify video frames by combining the GOP structure, and to perform intelligent preloading and cache clearing algorithms on the frame data of multiple video streams based on the encoding dependency characteristics of each type of video frame. According to the scheduling instructions of the playback control layer, it generates time-synchronized frame data of the two channels of the front-view camera and the in-vehicle camera, and feeds it back to the playback control layer. The user display layer is communicatively connected to the data processing layer. Used to receive and synchronously render and display frame data from the data processing layer.

2. The multi-channel video frame-by-frame synchronous playback system according to claim 1, characterized in that, The playback control layer includes: The state manager is used to monitor the operational status of the system and manage the operation sequence through a token mechanism to prevent data conflicts. The frame-by-frame controller is used to receive frame-by-frame playback instructions triggered by the user and convert the instructions into frame-by-frame control signals for processing by the time synchronizer. A time synchronizer is used to align the time points of the video stream collected by the vehicle terminal during frame-by-frame playback based on a preset time synchronization algorithm.

3. The multi-channel video frame-by-frame synchronous playback system according to claim 2, characterized in that, The time synchronizer executes the time synchronization algorithm process including: Obtain the timestamps of the current video frames to be processed from the front-view camera and the in-vehicle camera; Based on the frame-by-frame playback direction, determine the synchronization time point for time alignment between the two video channels; if it is a forward frame-by-frame playback, select the maximum value of the two timestamps as the synchronization time point; if it is a reverse frame-by-frame playback, select the minimum value of the two timestamps as the synchronization time point. For each video channel, find the video frame whose timestamp is closest to the synchronization time point, and complete the basic time alignment of the video frames of the two channels. Frame time difference detection is performed on the two video frames after basic time alignment, and synchronization error compensation is performed based on the detected time deviation value to achieve high-precision alignment.

4. The multi-channel video frame-by-frame synchronous playback system according to claim 3, characterized in that, The data processing layer includes: A multi-channel cache manager is used to cache video frames from the front-view camera and the in-vehicle camera into a frame sequence format based on the GOP structure, and to execute intelligent preloading algorithm and cache cleanup algorithm in combination with encoding dependency features. The GOP data processor is used to retrieve GOP frame sequence data from the multi-channel buffer manager, identify the frame type and detect GOP boundaries based on the feature identification information of the frame encoding header, and decode the single-frame video data corresponding to the target synchronization time point; the frame types include I-frames, P-frames and B-frames. The frame data renderer is used to process the time-synchronized video frames from the decoded front-view camera and in-vehicle camera channels according to the scheduling instructions of the playback control layer, generate frame data for direct rendering by the user display layer, and feed it back to the playback control layer.

5. The multi-channel video frame-by-frame synchronous playback system according to claim 4, characterized in that, The process by which the multi-channel cache manager executes the intelligent preloading algorithm includes: Detect the current amount of cached data; if it is less than a preset threshold, start the preloading process. If the playback is in a forward frame-by-frame manner and the current playback position is at the end boundary of the current GOP, then the I-frame data and the complete frame sequence of the next GOP are preloaded. If the playback is in reverse frame-by-frame and the current playback position is at the beginning boundary of the current GOP, then the I-frame data of the previous GOP is preloaded, and then the remaining frame data of the GOP is loaded.

6. The multi-channel video frame-by-frame synchronous playback system according to claim 5, characterized in that, The process of the multi-channel cache manager executing the cache cleanup algorithm includes: Traverse the buffered frame sequences of the front-view camera and the in-vehicle camera; If the timestamp of the frame data in the cached frame sequence is earlier than the difference between the current playback time and the preset cache window duration, then the frame data is determined to be expired frame data; and the expired frame data is cleaned up in units of GOP.

7. The multi-channel video frame-by-frame synchronous playback system according to claim 2, characterized in that, The time synchronizer also executes an adaptive frame time calculation algorithm to accurately calculate the target frame time when the video frame interval is irregular. This algorithm includes: Obtain the timestamp and frame rate information of the current video frame to be processed on the vehicle-mounted device; The theoretical frame interval of the video frame to be processed is calculated based on the frame rate information. The actual frame interval mode of the video stream at the vehicle terminal is detected. If the actual frame interval mode is an abnormal mode, the frame extraction interval is used as a candidate frame interval. If the actual frame interval mode is a normal mode, the theoretical frame interval is used as a candidate frame interval. Based on the candidate frame interval, the actual frame interval is determined in combination with the playback speed requirements of the video stream; if the playback speed is greater than or equal to 2, the actual frame interval is recalculated based on the GOP length; if the playback speed is equal to 1, the current actual frame interval remains unchanged. When moving forward frame by frame to the last frame of the current second, if the decimal part of the timestamp calculated based on the actual frame interval exceeds a preset threshold, the time will be forcibly rounded up to the next second, and the precise timestamp of the target frame will be output.

8. A method for synchronous playback of multiple video streams frame by frame, characterized in that, The method, applied to a multi-channel video frame-by-frame synchronous playback system as described in any one of claims 1 to 6, comprises: S1. Collect multiple video streams through the vehicle's front-view camera and in-vehicle camera; S2. Receive the frame-by-frame operation command triggered by the user, and calculate the synchronization time point of each video stream through the time synchronizer. S3. Based on the synchronization time point, obtain the target frame data from the corresponding video stream, and perform decoding and rendering processing based on the GOP structure; S4. The user display layer synchronously renders and displays the processed multiple video frames.

9. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the multi-channel video frame-by-frame synchronous playback method as described in claim 1.