Target detection method, system and equipment based on video frame synchronous reading

By configuring the NTP client and FFmpeg encoding thread for multi-camera devices, generating and synchronizing timestamp video frames, and using deep learning models for object detection, the problem of incomplete monitoring in industrial production of traditional single-camera monitoring methods is solved, and the monitoring abnormal recognition accuracy is improved.

CN120339893APending Publication Date: 2025-07-18CHONGQING UNIV OF ARTS & SCI
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202311374159.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-10-23
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In industrial production, traditional single camera monitoring methods are difficult to fully monitor the details of each process of the production line, and temporary occlusion of equipment or personnel from a single perspective cannot be identified in time, which affects the monitoring quality.

Method used

By configuring NTP clients, read threads and FFmpeg encoding threads for multiple camera devices in the LAN, generating and synchronizing timestamp video frames, using deep learning models for object detection, ensuring the consistency of timestamp difference values of video frames, and achieving synchronous reading and object detection of multi-camera video frames.

Benefits of technology

It improves the accuracy of monitoring abnormal recognition, ensures the time synchronization of video frames in multi-camera systems, and enhances the monitoring ability of industrial production processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339893A_ABST
    Figure CN120339893A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method, system and device based on video frame synchronous reading, and relates to the technical field of data extraction for target detection. The method is applied to a monitoring system, and comprises the following steps: determining a synchronous timestamp-containing video sequence according to a plurality of low-latency cache queues; timestamp difference values corresponding to any two timestamp-containing video frames in the synchronous timestamp-containing video sequence are smaller than a timestamp difference value threshold value; and inputting the synchronous timestamp-containing video sequence into the target detection model to obtain a target detection result corresponding to the center timestamp. According to the invention, the camera equipment is configured with one NTP client, one reading thread and one FFmpeg coding thread to obtain the synchronous video sequence containing the timestamp, and the target detection model is trained, so that the monitoring abnormity identification precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data extraction for target detection, and particularly to a target detection method, system and device based on synchronous reading of video frames. Background Art

[0002] With the continuous improvement of the level of industrial automation, video surveillance technology and intelligent algorithm detection have been widely applied in various links of industrial production, such as production line quality inspection and operation process monitoring. Moreover, the industrial environment has relatively high requirements for the real-time performance and accuracy of the application of video surveillance technology and intelligent algorithms. However, the complex production environment and conditions increase the difficulty of achieving this goal. In the traditional single-camera monitoring method, due to the limited information obtained from a single perspective in the industrial production process monitoring, it is difficult to comprehensively monitor the details of each process of the entire production line, and it is easy to miss some key problem points; and the temporary occlusion of equipment or personnel under a single perspective cannot be recognized in time, affecting the monitoring quality. Summary of the Invention

[0003] The object of the present invention is to provide a target detection method, system and device based on synchronous reading of video frames, which can improve the accuracy of monitoring anomaly recognition.

[0004] To achieve the above object, the present invention provides the following solutions:

[0005] A target detection method based on synchronous reading of video frames, the target detection method is applied to a monitoring system, the monitoring system includes: an NTP server and multiple camera devices; the NTP server and the multiple camera devices are all set in the same local area network; any one of the camera devices is configured with an NTP client, a reading thread and an FFmpeg encoding thread;

[0006] The NTP client is used to generate a timestamp when a video frame is generated by the corresponding camera device;

[0007] The NTP server is used to synchronize the timestamps generated by the multiple NTP clients at the same moment;

[0008] The reading thread is used to read the original video raw stream data of multiple video frames generated by the corresponding camera device and the timestamp corresponding to each video frame;

[0009] The FFmpeg encoding thread is used to encode the original video raw stream data of each video frame and its corresponding timestamp in chronological order to obtain multiple video frames with timestamps, and form a low-latency cache queue corresponding to the camera device;

[0010] The target detection method includes:

[0011] Determine a synchronized timestamp-containing video sequence according to multiple low-latency buffer queues; the timestamp difference between any two timestamp-containing video frames in the synchronized timestamp-containing video sequence is less than the timestamp difference threshold;

[0012] Input the synchronized timestamp-containing video sequence into a target detection model to obtain a target detection result corresponding to the central timestamp; the target detection model is obtained by training a deep learning model using multiple annotated synchronized timestamp-containing video historical sequences.

[0013] Optionally, the determining a synchronized timestamp-containing video sequence according to multiple low-latency buffer queues includes:

[0014] Extract one timestamp-containing video frame from each low-latency buffer queue in sequence according to the number order of the corresponding camera devices to obtain a pending synchronized timestamp-containing video sequence;

[0015] Judge whether there is a timestamp difference greater than the timestamp difference threshold between two timestamp-containing video frames in the pending synchronized timestamp-containing video sequence to obtain a first judgment result;

[0016] If the first judgment result is yes, update multiple low-latency buffer queues and return to the step "Extract one timestamp-containing video frame from each low-latency buffer queue in sequence according to the number order of the corresponding camera devices to obtain a pending synchronized timestamp-containing video sequence";

[0017] If the first judgment result is no, determine the pending synchronized timestamp-containing video sequence as the synchronized timestamp-containing video sequence.

[0018] Optionally, the updating of multiple low-latency buffer queues includes:

[0019] Determine the average timestamp corresponding to the pending synchronized timestamp-containing video sequence;

[0020] Determine two timestamp-containing video frames with a timestamp difference greater than the timestamp difference threshold as pending video frames;

[0021] Determine the pending video frame with the largest difference between the corresponding timestamp and the average timestamp as the reference video frame;

[0022] Judge whether the timestamp corresponding to the reference video frame is earlier than the average timestamp to obtain a second judgment result;

[0023] If the second judgment result is yes, delete the reference video frame from the low-latency buffer queue corresponding to the reference video frame;

[0024] If the second judgment result is no, determine the timestamp-containing video frames other than the reference video frame in the pending synchronized timestamp-containing video sequence as non-reference video frames;

[0025] Delete the corresponding non-reference video frames from the low-latency buffer queues corresponding to each non-reference video frame.

[0026] Optionally, before determining the synchronized timestamp-containing video sequence according to the multiple low-latency buffer queues, it further includes:

[0027] Obtain multiple low-latency buffer history queues

[0028] Determine a synchronized timestamp-containing video history sequence according to the multiple low-latency buffer history queues;

[0029] Label the targets in the synchronized timestamp-containing video history sequence respectively to obtain the labeled synchronized timestamp-containing video history sequence;

[0030] Use the labeled synchronized timestamp-containing video history sequence as the input and the target labeling result as the output to train a deep learning model to obtain a target detection model.

[0031] A target detection system based on synchronous reading of video frames, comprising:

[0032] A synchronized timestamp-containing video sequence acquisition module, configured to determine a synchronized timestamp-containing video sequence according to multiple low-latency buffer queues; the time difference between any two timestamp-containing video frames in the synchronized timestamp-containing video sequence is less than the timestamp difference threshold;

[0033] A target detection module, configured to input the synchronized timestamp-containing video sequence into the target detection model to obtain a target detection result corresponding to the central timestamp; the target detection model is obtained by training a deep learning model with multiple labeled synchronized timestamp-containing video history sequences.

[0034] An electronic device, comprising a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the above-mentioned target detection method based on synchronous reading of video frames.

[0035] Optionally, the memory is a readable storage medium.

[0036] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0037] A target detection method, system and device based on video frame synchronous reading provided by the present invention are applied to a monitoring system. According to a plurality of low-latency buffer queues, a synchronized video sequence with timestamps is determined; the timestamp difference between any two video frames with timestamps in the synchronized video sequence with timestamps is less than the timestamp difference threshold; the synchronized video sequence with timestamps is input into a target detection model to obtain a target detection result corresponding to the central timestamp. By configuring an NTP client, a reading thread and an FFmpeg encoding thread for each camera device to obtain a synchronized video sequence with timestamps and training the target detection model, the present invention can improve the accuracy of monitoring anomaly recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0039] Figure 1 It is a flowchart of the target detection method based on video frame synchronous reading in Embodiment 1 of the present invention;

[0040] Figure 2 It is a schematic diagram of the principle of the target detection method based on video frame synchronous reading in Embodiment 1 of the present invention;

[0041] Figure 3 It is a schematic diagram of the NTP architecture of the target detection method in Embodiment 1 of the present invention;

[0042] Figure 4 It is a flowchart of video frame acquisition in Embodiment 1 of the present invention;

[0043] Figure 5 It is a flowchart of video frame encoding in Embodiment 1 of the present invention;

[0044] Figure 6 It is a flowchart of video frame synchronization in Embodiment 1 of the present invention;

[0045] Figure 7 It is a flowchart of video frame synchronous inference in Embodiment 1 of the present invention

[0046] Figure 8 It is a schematic diagram of the structure of the terminal device in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0048] The purpose of the present invention is to provide a target detection method, system and device based on video frame synchronous reading, which can improve the accuracy of monitoring anomaly recognition.

[0049] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0050] Embodiment 1

[0051] This embodiment provides a target detection method based on video frame synchronous reading. The target detection method is applied to a monitoring system, and the monitoring system includes: an NTP server and multiple camera devices. The NTP server and multiple camera devices are both set in the same local area network. Any camera device is configured with an NTP client, a reading thread and an FFmpeg encoding thread. The NTP client is used to generate a timestamp when the corresponding camera device generates a video frame. The NTP server is used to synchronize the timestamps generated by multiple NTP clients at the same moment. The reading thread is used to read the original video raw stream data of multiple video frames generated by the corresponding camera device and the timestamp corresponding to each video frame. The FFmpeg encoding thread is used to encode the original video raw stream data of each video frame and its corresponding timestamp in chronological order to obtain multiple timestamp-containing video frames, and form a low-latency cache queue corresponding to the camera device.

[0052] As Figure 1 and Figure 2 shown, the target detection method includes:

[0053] Step 101: Determine a synchronized timestamp-containing video sequence according to multiple low-latency cache queues. The difference between the timestamps corresponding to any two timestamp-containing video frames in the synchronized timestamp-containing video sequence is less than the timestamp difference threshold.

[0054] Step 102: Input the synchronized timestamp-containing video sequence into the target detection model to obtain the target detection result corresponding to the central timestamp. The target detection model is obtained by training a deep learning model using multiple labeled synchronized timestamp-containing video historical sequences.

[0055] Step 101 includes:

[0056] Step 101-1: Extract one frame of timestamp-containing video from each low-latency cache queue in sequence according to the number order of the corresponding camera devices, and obtain a pending sequence of synchronized timestamp-containing videos.

[0057] Step 101-2: Determine whether there is a timestamp difference between two frames of timestamp-containing video frames in the pending sequence of synchronized timestamp-containing videos that is greater than the timestamp difference threshold, and obtain a first judgment result. If the first judgment result is yes, execute Step 101-3; if the first judgment result is no, execute Step 101-4.

[0058] Step 101-3: Update multiple low-latency cache queues, and return to Step 101-1.

[0059] Step 101-4: Determine the pending sequence of synchronized timestamp-containing videos as the synchronized timestamp-containing video sequence.

[0060] Step 101-3 includes:

[0061] Step 101-3-1: Determine the average timestamp corresponding to the pending sequence of synchronized timestamp-containing videos.

[0062] Step 101-3-2: Determine the two frames of timestamp-containing video frames with a timestamp difference greater than the timestamp difference threshold as pending video frames.

[0063] Step 101-3-3: Determine the pending video frame with the largest difference between the corresponding timestamp and the average timestamp as the reference video frame.

[0064] Step 101-3-4: Determine whether the timestamp corresponding to the reference video frame is earlier than the average timestamp, and obtain a second judgment result. If the second judgment result is yes, execute Step 101-3-5; if the second judgment result is no, execute Step 101-3-6.

[0065] Step 101-3-5: Delete the reference video frame from the low-latency cache queue corresponding to the reference video frame.

[0066] Step 101-3-6: Determine the timestamp-containing video frames in the pending sequence of synchronized timestamp-containing videos other than the reference video frame as non-reference video frames.

[0067] Step 101-3-7: Delete the corresponding non-reference video frames from the low-latency cache queues corresponding to each non-reference video frame.

[0068] Before Step 101, it also includes:

[0069] Step 103: Obtain multiple low-latency cache history queues.

[0070] Step 104: Determine the synchronized video history sequence with timestamps based on multiple low-latency cache history queues.

[0071] Step 105: Label the targets in the synchronized video history sequence with timestamps respectively to obtain the labeled synchronized video history sequence with timestamps.

[0072] Step 106: Use the labeled synchronized video history sequence with timestamps as the input and the target labeling result as the output to train the deep learning model to obtain the target detection model.

[0073] Specifically, a target detection method based on video frame synchronous reading provided by the present invention includes the steps:

[0074] S1: NTP calibration module: Set up a local area network NTP server as the time center of the local area network, configure an NTP client for each camera device in the local area network to achieve accurate NTP time synchronization. This step is implemented by the NTP calibration module.

[0075] S2: Configure an independent reading thread for each camera to directly read uncompressed low-latency raw video bitstream data and native millisecond-level timestamps from the camera. This step is implemented by the video data acquisition module.

[0076] S3: Start multiple FFmpeg encoding threads, accelerate through hardware encoding, perform H.264 high-efficiency low-latency encoding on each bitstream, output video frames with timestamps, and send them to the low-latency cache queue. This step is implemented by the video frame encoding module.

[0077] S4: Regularly retrieve video frames from the cache queues of each encoding thread in the main thread, perform precise time axis alignment according to the timestamp information in the frames, and output synchronized video frames with timestamps within the deviation range. This step is implemented by the video frame synchronization module.

[0078] S5: Form a sample batch of the video frames synchronized once, and sequentially input them into deep learning models such as the target detection and recognition models for parallel inference to obtain synchronized detection results. This step is implemented by the video frame synchronous inference module.

[0079] Step S1 specifically includes:

[0080] In the local area network, first, a computer with relatively strong performance needs to be selected as the Network Time Protocol (NTP) server. This server will act as the time source and provide accurate time information. Install and configure the NTP server software so that the server can provide time synchronization services to other devices.

[0081] Next, it is necessary to configure the camera devices that need to perform time synchronization. In the operating system of each camera device, configure the NTP client. By configuring the NTP client, the camera device can establish a connection with the NTP server and perform time synchronization and calibration regularly. In this way, the camera device can obtain accurate time information from the NTP server and update and calibrate its own system time.

[0082] Through NTP time synchronization, all camera devices within the local area network can maintain a time synchronization accuracy of milliseconds with the clock of the NTP server. This means that the system time of the camera device can be highly consistent with the clock of the NTP server and be synchronized with millisecond-level accuracy. This precision of time synchronization ensures that the camera device has consistent time stamps when recording and processing video data, providing an accurate time reference for subsequent data analysis and processing.

[0083] NTP (Network Time Protocol) is a protocol specifically used to achieve time synchronization in computer networks. Its goal is to coordinate the clocks of various device nodes in the network to ensure that they have the same time reference and provide high-precision time calibration, with an accuracy of up to millisecond level and above.

[0084] For example Figure 3 , set up a computer device that does not stop for 24 hours within the local area network as the NTP server, and at the same time connect the relevant cameras to the current local area network and configure the NTP client, which can ensure that accurate time references are provided for all camera devices connected to the NTP server.

[0085] Step S2 specifically includes:

[0086] To avoid the additional encoding and decoding delays introduced by the video compression algorithm and obtain uncompressed high-quality low-latency raw bitstream data and millisecond-level timestamps, for each camera that needs to read video frames, create a separate thread specifically.

[0087] Each thread is specifically responsible for reading the raw bitstream data from the corresponding camera and using the NTP service to obtain accurate millisecond-level timestamps. In this way, additional encoding and decoding delays are avoided, and high-quality and extremely low-latency video data is obtained for subsequent use.

[0088] A raw video stream refers to the original video data that has been collected or obtained without compression or encoding processing. The raw bitstream retains the complete information of the video signal, including the accurate values and color information of each pixel, as well as the timestamp information related to the video frame.

[0089] A timestamp is a numerical value used to precisely record the time when an event occurs. In the field of video processing, timestamps are used to mark the exact time points at which video frames are acquired. For each video frame, the timestamp represents the exact time when the frame is captured or acquired. Timestamps are typically expressed in specific time units (such as milliseconds) to represent the amount of time elapsed since a reference time point (such as when the video starts playing or the system boots up).

[0090] As Figure 4 , each camera that needs to acquire video frames is responsible for reading by a separate thread, enabling parallel processing and reading operations without interfering with each other. At the same time, by leveraging the accurate timestamps provided by the NTP service, it is ensured that the video data has high-quality millisecond-level timestamp markings, which helps with subsequent video processing and analysis.

[0091] Step S3 specifically includes:

[0092] To perform H.264 encoding processing on each acquired raw bitstream, an independent FFmpeg thread is used to encode each video stream. By leveraging applicable hardware encoding acceleration technologies, the efficiency and performance of the encoding process can be improved. Once a video frame is encoded, the encoded video frame and its corresponding timestamp are stored in a cache queue.

[0093] FFmpeg is a cross-platform open-source multimedia processing toolset that encompasses a wide range of audio and video processing tools and libraries. FFmpeg has the ability to handle multiple audio and video formats and provides rich functions and function libraries for performing various multimedia processing tasks such as decoding, encoding, transcoding, clipping, merging, and streaming transmission.

[0094] H.264 encoding is a general and widely used video compression standard and encoding format. It aims to achieve high-quality video transmission and storage by effectively compressing video data and reducing the bitrate requirements.

[0095] Hardware encoding is a process of using specialized hardware encoders for video compression and encoding. A hardware encoder is a dedicated hardware device with highly optimized encoding algorithms and circuit designs that can achieve fast and efficient video encoding.

[0096] The hardware encoder is an embedded hardware module, usually integrated into the CPU, GPU, or dedicated encoder chips, with highly optimized algorithms and parallel processing capabilities, capable of completing video encoding in a more efficient manner, aiming to provide fast and efficient video encoding functions.

[0097] As Figure 5, a dedicated FFmpeg thread is set for each video stream to handle H.264 encoding. By using independent threads, multiple video streams can be processed in parallel, thus improving the overall encoding efficiency.

[0098] During the encoding process, appropriate hardware encoding acceleration techniques can be used to improve performance. These techniques utilize dedicated hardware (such as GPU) to accelerate the encoding algorithm, thereby reducing the encoding time and resource consumption. By making full use of hardware resources, a more efficient encoding process can be achieved.

[0099] Once the video frame is encoded, the encoded video frame and its corresponding timestamp will be stored in the cache queue. This cache queue is used to store video frames to be processed for subsequent processing or transmission. Through the cache queue, the encoded video data can be effectively managed, and the correspondence between the timestamp and the video frame can be ensured not to be lost or confused.

[0100] Step S4 specifically includes:

[0101] In the main thread, the encoded multi-channel camera video frames and their corresponding timestamps are sequentially retrieved from the cache queue. Next, by comparing these timestamps with each other and calculating the absolute value of the time difference between them, the time deviation value is obtained. By setting an appropriate time deviation range, it can be determined whether the time deviation value is within the set deviation range, and based on this, it can be determined whether the current video frame is a synchronous video frame, that is, has a consistent time mark.

[0102] The time stamp deviation range is a time value in milliseconds, used to determine the synchronization of video frames. Specifically, through the time difference calculated from the timestamps of multiple video frames, if its value is less than or equal to the set time stamp deviation range, it can be determined that these video frames are synchronous.

[0103] Such as Figure 6 , in the main thread, in the order of video frames, the video frames and the corresponding timestamps are extracted one by one from the cache queue. The timestamp records the capture or encoding time of the video frame. By comparing these timestamps, the absolute value of the time difference between the video frames, that is, the time deviation, can be calculated. The time deviation value reflects the time relationship between the video frames.

[0104] To determine whether the video frames are synchronous, a time deviation range in milliseconds needs to be set. Within this range, if the time deviation value falls within the set range, it can be considered that the current video frame has a consistent time mark, that is, is synchronized with other video frames. This judgment of time synchronization can help ensure the consistency of the video frames of the multi-channel cameras during playback or processing, thereby determining whether it is a synchronous video frame.

[0105] Step S5 specifically includes:

[0106] After the operation of synchronizing video frames aligned on the time axis, video frames of the same batch will be combined into a sample batch. This sample batch can contain video frames from multiple cameras, which capture the same scene at the same moment.

[0107] This sample batch can be passed to deep learning models such as object detection for parallel inference. The deep learning model will process each video frame in the sample batch simultaneously to identify and detect the objects therein. Since these video frames are captured at the same moment and are processed with time axis alignment, the deep learning model can perform parallel inference on these video frames. In this way, the recognition results of objects can be obtained from multiple camera video frames at the same moment.

[0108] Such as Figure 7 , the synchronized video frames aligned on the time axis will form a sample batch and be sent to deep learning models such as object detection for parallel inference. By processing these video frames simultaneously, the model can obtain the recognition results of objects on multiple camera video frames at the same moment. Since these video frames are captured at the same moment, the model can effectively share information and understand the context among them, improving the efficiency and performance of the model, reducing the inference time, and enabling the model to perform more accurate object detection through time axis alignment.

[0109] The present invention implements time synchronization calibration services for each camera in the local area network by setting up an NTP server locally; Video data acquisition module: Obtain the original raw stream and millisecond-level timestamps by setting up independent threads for each camera; Video frame encoding module: Efficiently encode the raw stream data through video encoding tools; Video frame synchronization module: Synchronize the encoded video frames at the millisecond level; Video frame synchronous inference module: Combine the video frames synchronized once into a batch and send them to deep learning models for parallel inference detection. It can achieve more efficient and accurate video content understanding and detection recognition, and can effectively improve the accuracy and reliability of video-based visual inference tasks.

[0110] Embodiment 2

[0111] To execute the method corresponding to the above Embodiment 1 to achieve the corresponding functions and technical effects, the following provides an object detection system based on synchronous reading of video frames, including:

[0112] Synchronized timestamp-containing video sequence acquisition module, used to determine a synchronized timestamp-containing video sequence according to multiple low-latency buffer queues. The timestamp difference between any two timestamp-containing video frames in the synchronized timestamp-containing video sequence is less than the timestamp difference threshold.

[0113] A target detection module, which is configured to input a synchronized video sequence with timestamps into a target detection model to obtain a target detection result corresponding to the central timestamp. The target detection model is obtained by training a deep learning model using multiple annotated synchronized video historical sequences with timestamps.

[0114] Embodiment 3

[0115] This embodiment provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the target detection method based on synchronous video frame reading described in Embodiment 1. Among them, the memory is a readable storage medium. As Figure 8 , the terminal device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a real-time video synchronization reading program and a synchronous inference detection program. When the processor executes the computer program, it implements the steps in the above-mentioned embodiments of the real-time video synchronization reading and synchronous inference detection methods, such as Figure 1 the steps S1 to S5 shown. Or when the processor executes the computer program, it implements the functions of each module in the above-mentioned device embodiments.

[0116] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device. For example, the computer program can be divided into various modules, and the specific functions of each module will not be elaborated again.

[0117] The terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the schematic diagram is only an example of the terminal device, and does not constitute a limitation on the terminal device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the terminal device may further include input / output devices, network access devices, a bus, etc.

[0118] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The processor is the control center of the terminal device, and connects various parts of the entire terminal device through various interfaces and lines.

[0119] The memory can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and invoking the data stored in the memory, the processor realizes various functions of the terminal device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the computer (such as audio and video data, files, etc.). In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a flash memory card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0120] Among them, if the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present invention, it can also be completed by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0121] In this specification, the various embodiments are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0122] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A target detection method based on synchronous reading of video frames, characterized in that, The target detection method is applied to a monitoring system, which includes: an NTP server and multiple camera devices; the NTP server and the multiple camera devices are both set in the same local area network; any one of the camera devices is configured with an NTP client, a reading thread, and an FFmpeg encoding thread; The NTP client is used to generate a timestamp when the corresponding camera device generates a video frame; The NTP server is used to synchronize the timestamps generated by multiple NTP clients at the same moment; The reading thread is used to read the raw video stream data of multiple video frames generated by the corresponding camera device and the timestamp corresponding to each video frame; The FFmpeg encoding thread is used to encode the raw video stream data of each video frame and its corresponding timestamp in chronological order to obtain multiple timestamp-containing video frames, and form a low-latency cache queue for the corresponding camera device; The target detection method includes: Determine a synchronized timestamp-containing video sequence according to multiple low-latency cache queues; the timestamp difference between any two timestamp-containing video frames in the synchronized timestamp-containing video sequence is less than the timestamp difference threshold; Input the synchronized timestamp-containing video sequence into the target detection model to obtain the target detection result corresponding to the central timestamp; the target detection model is obtained by training a deep learning model using multiple labeled synchronized timestamp-containing video historical sequences.

2. The object detection method based on video frame synchronous reading according to claim 1, characterized in that, The determining of the synchronized timestamp-containing video sequence according to multiple low-latency cache queues includes: Extract one timestamp-containing video from each low-latency cache queue in turn according to the serial number order of the corresponding camera devices to obtain a pending synchronized timestamp-containing video sequence; Judge whether there is a timestamp difference between two timestamp-containing video frames in the pending synchronized timestamp-containing video sequence that is greater than the timestamp difference threshold, and obtain a first judgment result; If the first judgment result is yes, update multiple low-latency cache queues, and return to the step of "extracting one timestamp-containing video from each low-latency cache queue in turn according to the serial number order of the corresponding camera devices to obtain a pending synchronized timestamp-containing video sequence"; If the first judgment result is no, determine the pending synchronized timestamp-containing video sequence as the synchronized timestamp-containing video sequence.

3. The object detection method based on video frame synchronous reading according to claim 2, wherein, The updating of multiple low-latency cache queues includes: Determine the average timestamp corresponding to the pending synchronized timestamp-containing video sequence; Determine the two timestamp-containing video frames with a timestamp difference greater than the timestamp difference threshold as pending video frames; Determine the pending video frame with the largest difference between the corresponding timestamp and the average timestamp as the reference video frame; Judge whether the timestamp corresponding to the reference video frame is earlier than the average timestamp, and obtain a second judgment result; If the second judgment result is yes, delete the reference video frame from the low-latency cache queue corresponding to the reference video frame; If the second judgment result is no, determine the timestamp-containing video frames other than the reference video frame in the pending synchronized timestamp-containing video sequence as non-reference video frames; Delete the corresponding non-reference video frames from the low-latency cache queues corresponding to each non-reference video frame.

4. The object detection method based on video frame synchronous reading according to claim 1, wherein, Before determining the synchronized timestamp-containing video sequence according to the multiple low-latency cache queues, the following steps are further included: Obtain multiple low-latency cache history queues Determine a synchronized timestamp-containing video history sequence according to the multiple low-latency cache history queues; Label the targets in the synchronized timestamp-containing video history sequence respectively to obtain the labeled synchronized timestamp-containing video history sequence; Use the labeled synchronized timestamp-containing video history sequence as the input and the target labeling result as the output to train a deep learning model to obtain a target detection model.

5. A target detection system based on synchronous reading of video frames, characterized in that, It includes: A synchronized timestamp-containing video sequence acquisition module, configured to determine a synchronized timestamp-containing video sequence according to multiple low-latency cache queues; the time difference between any two timestamp-containing video frames in the synchronized timestamp-containing video sequence is less than the timestamp difference threshold; A target detection module, configured to input the synchronized timestamp-containing video sequence into the target detection model to obtain a target detection result corresponding to the central timestamp; The target detection model is obtained by training a deep learning model with multiple labeled synchronized timestamp-containing video history sequences.

6. An electronic device, characterized in that, It includes a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute a target detection method according to any one of claims 1 to 4, which is based on synchronous reading of video frames.

7. An electronic device according to claim 6, characterized in that, The memory is a readable storage medium.

Citation Information

Cited By

  • Multi-camera frame synchronization method and device and storage medium

    CN120856842A

  • Real-time multi-camera synchronization method and system based on multiple processes

    CN121037514A

  • Video reinjection method, system and equipment based on memory and readable storage medium

    CN121996444A