Video stream frame marking method, system and device based on ordered dictionary and OpenCV and medium
Through multi-threaded parallel processing and ordered dictionary storage technology, video stream frames are periodically sampled and inference are solved, and the problem of redundant inference in high-frame-rate video stream processing is significantly improved.
Patent Information
- Application Number
- CN202510193148.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
AI Technical Summary
In high-frame-rate video streaming scenarios, traditional frame-by-frame processing methods lead to excessive computing burden, slow system response speed, and unable to meet real-time requirements. How to reduce redundant inference in the video stream analysis process and improve system performance and response speed.
Multithreading technology is used to process the reading, inference, labeling and streaming process of video frames in parallel. By periodically sampling frames in the video stream, the video data of each frame is stored using an ordered dictionary structure. The reasoning service infers the main frame and applies the inference results to the secondary frame.
It effectively reduces redundant inference in the video stream analysis process, significantly improves the efficiency of video stream annotation, improves the system's real-time and computing optimization effect, and reduces the demand for computing resources.
Smart Images

Figure CN120075490A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of video streams, in particular to a video stream frame marking method, system, device and medium based on an ordered dictionary and OpenCV. Background Art
[0002] With the increasing demand for video stream processing, especially in high frame rate video stream scenarios, the traditional frame-by-frame processing method easily leads to excessive computational burden and slow system response speed, which cannot meet the real-time requirements. Especially in applications that require real-time analysis of large amounts of video stream data, how to reduce the pressure of inference calculations and improve efficiency has become an important technical issue. At present, many existing technologies use frame-by-frame inference methods to process video streams, but due to the high video frame rate, this method consumes a lot of computing resources and has a long delay time.
[0003] Therefore, how to reduce redundant reasoning in the video stream analysis process and improve system performance and response speed is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The technical task of the present invention is to provide a video stream frame marking method, system, device and medium based on an ordered dictionary and OpenCV to solve the problem of how to reduce redundant reasoning in the video stream analysis process and improve system performance and response speed.
[0005] The technical task of the present invention is achieved in the following way: a video stream marking method based on an ordered dictionary and OpenCV, which uses multi-threading technology to parallelly process the reading, reasoning, marking and streaming of video frames, that is, by periodically sampling frames in the video stream, using an ordered dictionary structure to store the video data of each frame, using an inference service to infer the main frame in the video data, and applying the inference result to the secondary frame in the video data, and the ordered queue of each frame of video data stored in the ordered dictionary is pushed as a video stream through FFmpeg.
[0006] Preferably, the method is as follows:
[0007] Video frame sampling: Use OpenCV's VideoCapture object to read video frame data from the video stream frame by frame, select some frames as main frames according to the set time interval or frame number interval, and the rest as sub-frames;
[0008] Ordered dictionary storage: Each frame of data is encapsulated and stored in an ordered dictionary through the Frame class;
[0009] Main frame reasoning: Whenever a main frame is extracted from a video stream, OpenCV combined with a deep learning reasoning model is used to perform target detection on the main frame and generate frame information;
[0010] Bounding box information synchronization: Apply the inference result of the main frame to the sub-frame corresponding to the main frame, that is, the sub-frame corresponding to the main frame directly reuses the bounding box information of the main frame;
[0011] Video streaming: Push the ordered queue of each frame of video data stored in the ordered dictionary as a video stream through FFmpeg.
[0012] Preferably, the key of each dictionary item in the ordered dictionary is the timestamp of the video frame, and the value is a Frame object containing the video frame data.
[0013] More preferably, the ordered dictionary stores the video frame data in chronological order.
[0014] A video stream bounding box system based on an ordered dictionary and OpenCV, the system includes:
[0015] Video frame sampling module, used to read video frame data frame by frame from the video stream using the VideoCapture object of OpenCV, and select some frames as main frames according to the set time interval or frame number interval, and the remaining frames as sub-frames;
[0016] Ordered dictionary storage module, used to encapsulate and store each frame of data in the ordered dictionary through the Frame class;
[0017] Main frame inference module, used to perform object detection on the main frame using OpenCV combined with a deep learning inference model and generate bounding box information whenever a main frame is extracted from the video stream;
[0018] Bounding box information synchronization module, used to apply the inference result of the main frame to the sub-frame corresponding to the main frame, that is, the sub-frame corresponding to the main frame directly reuses the bounding box information of the main frame;
[0019] Video streaming module, used to push the ordered queue of each frame of video data stored in the ordered dictionary as a video stream through FFmpeg.
[0020] Preferably, the system uses multi-threaded technology to parallelize the processes of reading, inferring, annotating, and streaming video frames.
[0021] Preferably, the key of each dictionary item in the ordered dictionary is the timestamp of the video frame, and the value is a Frame object containing the video frame data.
[0022] More preferably, the ordered dictionary stores the video frame data in chronological order.
[0023] An electronic device, including: a memory and at least one processor;
[0024] Wherein, a computer program is stored on the memory;
[0025] The at least one processor executes the computer program stored in the memory, such that the at least one processor executes the video stream bounding box method based on an ordered dictionary and OpenCV as described above.
[0026] A computer-readable storage medium stores a computer program, and the computer program can be executed by a processor to implement the video stream bounding box method based on an ordered dictionary and OpenCV as described above.
[0027] The video stream bounding box method, system, device, and medium of the present invention based on an ordered dictionary and OpenCV have the following advantages:
[0028] (1) By periodically sampling frames in the video stream and storing the video data of each frame using an ordered dictionary structure, the inference service infers the main frame and applies the inference result to the secondary frame, thereby effectively reducing the inference calculation burden of the secondary frame.
[0029] (2) The present invention can significantly improve the efficiency of video stream annotation, is widely applicable to fields such as video surveillance, intelligent transportation, and autonomous driving, has high real-time performance and calculation optimization effects. At the same time, the present invention can achieve rapid recognition and annotation of targets in the video stream, greatly improving the efficiency and accuracy of video analysis, and providing new technical means and solutions for related fields.
[0030] (3) The present invention relates to video stream processing technology, especially video stream bounding box technology, and combines the optimization of computer vision and inference service. Specifically, it is applicable to systems that need to process video stream data in real time, such as intelligent transportation, video surveillance, and autonomous driving, meeting the requirements of real-time performance and high efficiency. Usually, it is required to quickly and accurately process a large amount of video data for tasks such as real-time monitoring, target tracking, and anomaly detection. It can also reduce the system's demand for hardware resources and improve the scalability and stability of the system.
[0031] (4) The use of an ordered dictionary structure in the present invention can ensure that frame data is stored in chronological order, facilitating quick access during subsequent processing. The use of an ordered dictionary not only improves the efficiency of data retrieval but also ensures the orderliness of the data, providing strong support for the real-time processing of video streams. At the same time, the storage and processing methods of the ordered dictionary accelerate the access of frame data and improve the system processing efficiency.
[0032] (5) The image content difference between the secondary frame and the main frame of the present invention is small, so the secondary frame can directly reuse the bounding box information of the main frame, avoiding repeated inference calculations for the secondary frame, significantly improving the efficiency of video stream processing, and reducing unnecessary consumption of computing resources.
[0033] (6) During the inference process of the present invention, the bounding box information of the main frame is applied to the main frame and its corresponding sub-frame. The sub-frame does not need to be re-inferred and directly uses the inference result of the main frame for annotation. This process greatly speeds up the processing speed of the video stream and improves the real-time response ability of the system;
[0034] (7) The present invention reduces the inference frequency: by periodically sampling video frames (for example, sampling every certain time or several frames), the number of video frames that need to be processed per second is reduced, thereby reducing the computational burden; at the same time, periodic sampling and reuse of bounding box information enable the annotation process of the video stream to meet the real-time requirements;
[0035] (8) The sub-frame of the present invention reuses the bounding box information of the main frame, avoiding repeated inference of the sub-frame, greatly reducing the computational burden, further optimizing the inference calculation, improving the system performance, and being able to reduce the demand for computing resources while maintaining high efficiency;
[0036] (9) The present invention adopts a multi-threaded mechanism to process the operations of reading, inferring, and annotating video frames in parallel, ensuring that the system can maintain high-efficiency processing ability and real-time performance in a high-frame-rate video stream; the application of multi-threaded technology enables the system to make full use of the advantages of modern multi-core processors, significantly improving the throughput and response speed of video stream processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The present invention will be further described below with reference to the accompanying drawings.
[0038] Attached Figure 1 is a flowchart of a video stream bounding box method based on an ordered dictionary and OpenCV. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The video stream bounding box method, system, device, and medium of the present invention based on an ordered dictionary and OpenCV will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.
[0040] Embodiment 1:
[0041] As shown in the attached Figure 1 figure, this embodiment provides a video stream bounding box method based on an ordered dictionary and OpenCV. This method uses multi-threaded technology to process the operations of reading, inferring, annotating, and pushing video frames in parallel, that is, by periodically sampling the frames in the video stream, storing the video data of each frame using an ordered dictionary structure, inferring the main frame in the video data through an inference service, and applying the inference result to the sub-frame in the video data. The ordered queue of the video data of each frame stored in the ordered dictionary is pushed as a video stream through FFmpeg; specifically as follows:
[0042] S1. Video frame sampling: Use the VideoCapture object in OpenCV to read video frame data frame by frame from the video stream. Select some frames as main frames at a set time interval or frame interval, and the remaining frames as secondary frames;
[0043] S2. Ordered dictionary storage: Each frame of data is encapsulated and stored in an ordered dictionary through the Frame class. Among them, the key of each dictionary item in the ordered dictionary is the timestamp of the video frame, and the value is a Frame object containing the video frame data. Through the ordered dictionary structure, it can ensure that the frame data is stored in chronological order, facilitating quick access during subsequent processing, not only improving the efficiency of data retrieval, but also ensuring the orderliness of the data, providing strong support for the real-time processing of the video stream;
[0044] S3. Main frame inference: Whenever a main frame is extracted from the video stream, use OpenCV combined with a deep learning inference model to perform object detection on the main frame and generate bounding box information. Among them, deep learning inference models such as NCNN, OpenVino, TensorRT, MediaPipe, etc. are used;
[0045] S4. Bounding box information synchronization: The image content difference between the secondary frame and the main frame is small. Apply the main frame inference result to the secondary frame corresponding to the main frame, that is, the secondary frame corresponding to the main frame directly reuses the bounding box information of the main frame, avoiding repeated inference calculations for the secondary frame, significantly improving the efficiency of video stream processing, and reducing unnecessary consumption of computing resources;
[0046] S5. Video streaming: Push the ordered queue of each frame of video data stored in the ordered dictionary as a video stream through FFmpeg.
[0047] The inference calculation optimization in this embodiment is reflected in the following two aspects:
[0048] ① Reduce the inference frequency: By periodically sampling video frames (for example, sampling every certain time or several frames), the number of video frames that need to be processed per second is reduced, thereby reducing the computational burden.
[0049] ② Bounding box information reuse: The secondary frame reuses the bounding box information of the main frame, further optimizing the inference calculation and improving the system performance. In this way, the system can reduce the demand for computing resources while maintaining high efficiency.
[0050] Embodiment 2:
[0051] This embodiment provides a video stream bounding box system based on an ordered dictionary and OpenCV. This system uses multi-threaded technology to parallelize the processes of reading, inferring, annotating, and streaming video frames; This system includes:
[0052] Video frame sampling module, which is used to read video frame data frame by frame from a video stream using the VideoCapture object of OpenCV, and select some frames as main frames and the remaining frames as secondary frames according to the set time interval or frame number interval;
[0053] Ordered dictionary storage module, which is used to encapsulate and store each frame of data in an ordered dictionary through the Frame class; among them, the key of each dictionary item in the ordered dictionary is the timestamp of the video frame, and the value is the Frame object containing the video frame data; the ordered dictionary stores the video frame data in chronological order;
[0054] Main frame inference module, which is used to perform object detection on the main frame using OpenCV combined with a deep learning inference model and generate bounding box information whenever a main frame is extracted from the video stream;
[0055] Bounding box information synchronization module, which is used to apply the main frame inference result to the secondary frame corresponding to the main frame, that is, the secondary frame corresponding to the main frame directly reuses the bounding box information of the main frame;
[0056] Video streaming module, which is used to push the ordered queue of each frame of video data stored in the ordered dictionary as a video stream through FFmpeg.
[0057] Embodiment 3:
[0058] This embodiment also provides an electronic device, including: a memory and a processor;
[0059] Among them, the memory stores computer execution instructions;
[0060] The processor executes the computer execution instructions stored in the memory, so that the processor executes the video stream bounding box method based on an ordered dictionary and OpenCV in any embodiment of the present invention.
[0061] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0062] The memory can be used to store computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory and invoking the data stored in the memory, the processor realizes various functions of the electronic device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash memory card, at least one magnetic disk storage period, flash memory device, or other volatile solid-state storage devices.
[0063] Embodiment 4:
[0064] This embodiment also provides a computer-readable storage medium, which stores multiple instructions. The instructions are loaded by the processor to cause the processor to execute the video stream bounding box method based on an ordered dictionary and OpenCV in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided. The software program code for realizing the functions in any one of the above embodiments is stored on the storage medium, and the computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.
[0065] In this case, the program code read from the storage medium itself can realize the functions in any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.
[0066] Examples of the storage medium for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Optionally, the program code can be downloaded from a server computer via a communication network.
[0067] In addition, it should be clear that not only can the functions in any one of the above embodiments be realized by executing the program code read by the computer, but also by causing an operating system operating on the computer based on the instructions of the program code to complete part or all of the actual operations.
[0068] In addition, it can be understood that the program code read out from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU etc. installed on the expansion board or the expansion unit are made to execute part or all of the actual operations, thereby implementing the functions of any one of the above embodiments.
[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A video stream frame marking method based on ordered dictionary and OpenCV, characterized in that: The method uses multi-threading technology to parallelly process the reading, reasoning, annotation and streaming of video frames, that is, by periodically sampling frames in the video stream, using an ordered dictionary structure to store the video data of each frame, using an inference service to infer the main frame in the video data, and applying the inference result to the secondary frame in the video data. The ordered queue of each frame of video data stored in the ordered dictionary is pushed as a video stream through FFmpeg.
2. The video stream marking method based on ordered dictionary and OpenCV according to claim 1, characterized in that: The method is as follows: Video frame sampling: Use OpenCV's VideoCapture object to read video frame data from the video stream frame by frame, select some frames as main frames according to the set time interval or frame number interval, and the rest as sub-frames; Ordered dictionary storage: Each frame of data is encapsulated and stored in an ordered dictionary through the Frame class; Main frame reasoning: Whenever a main frame is extracted from a video stream, OpenCV combined with a deep learning reasoning model is used to perform target detection on the main frame and generate frame information; Frame information synchronization: The main frame inference result is applied to the sub-frame corresponding to the main frame, that is, the sub-frame corresponding to the main frame directly reuses the frame information of the main frame; Video streaming: The ordered queue of each frame of video data stored in the ordered dictionary is pushed as a video stream through FFmpeg.
3. The video stream marking method based on ordered dictionary and OpenCV according to claim 1, characterized in that: The key of each dictionary item in the ordered dictionary is the timestamp of the video frame, and the value is the Frame object containing the video frame data.
4. The video stream frame marking method based on ordered dictionary and OpenCV according to any one of claims 1 to 3, characterized in that: The ordered dictionary stores video frame data in chronological order.
5. A video stream frame marking system based on ordered dictionary and OpenCV, characterized in that: The system includes: The video frame sampling module is used to use the VideoCapture object of OpenCV to read the video frame data from the video stream frame by frame, and select some frames as main frames and the rest as sub-frames according to the set time interval or frame number interval; The ordered dictionary storage module is used to encapsulate each frame of data and store it in an ordered dictionary through the Frame class; The main frame reasoning module is used to detect the target of the main frame and generate the frame information by using OpenCV combined with the deep learning reasoning model whenever the main frame is extracted from the video stream; The frame information synchronization module is used to apply the main frame inference result to the sub-frame corresponding to the main frame, that is, the sub-frame corresponding to the main frame directly reuses the frame information of the main frame; The video streaming module is used to push the ordered queue of each frame of video data stored in the ordered dictionary as a video stream through FFmpeg.
6. The video stream frame marking system based on ordered dictionary and OpenCV according to claim 5, characterized in that: The system uses multi-threading technology to parallelize the reading, reasoning, annotation and streaming of video frames.
7. The video stream frame marking system based on ordered dictionary and OpenCV according to claim 5, characterized in that: The key of each dictionary item in the ordered dictionary is the timestamp of the video frame, and the value is the Frame object containing the video frame data.
8. The video stream frame marking system based on ordered dictionary and OpenCV according to any one of claims 5 to 7, characterized in that: The ordered dictionary stores video frame data in chronological order.
9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores a computer program; The at least one processor executes the computer program stored in the memory, so that the at least one processor performs the video stream frame marking method based on an ordered dictionary and OpenCV according to any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which can be executed by a processor to implement the video stream frame marking method based on an ordered dictionary and OpenCV as claimed in any one of claims 1 to 4.