An image conversion and streaming method and device for enhancing real-time performance

Through multi-threaded parallel processing and fixed-point computing optimization methods, the problems of insufficient real-time performance and low resource utilization of edge computing devices such as drones in high-pixel image processing scenarios are solved, and efficient video streaming and low-latency streaming are achieved.

CN119893191BActive Publication Date: 2025-06-03ROCK AI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510368789.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-03
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

In the high-pixel image processing scenario, the existing technology lacks real-time performance and low resource utilization rate, resulting in high video streaming delays and poor user experience for edge computing devices such as drones.

Method used

The multi-threaded parallel processing architecture is adopted to process image frames in parallel through thread pool model, and the YUV conversion process is optimized through fixed-point calculation, combining shared memory mechanism and dynamic load balancing to improve processing efficiency and resource utilization.

Benefits of technology

It significantly improves image processing speed, reduces computing load, ensures video stream consistency and low latency transmission, and is suitable for high-resolution image processing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119893191B_ABST
    Figure CN119893191B_ABST
Patent Text Reader

Abstract

The present application discloses a method and device for enhancing real-time image conversion and streaming. Adopting a multi-threaded architecture, the producer thread collects images and stores them in a double-ended queue. After the conveyor thread extracts image frames in batches into a temporary container, the wake-up thread pool is awakened for parallel processing. In the image conversion optimization, integer fixed-point calculation is designed to replace traditional floating-point operations. The coefficients of the YUV conversion formula are scaled by 256 times and then rounded to reduce the calculation time consumption. Through the coexistence internal sharing mechanism, the YUV data pointer is directly mapped to the AVFrame structure of FFmpeg to eliminate the memory deep copy overhead. Based on the error accumulator, the capacity of the temporary container is dynamically adjusted to match the production and consumption rates in real time, avoiding queue overflow or thread starvation. The present application realizes real-time and efficient streaming through a technical solution that combines multi-threaded parallel processing and image conversion optimization, and is particularly applicable to resource-constrained scenarios such as unmanned aerial vehicles and embedded devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing and video stream transmission, and particularly relates to an image conversion and streaming method and device for enhancing real-time performance, which is particularly suitable for low-latency video stream transmission in high-pixel image processing scenarios of edge computing devices such as unmanned aerial vehicles (UAVs). Background Art

[0002] With the deep application of UAVs in fields such as agricultural inspection, disaster monitoring, and logistics transportation, the technologies for real-time video stream acquisition, conversion, and streaming have become key supports. However, when dealing with high-resolution images, existing solutions generally face technical bottlenecks such as insufficient real-time performance and low resource utilization. The current mainstream implementation methods mainly rely on the following two technical solutions, but both have significant limitations:

[0003] 1. Solution 1: Transmission after local storage ( Figure 1 as shown):

[0004] Taking the DJI UAV M30 as an example, after the UAV video stream is captured in the Mat format of OpenCV to form a Mat video frame, it is stored as a local MP4 video file frame by frame using the VideoWrite tool, and then uploaded to a remote server through the HTTP / FTP protocol for the client to download and play.

[0005] This solution has a simple and convenient interface and does not require too many real-time encoding and decoding operations. Since it will be saved in advance, there will be no loss of data due to inconsistent streaming and decoding speeds. During the video playback process, the smoothness of local playback mainly depends on whether the video file has been completely transmitted to the local. Specifically, only when all the data of the video file is completely transmitted to the local device will the playback start. This means that during the playback process, the network condition will not affect the local playback because the content being played has been stored locally and does not rely on real-time network transmission. However, this solution has the following core defects:

[0006] a) Performance collapse caused by I / O blocking: During the video transmission process, the transmission of video frames through the network itself is an I / O-intensive operation, and when these video frames arrive locally, they need to be further written to the disk, which involves another I / O-intensive operation. Since these operations all consume a large amount of time, it is difficult to ensure real-time performance throughout the process. Saving the video requires writing to the disk frame by frame, causing the program to be frequently blocked in I / O operations, and the CPU cannot efficiently execute subsequent tasks. The overall processing rate is limited by the dual performance bottlenecks of network transmission and disk throughput.

[0007] b) Uncontrollable end-to-end latency: The video needs to be completely stored before it can be transmitted, resulting in a link latency of up to several seconds from acquisition to playback, which cannot meet the requirements of real-time monitoring (such as the identification of emergencies in UAV patrol).

[0008] 2. Solution 2: Real-time streaming based on the producer-consumer model ( Figure 2 , Figure 3 as shown)

[0009] To improve real-time performance, the existing technology uses the FFmpeg toolchain combined with the RTMP protocol. After obtaining the Mat frame, it is converted to the YUV420 format, the AVFrame structure is filled and encoded, and then pushed to the streaming media server through the RTMP protocol (such as SRS / Nginx-RTMP-Module / Nginx-HTTP-FLV-Module). The client pulls the stream from the streaming media server through protocols such as HTTP-FLV, WebRTC, or HLS and displays it (such as Figure 2 as shown). This solution optimizes the processing flow through the producer-consumer model (coordination of two threads):

[0010] Thread A (producer) maintains a double-ended queue (Double-Ended Queue) of size N, inserts data from the tail of the queue, and data exceeding N will overflow from the head of the queue. Thread B (consumer) mutually extracts the head data from the queue, performs the conversion from Mat to YUV420 format, then fills the data of the AVFrame structure, and finally pushes the stream after encoding. This process repeats, and the processing logic is as Figure 3 shown, taking N = 5 as an example.

[0011] In an ideal state, the speed of data filling into the queue and the speed of data extraction should be basically matched or the former is slightly greater than the latter (if the former is less than the latter, the consumer starvation phenomenon will occur, that is, the video frames have been consumed, and the consumer cannot obtain the latest video frames, resulting in stuttering and incoherence of the pushed video), and the length of the queue should be maintained within a small range to ensure real-time performance and the coherence of the pushed video.

[0012] In this design, when the pixel amount of a single Mat data is small, the speed of data extraction and processing will be greater than the speed of data filling into the queue. At this time, an appropriate delay can be added to the consumer thread B to ensure that the data processing speed and the data filling speed into the queue are matched, so as to avoid the situation that there is no new data coming in thread B and the pushed video is briefly interrupted due to waiting in vain.

[0013] However, this solution has the following core defects:

[0014] a) Queue overflow and frame skipping in large-pixel scenarios: Such as Figure 3As shown, when the pixels of the image to be transmitted are large and the image quality is not desired to be sacrificed, the speed at which thread B retrieves and processes data from the queue will be much slower than the speed at which thread A fills it. Thread A still continuously fills in data to ensure the real-time nature of the video frames, resulting in many data in the queue overflowing actively before they can be processed, such as data 1. At this time, during the video push process, due to the excessive interval between two frames (the overflowed data is not processed, and the subsequent processed data is several frames apart), it causes stuttering and frame skipping phenomena when the pushed video is played, reducing the user's viewing experience.

[0015] b) Waste of resources in single-threaded serial processing: As Figure 3 shown in the process, the consumer thread (thread B) only relies on a single core to perform serial operations and cannot utilize the parallel computing power of a multi-core CPU (for example, an 8-core CPU only occupies 1 core), resulting in idle hardware resources and efficiency bottlenecks.

[0016] c) Contradiction between high-precision conversion and computing power of embedded devices: The existing YUV conversion interface of OpenCV uses floating-point operations (such as coefficients 0.299, 0.587, etc.). Although it ensures the image quality, its computational complexity causes the single-frame conversion time-consuming to account for more than 60% (such as 50 ms / frame), exacerbating the processing delay and energy consumption pressure in edge devices such as drones.

[0017] d) Redundant memory copying exacerbates latency: As Figure 2 in the step of "filling the AVFrame structure" in, the existing solution copies the YUV420 data to the AVFrame structure through deep copying (memcpy), resulting in an increase in the single-frame memory operation time-consuming (such as from 1 ms to 5 ms), further weakening the real-time nature.

[0018] In summary, as Figures 1 to 3 shown, the existing solutions have significant defects in terms of real-time nature, resource utilization rate, and algorithm efficiency. Especially in high-resolution image processing scenarios, queue overflow, single-thread bottlenecks, and memory redundancy problems severely restrict the real-time transmission ability of drone video streams. How to systematically solve the above technical pain points through multi-threaded scheduling optimization, lightweight algorithm design, and memory management innovation, and provide an efficient solution for the real-time video stream processing of edge devices, is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention

[0019] The present invention aims to solve the problems of insufficient real-time nature and low resource utilization rate in high-pixel image pushing in the prior art. Through multi-threaded parallel processing and image conversion optimization, it significantly improves the processing speed and reduces the computational load, ensuring the coherence of the video stream, and is applicable to embedded and edge computing scenarios.

[0020] To achieve the above object, the present invention adopts the following technical solutions:

[0021] In a first aspect, the present application discloses a method for enhancing real-time image conversion and streaming, including the following steps:

[0022] Obtain the original image data from the image source through the producer thread, and store the original image data into the double-ended queue in sequence;

[0023] Extract multiple image frames from the double-ended queue in sequence through the delivery thread model, assign a unique serial number to each image frame according to the order in the queue, and then store them into the temporary container;

[0024] Wake up multiple processing threads in the thread pool, and perform the following parallel processing on the image frames in the temporary container:

[0025] a) Convert the image frame from Mat format to YUV420 format, and optimize the conversion process using fixed-point calculation;

[0026] b) Directly map the converted YUV420 data to the data field of the AVFrame structure of FFmpeg through the shared memory mechanism;

[0027] Store the processed image frames into the global sorting container according to the serial numbers, where the global sorting container is automatically sorted in ascending order according to the serial numbers, control the atomicity of the multi-threaded insertion operation through the global mutex lock, and at the same time count the number of completed frames through the atomic counter;

[0028] When the count of the atomic counter reaches the preset threshold, the streaming thread obtains the data from the global sorting container in the order of the serial numbers and streams it to the remote server;

[0029] Dynamically adjust the capacity of the temporary container and the scale of the thread pool according to the difference between the current processing speed and the target speed recorded by the error accumulator, so as to match the data extraction speed with the processing speed.

[0030] In a preferred embodiment, the optimization process of the fixed-point calculation includes:

[0031] Multiply the floating-point coefficient in the YUV conversion formula by 256 and then take the integer, and divide the result by 256 to restore the precision after the conversion. The specific conversion formula is:

[0032] Y = (77R + 150G + 29B) / 256

[0033] U = (-38R - 74G + 112B) / 256

[0034] V = (157R - 132G - 26B) / 256

[0035] Among them, R, G, and B are the three components of a Mat-format image, and Y, U, and V are the three components of the YUV420 format.

[0036] In a preferred embodiment, the wake-up strategy of the thread pool is as follows: when the number of image frames in the temporary container reaches a preset batch threshold, wake up a number of threads equal to the batch quantity for parallel processing.

[0037] In a preferred embodiment, the process of dynamically adjusting the capacity of the temporary container and the scale of the thread pool includes: judging according to the difference between the current processing speed and the target speed recorded by the error accumulator:

[0038] When the difference exceeds the set threshold range, trigger an adjustment of the number of image frames extracted from the temporary container in the next batch, and synchronously adjust the scale of the thread pool to make the number of threads match the capacity of the temporary container;

[0039] When the difference is within the set threshold range, keep the current capacity of the temporary container and the scale of the thread pool.

[0040] In a preferred embodiment, the steps of converting an image frame from the Mat format to the YUV420 format further include:

[0041] Call the GPU parallel accelerator for fixed-point calculation, and implement pixel-level parallel processing through CUDA kernel functions;

[0042] Dynamically enable or disable the GPU acceleration function according to the system power or computing power load.

[0043] In a second aspect, the present application discloses a device for enhancing real-time image conversion and streaming, including:

[0044] A data acquisition module configured to obtain raw image data from an image source through a producer thread and store the raw image data in a double-ended queue in sequence;

[0045] A batch extraction module configured to sequentially extract multiple image frames from the double-ended queue through a delivery thread model, assign a unique serial number to each image frame according to the order in the queue, and store them in a temporary container;

[0046] A parallel processing module configured to wake up multiple processing threads in a thread pool and perform parallel processing on the image frames in the temporary container, including:

[0047] A format conversion unit for converting an image frame from the Mat format to the YUV420 format and optimizing the conversion process by fixed-point calculation;

[0048] A memory mapping unit for directly mapping the converted YUV420 data to the data field of the AVFrame structure of FFmpeg through a shared memory mechanism;

[0049] A data sorting module configured to store the processed image frames into a global sorting container in sequence number, where the global sorting container is automatically sorted in ascending order according to the sequence number, and the atomicity of multi-threaded insertion operations is controlled by a global mutex lock, and at the same time, the number of completed frames is counted by an atomic counter;

[0050] A dynamic streaming module configured to, when the atomic counter reaches a preset threshold, obtain data from the global sorting container in sequence order and stream it to a remote server, and dynamically adjust the capacity of the temporary container and the scale of the thread pool according to the difference between the current processing speed and the target speed recorded by the error accumulator.

[0051] In a preferred embodiment, the format conversion unit includes a fixed-point calculation module configured to:

[0052] Multiply the floating-point coefficients in the YUV conversion formula by 256 and then take the integer part, and divide the result by 256 to restore the precision after the conversion. The specific conversion formula is:

[0053] Y = (77R + 150G + 29B) / 256

[0054] U = (-38R - 74G + 112B) / 256

[0055] V = (157R - 132G - 26B) / 256

[0056] where R, G, and B are the three components of the Mat format image, and Y, U, and V are the three components of the YUV420 format.

[0057] In a preferred embodiment, the thread pool wake-up strategy of the parallel processing module is configured to:

[0058] When the number of image frames in the temporary container reaches a preset batch threshold, wake up the same number of threads as the batch quantity for parallel processing.

[0059] In a preferred embodiment, the dynamic streaming module includes an error accumulator configured to:

[0060] Record the difference between the current processing speed and the target speed, and trigger an adjustment according to whether the difference exceeds a set threshold range. When the difference exceeds the set threshold range, adjust the number of image frames extracted from the temporary container in the next batch in proportion, and synchronously adjust the scale of the thread pool.

[0061] In a preferred embodiment, the format conversion unit further includes a GPU acceleration module, configured to:

[0062] Call the GPU parallel accelerator for fixed-point calculation and implement pixel-level parallel processing through CUDA kernel functions;

[0063] Dynamically enable or disable the GPU acceleration function according to the system power or computing power load.

[0064] In a preferred embodiment, the global sorting container is a thread-safe ordered map, configured to:

[0065] Store data in a key-value pair structure, where the key is the sequence number of the image frame and the value is the corresponding processing result;

[0066] Arrange in ascending order automatically by sequence number;

[0067] Control the atomicity of multi-threaded insertion operations through mutex locks.

[0068] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0069] The present invention provides a method and device for enhancing real-time image conversion and streaming, which solves the problems of insufficient real-time performance and low resource utilization rate in the processing of high-resolution video streams such as drones. The technical means include:

[0070] Multi-threaded parallel processing architecture: Adopt a thread pool model, expand the consumer threads into multi-threaded parallel tasks, and realize the full utilization of multi-core CPU resources by dynamically allocating image frames to multiple threads, significantly improving the processing throughput.

[0071] Lightweight fixed-point calculation optimization: Replace the traditional floating-point operation YUV conversion formula (such as Y = 0.299R + 0.587G + 0.114B) with integer fixed-point calculation (such as Y = (77R + 150G + 29B) / 256), and reduce the conversion time per frame while ensuring visual quality.

[0072] Shared memory mechanism: Eliminate deep memory copies by directly passing the YUV data pointer to the AVFrame structure, reducing memory bandwidth occupancy and processing latency.

[0073] Dynamic load balancing: Based on the feedback mechanism of the error accumulator, dynamically adjust the capacity of the temporary container, balance the production and consumption rates, and avoid queue overflow and thread starvation. Description of the Drawings

[0074] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0075] Figure 1 It is a schematic flowchart of the prior art for locally storing MP4 video files and transmitting them using the HTTP protocol.

[0076] Figure 2 It is a schematic flowchart of the prior art for pushing a stream to a remote server using the FFmpeg pushing tool via the RTMP protocol and pulling the stream for display by the client via HTTP-FLV / WebRTC / HLS.

[0077] Figure 3 It is a schematic diagram of the dual-end queue data processing logic based on the producer and consumer design concept in the prior art.

[0078] Figure 4 It is a schematic diagram of the system architecture for multi-threaded parallel processing of image frames and pushing a stream proposed by the present invention.

[0079] Figure 5 It is a schematic diagram of the principle for directly mapping the converted YUV420 data to the data fields of the AVFrame structure through the shared memory mechanism proposed by the present invention. Specific Embodiments

[0080] To make the above and other features and advantages of the present invention clearer, the present invention will be further described below with reference to the drawings. It should be understood that the specific embodiments given herein are for the purpose of explaining to those skilled in the art and are merely exemplary, not restrictive.

[0081] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0082] Aiming at the defects existing in the prior art, the present invention optimizes the real-time performance of pushing a large image stream from the following two dimensions:

[0083] In the first dimension, to address the problem of excessive time consumption in single-threaded serial processing, the dynamic thread pool technology is adopted to parallelize the processing process. By pre-caching a certain number of image frames (such as 5 frames), when the cache reaches the threshold, a multi-thread parallel processing mechanism is triggered, compressing the processing time of multiple frames of images to be close to the single-frame processing time, and fully utilizing the multi-core resources of the CPU. Although multi-thread synchronization will introduce additional overhead, it can be ignored compared with serial frame-by-frame processing.

[0084] In the second dimension, to address the problem of excessive latency in single-frame processing, the processes such as memory copying and complex calculations during the processing can be optimized. For example, the computational load can be reduced and the real-time requirements can be improved through the shared memory mechanism and fixed-point calculation optimization.

[0085] In view of the above considerations, the present invention proposes an image conversion and streaming method and device for enhancing real-time performance, which solves the above problems through the following technical solutions:

[0086] Multi-thread parallel processing: Design a dynamic thread pool, allocate image frames to multiple threads for parallel processing, and ensure the order of processing results through a global sorting container.

[0087] Fixed-point calculation optimization: Convert floating-point coefficients to integers for calculation, and combine GPU acceleration and shared memory mechanism to significantly improve the conversion efficiency.

[0088] Dynamic load balancing: Based on the error accumulator and adaptive adjustment strategy, match the processing speed with the data generation speed in real time, and optimize the resource utilization rate by dynamically adjusting the scale of the thread pool.

[0089] Embodiment 1:

[0090] This embodiment provides a method for image conversion and streaming with enhanced real-time performance, which solves the problems of high latency and large resource consumption in high-resolution image streaming through multi-thread parallel processing and image conversion optimization.

[0091] Among them, the multi-thread parallel processing can be solved by using a multi-thread model. The main idea is to design a thread pool, parallelize the processing of image frames, improve the speed of fetching data, and thus increase the continuity and real-time performance of the images. The system mainly includes a thread pool, a delivery thread model, a streaming thread model, a data delivery container a, a global sorting container map, an atomic counter (to ensure the mutual exclusion of shared variables), and an error accumulator. The principle is as Figure 4 shown.

[0092] The specific implementation steps are as follows:

[0093] Step S1: The producer thread obtains the original image data from the image source and stores the original image data into the double-ended queue in sequence.

[0094] The producer thread continuously obtains the raw image data from an image source (such as a drone camera), in the Mat format of OpenCV.

[0095] The image data is stored in a double-ended queue (Double-Ended Queue) in the acquisition order. The double-ended queue is designed to be thread-safe and supports concurrent operations of the producer and consumer threads.

[0096] Step S2: Extract multiple image frames from the double-ended queue in sequence through the transfer thread model, and then assign a unique serial number according to the order of the image frames in the queue and store them in a temporary container (container a). At this time, there are multiple image frames with serial numbers in container a.

[0097] The transfer thread model extracts multiple image frames (for example, 5 frames each time) from the head of the double-ended queue in sequence, and assigns a globally unique increasing serial number to each frame (such as incrementing frame by frame starting from 0).

[0098] The extracted image frames are stored in container a, and container a supports dynamic capacity adjustment (the initial capacity is N, such as N = 5).

[0099] Step S3: When container a reaches the threshold capacity, wake up multiple processing threads in the thread pool, and distribute the image frames in container a to different threads for processing, so as to process the image frames in parallel. Specifically, it includes the following processing steps:

[0100] 1. Thread pool management and task allocation

[0101] When the number of image frames in container a reaches the preset threshold (such as 5 frames), wake up the same number of processing threads in the thread pool as the number of image frames (for example, 5 frames correspond to 5 threads).

[0102] The scale of the thread pool is dynamically adjusted according to the container capacity: if the container expands, the number of threads is increased; if the container shrinks, the number of threads is decreased.

[0103] 2. Processing tasks

[0104] Each thread allocates one frame of data from container a for processing. The processing task includes the following two sub-steps:

[0105] a) Convert the image frame from the Mat format to the YUV420 format, and optimize the conversion process using fixed-point calculation.

[0106] In the existing OpenCV library, the provided interfaces can achieve a relatively high-precision conversion. However, the power consumption will increase linearly with the excessive precision, and the speed of single calculation will also decrease. Video streams are generally captured by embedded devices and processed in real time on edge devices. Most of these devices have limited computing power and have high requirements for battery life. Especially in the scenario of drones, high-performance computing devices cannot be deployed, and the battery resources carried are also very limited. In addition, the real-time performance after the drone pushes the stream is the most important, and the requirement for the image quality is not particularly high.

[0107] In the traditional method, the conversion principle of converting an image frame from Mat format to YUV420 format is as follows. Here, R, G, and B are the three components of the Mat format image, and Y, U, and V are the three components of the YUV420 format.

[0108] Y = 0.299 * R + 0.587 * G + 0.114 * B

[0109] U = -0.147 * R - 0.289 * G + 0.436 * B

[0110] V = 0.615 * R - 0.515 * G - 0.100 * B

[0111] The existing API interfaces in OpenCV internally complete the conversion using the above formulas, achieving a high-precision conversion. And this API does not reserve optional precision parameters, so it will consume more time.

[0112] In view of the above situation, this embodiment optimizes the calculation speed by using fixed-point calculation. Fixed-point calculation is a way to replace floating-point calculation. It can be achieved by magnifying the fractional part into an integer, which can meet the requirements of faster calculation speed and lower power consumption. For example, to represent a number with a precision of 4 decimal places, the floating-point number can be multiplied by a constant (such as 10000) to convert it into an integer. After calculation, the result is divided by this constant (such as 10000) to restore the original precision.

[0113] For the above formulas, in order to avoid losing precision, this embodiment multiplies the floating-point numbers by 256 and then rounds the results to integers. The coefficients are respectively:

[0114] Ycoeff 1 = 0.299 * 256 = 76.544 ≈ 77

[0115] Ycoeff 2 = 0.587 * 256 = 150.272 ≈ 150

[0116] Ycoeff 3 = 0.114 * 256 = 29.184 ≈ 29

[0117] The formula for converting Mat format to YUV420P format is reconstructed as:

[0118] Y = (77R + 150G + 29B) / 256

[0119] U = (-38R - 74G + 112B) / 256

[0120] V = (157R - 132G - 26B) / 256

[0121] Using fixed-point number calculation can significantly improve the calculation efficiency, especially in embedded systems or scenarios that require efficient image processing. Although there is a loss of precision, in many scenarios, this loss is acceptable, especially when the conversion precision is not required to be very high.

[0122] In a preferred embodiment, a GPU parallel accelerator can also be configured to call the underlying GPU resources for accelerated calculation. When it is detected that the processing delay is greater than the threshold (i.e., when real-time requirements need to be improved), GPU acceleration is automatically started. After the real-time performance meets the standard, the accelerated calculation is turned off and switched back to the CPU mode to reduce power consumption.

[0123] b) Directly map the YUV420 data to the data field of the AVFrame structure of FFmpeg through the shared memory mechanism to avoid deep copy (memcpy) operations.

[0124] The previous approach was to use memcpy for deep copying of memory to avoid pointer management and complex variable lifecycle issues, but it would introduce a large memory overhead. However, the data is read-only and does not need to be repeatedly created and destroyed.

[0125] In this embodiment, the data address of YUV420P is directly passed to the elements of data of AVFrame through the shared memory mechanism, thereby realizing the reuse of Mat frame data. The principle after improvement is as Figure 5 shown. By improving this method, the push stream speed can be greatly increased to ensure real-time performance.

[0126] 3. Storage of processing results

[0127] After the processing is completed, the result and the serial number form a key-value pair, which is inserted into the global sorted container (map) through the global mutex, so as to be automatically sorted in ascending order according to the serial number, ensuring that the push stream is carried out in the order of image generation.

[0128] Step S4: Store the processed image frames into the global sorting container according to the sequence numbers, and count the number of completed frames through an atomic counter.

[0129] The global sorting container stores all processed image frames. The key is the sequence number, and the value is the processing result (AVFrame pointer).

[0130] The atomic counter counts the number of completed frames, and the count of the counter increases by one after each frame is inserted.

[0131] Step S5: When the count of the atomic counter reaches the preset threshold, the streaming thread retrieves data from the global sorting container in the order of sequence numbers and streams it to the remote server.

[0132] The streaming thread continuously monitors the value of the atomic counter. When the count reaches the preset threshold (such as 5 frames), perform the following operations:

[0133] Retrieve AVFrame data (such as sequence numbers 0 to 4) from the global sorting container (map) in ascending order of sequence numbers.

[0134] Push the data to the remote server through the streaming interface of FFmpeg (such as av_interleaved_write_frame), supporting protocols such as RTMP and HTTP-FLV.

[0135] After the streaming is completed, clear the global sorting container (map) and reset the atomic counter to 0.

[0136] Step S6: The error accumulator dynamically adjusts the capacity of container a and the scale of the thread pool according to the processing speed of this time, so as to ensure that the storage and retrieval speeds from the queue match.

[0137] The error accumulator records in real time the difference between the processing speed (such as the number of frames processed per second) and the target speed (Δ = target speed - actual speed).

[0138] If the Δ value exceeds the set threshold range (such as |Δ|>2 frames / second), trigger:

[0139] 1. Adjust the capacity of the current container a:

[0140] For example, adjust the capacity of the current container a proportionally (such as expanding by Δ×k, where k is an empirical coefficient).

[0141] If the current capacity is 5 frames, Δ = 4 frames / second, and k = 0.5, then the new capacity is: 5 + 0.5×4 = 7 frames.

[0142] 2. Synchronously adjust the scale of the thread pool: Dynamically increase or decrease the number of threads according to the new capacity to ensure that all image frames in the container can be processed in parallel.

[0143] For example, if the new capacity is 10 frames, the thread pool activates 10 threads to ensure that all frames can be processed in parallel.

[0144] If the Δ value is within the set threshold range (e.g., |Δ| ≤ 2 frames / second), then:

[0145] Maintain the current container capacity and thread pool size to avoid the system overhead caused by frequent adjustments.

[0146] Embodiment 2:

[0147] Based on the same design concept, this embodiment also provides an image conversion and streaming pushing device for enhancing real-time performance, which specifically includes the following modules:

[0148] A data acquisition module, configured to obtain raw image data from an image source through a producer thread and store the raw image data into a double-ended queue in sequence;

[0149] A batch extraction module, configured to sequentially extract multiple image frames from the double-ended queue through a delivery thread model, assign a unique serial number to each image frame according to the order in the queue, and then store them into a temporary container;

[0150] A parallel processing module, configured to wake up multiple processing threads in the thread pool and process the image frames in the temporary container in parallel, including:

[0151] A format conversion unit, used to convert the image frame from Mat format to YUV420 format, and optimize the conversion process by fixed-point calculation;

[0152] A memory mapping unit, used to directly map the converted YUV420 data to the data field of the AVFrame structure of FFmpeg through a shared memory mechanism;

[0153] A data sorting module, configured to store the processed image frames into a global sorting container according to the serial numbers, where the global sorting container is automatically sorted in ascending order according to the serial numbers, control the atomicity of multi-threaded insertion operations through a global mutex lock, and at the same time count the number of completed frames through an atomic counter;

[0154] A dynamic streaming pushing module, configured to, when the atomic counter reaches a preset threshold, obtain data from the global sorting container in sequence and push it to a remote server, and dynamically adjust the capacity of the temporary container and the thread pool size according to the difference between the current processing speed and the target speed recorded by an error accumulator.

[0155] In a preferred embodiment, the format conversion unit includes a fixed-point calculation module, configured to:

[0156] Round the floating-point coefficients in the YUV conversion formula after multiplying them by 256, and divide the result by 256 to restore the precision after the conversion. The specific conversion formula is as follows:

[0157] Y = (77R + 150G + 29B) / 256

[0158] U = (-38R - 74G + 112B) / 256

[0159] V = (157R - 132G - 26B) / 256

[0160] Wherein, R, G, and B are the three components of the Mat-format image, and Y, U, and V are the three components of the YUV420 format.

[0161] In a preferred embodiment, the dynamic streaming module includes an error accumulator, configured to:

[0162] Record the difference between the current processing speed and the target speed, and trigger an adjustment according to whether the difference exceeds the set threshold range. When the difference exceeds the set threshold range, adjust the number of image frames extracted from the next batch of temporary containers proportionally, and synchronously adjust the size of the thread pool.

[0163] In a preferred embodiment, the thread pool wake-up strategy of the parallel processing module is configured to:

[0164] When the number of image frames in the temporary container reaches the preset batch threshold, wake up the same number of threads as the batch number for parallel processing.

[0165] In a preferred embodiment, the format conversion unit further includes a GPU acceleration module, configured to:

[0166] Call the GPU parallel accelerator for fixed-point calculation, and implement pixel-level parallel processing through CUDA kernel functions;

[0167] Dynamically enable or disable the GPU acceleration function according to the system power or computing power load.

[0168] In a preferred embodiment, the global sorting container is a thread-safe ordered mapping table (map), configured to:

[0169] Store data in a key-value pair structure, where the key is the serial number of the image frame and the value is the corresponding processing result;

[0170] Automatically sort in ascending order by serial number;

[0171] Control the atomicity of multi-threaded insertion operations through mutex locks.

[0172] It can be understood that the various modules described in the device for image conversion and streaming with enhanced real-time performance andFigure 4 Corresponding to each step in the described method for image conversion and streaming with enhanced real-time performance. Thus, the steps, features, and beneficial effects described above for the method of image conversion and streaming based on enhanced real-time performance are equally applicable to the apparatus for image conversion and streaming with enhanced real-time performance and the modules included therein, and will not be elaborated herein.

[0173] In summary, through the combination of multi-thread synchronization, fixed-point calculation optimization, shallow copy memory reuse, and dynamic adjustment strategy, the present invention realizes the efficient conversion of high-resolution video streams and low-latency streaming, providing a reliable technical solution for real-time video transmission scenarios such as unmanned aerial vehicles and edge computing devices.

[0174] The specific embodiments of the present invention have been described in detail above, but they are only examples, and the present invention is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications and substitutions made to the present invention are also within the scope of the present invention. Therefore, all equivalent transformations and modifications made without departing from the spirit and scope of the present invention should be covered within the scope of the present invention.

Claims

1. A method for image conversion and streaming with enhanced real-time performance, characterized in that: The steps include: Acquire original image data from an image source through a producer thread, and store the original image data in a double-ended queue in order; Extracting multiple image frames from the double-ended queue in sequence through a conveying thread model, assigning a unique serial number to each image frame according to the order in the queue, and storing them in a temporary container; Awaken multiple processing threads in the thread pool and perform the following processing on the image frames in the temporary container in parallel: Convert image frames from Mat format to YUV420 format, using fixed-point calculation to optimize the conversion process; Map the converted YUV420 data directly to the data field of FFmpeg's AVFrame structure through the shared memory mechanism; The processed image frames are stored in a global sorting container according to the sequence numbers, wherein the global sorting container automatically arranges the frames in ascending order according to the sequence numbers, controls the atomicity of the multi-threaded insertion operation through a global mutex lock, and counts the number of completed frames through an atomic counter; When the count of the atomic counter reaches a preset threshold, the streaming thread obtains data from the global sorting container in sequence and streams it to the remote server; According to the difference between the current processing speed recorded by the error accumulator and the target speed, the capacity of the temporary container and the size of the thread pool are dynamically adjusted to match the data extraction speed with the processing speed; wherein the processing speed is the actual number of frames processed per second, and the target speed is the target number of frames processed per second; the process of dynamically adjusting the capacity of the temporary container and the size of the thread pool includes: judging according to the difference between the current processing speed recorded by the error accumulator and the target speed: When the difference exceeds the set threshold range, the number of image frames extracted by the temporary container in the next batch is adjusted, and the thread pool size is adjusted synchronously so that the number of threads matches the capacity of the temporary container; When the difference is within the set threshold range, the current temporary container capacity and thread pool size are maintained.

2. The method for image conversion and streaming with enhanced real-time performance according to claim 1, characterized in that: The optimization process of the fixed-point calculation includes: Multiply the floating-point coefficient in the YUV conversion formula by 256 and round it to an integer. After the conversion is complete, divide the result by 256 to restore the accuracy. The specific conversion formula is: Y = (77R + 150G + 29B) / 256 U=(-38R-74G+112B) / 256 V=(157R-132G-26B) / 256 Among them, R, G, B are the three components of the Mat format image, and Y, U, V are the three components of the YUV420 format.

3. The method for image conversion and streaming with enhanced real-time performance according to claim 1, characterized in that: The wake-up strategy of the thread pool is: when the number of image frames in the temporary container reaches a preset batch threshold, threads equal to the batch number are woken up for parallel processing.

4. The method for image conversion and streaming with enhanced real-time performance according to claim 1, characterized in that: The step of converting the image frame from Mat format to YUV420 format also includes: Call the GPU parallel accelerator for fixed-point calculations and implement pixel-level parallel processing through CUDA kernel functions; Dynamically enable or disable GPU acceleration based on system power or computing load.

5. A device for image conversion and streaming with enhanced real-time performance, characterized in that: include: A data acquisition module, configured to obtain raw image data from an image source through a producer thread, and store the raw image data in a double-ended queue in sequence; A batch extraction module is configured to extract multiple image frames from the double-ended queue in sequence through a conveying thread model, assign a unique serial number to each image frame according to the order in the queue, and then store them in a temporary container; The parallel processing module is configured to wake up multiple processing threads in the thread pool and process the image frames in the temporary container in parallel, including: The format conversion unit is used to convert the image frame from Mat format to YUV420 format, and the fixed-point calculation is used to optimize the conversion process; The memory mapping unit is used to map the converted YUV420 data directly to the data field of the FFmpeg AVFrame structure through the shared memory mechanism; A data sorting module is configured to store the processed image frames into a global sorting container according to the sequence number, wherein the global sorting container automatically arranges the frames in ascending order according to the sequence number, controls the atomicity of the multi-threaded insertion operation through a global mutex lock, and counts the number of completed frames through an atomic counter; The dynamic streaming module is configured to obtain data from the global sorting container in sequence and stream it to the remote server when the atomic counter reaches a preset threshold, and dynamically adjust the capacity of the temporary container and the thread pool size according to the difference between the current processing speed recorded by the error accumulator and the target speed; wherein the processing speed is the actual number of frames processed per second, and the target speed is the target number of frames processed per second; the process of dynamically adjusting the capacity of the temporary container and the thread pool size includes: judging according to the difference between the current processing speed recorded by the error accumulator and the target speed: When the difference exceeds the set threshold range, the number of image frames extracted by the temporary container in the next batch is adjusted, and the thread pool size is adjusted synchronously so that the number of threads matches the capacity of the temporary container; When the difference is within the set threshold range, the current temporary container capacity and thread pool size are maintained.

6. The device for image conversion and streaming with enhanced real-time performance according to claim 5, characterized in that: The format conversion unit includes a fixed-point calculation module configured as follows: Multiply the floating-point coefficient in the YUV conversion formula by 256 and round it to an integer. After the conversion is complete, divide the result by 256 to restore the accuracy. The specific conversion formula is: Y = (77R + 150G + 29B) / 256 U=(-38R-74G+112B) / 256 V=(157R-132G-26B) / 256 Among them, R, G, B are the three components of the Mat format image, and Y, U, V are the three components of the YUV420 format.

7. The device for image conversion and streaming with enhanced real-time performance according to claim 5, characterized in that: The thread pool wake-up strategy of the parallel processing module is configured as follows: When the number of image frames in the temporary container reaches a preset batch threshold, threads equal to the batch number are awakened for parallel processing.

8. The device for image conversion and streaming with enhanced real-time performance according to claim 5, characterized in that: The format conversion unit also includes a GPU acceleration module configured as follows: Call the GPU parallel accelerator for fixed-point calculations and implement pixel-level parallel processing through CUDA kernel functions; Dynamically enable or disable GPU acceleration based on system power or computing load.

9. The device for image conversion and streaming with enhanced real-time performance according to claim 5, characterized in that: The global sorting container is a thread-safe ordered mapping table, configured as follows: The data is stored in a key-value pair structure, where the key is the sequence number of the image frame and the value is the corresponding processing result; Automatically sort in ascending order by serial number; The atomicity of multi-threaded insert operations is controlled by mutex locks.

Citation Information

Patent Citations

  • Multi-view vision camera synchronous acquisition method based on multiple threads

    CN113434304A

  • Unmanned aerial vehicle inspection method and system

    CN114895701A