Image preprocessing device and image processing system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-27
- Publication Date
- 2026-08-11
AI Technical Summary
[0002]在对图像进行预处理的过程中,例如,对图像部分区域进行截取时,处理效率低,延迟较高,难以满足低延迟、高吞吐的应用场景
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical terms in the embodiments of this application will be explained below.
Smart Images

Figure CN122269045B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and more specifically to an image preprocessing apparatus and an image processing system. Background Technology
[0002] In image preprocessing, such as cropping a portion of an image, the processing efficiency is low and the latency is high, making it difficult to meet the needs of low-latency, high-throughput applications. Therefore, how to achieve low-latency, high-throughput image region cropping has become an urgent problem to solve. Summary of the Invention
[0003] In view of the above problems, embodiments of this application provide an image preprocessing apparatus and an image processing system for improving image processing speed.
[0004] According to a first aspect of the embodiments of this application, an image preprocessing apparatus is provided, comprising: a streaming image slicing module configured to synchronously perform the following operations in a pipeline manner during the transmission of a video pixel stream carrying a video frame on a bus in the form of a parallel bus pixel stream: slicing data segments from the parallel bus pixel stream according to predetermined slicing parameters; reassembling the slicing data segments into continuous data blocks matching the bit width format of the bus; and outputting a sub-image pixel stream corresponding to the predetermined slicing parameters based on the continuous data blocks; wherein the streaming image slicing module is configured such that the processing delay from slicing data segments to outputting the sub-image pixel stream is a predetermined number of clock cycles, and the predetermined number of clock cycles is independent of the spatial resolution of the video frame.
[0005] A second aspect of this application provides an image processing system, comprising: an image preprocessing device, which is an apparatus according to the first aspect of this application, configured to output sub-image pixel streams corresponding to at least one source video stream; and a timing device configured to, for at least one source video stream, obtain a time offset based on the processing delays of the image preprocessing device and the target detection device, and the output delay of the image fusion device for buffering source video frames; generate a frame data extraction instruction based on the time offset and the current time, instructing the image fusion device to extract source video frames earlier than the time offset than the current time, so that the image fusion device extracts the source video frames corresponding to the frame data extraction time from the frame buffer queue according to the frame data extraction time indicated by the frame data extraction instruction, and outputs a fused video frame based on the source video frames corresponding to at least one source video stream and second position information, wherein the second position information is obtained by the target detection device by converting first position information into object position information in the source video frame according to a predetermined position mapping relationship, and the first position information is the object position information in the sub-image obtained by the target detection device based on the sub-image pixel stream, and the predetermined position mapping relationship characterizes the position mapping relationship between the pixels of the sub-image and the pixels of the source video frame.
[0006] According to a third aspect of the embodiments of this application, an image preprocessing method is provided, comprising: using a streaming image slicing module to synchronously perform the following operations in a pipeline manner during the transmission of a video pixel stream carrying a video frame on a bus in the form of a parallel bus pixel stream: slicing data segments from the parallel bus pixel stream according to predetermined slicing parameters; reassembling the slicing data segments into continuous data blocks that match the bit width format of the bus; and outputting a sub-image pixel stream corresponding to the predetermined slicing parameters based on the continuous data blocks; wherein the streaming image slicing module is configured such that the processing delay from slicing data segments to outputting the sub-image pixel stream is a predetermined number of clock cycles, and the predetermined number of clock cycles is independent of the spatial resolution of the video frame.
[0007] According to a fourth aspect of the present application, an electronic device is provided, comprising: the apparatus described in the first aspect of the present application. Attached Figure Description
[0008] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0009] Figure 1 A schematic diagram of an image preprocessing apparatus provided according to an embodiment of this application is shown;
[0010] Figure 2 A schematic diagram illustrating the process of obtaining a parallel bus effective pixel stream from a parallel bus pixel stream according to an embodiment of this application is shown.
[0011] Figure 3 A schematic diagram of a color component separation operation provided according to an embodiment of this application is shown;
[0012] Figure 4 A schematic diagram of a data segment extraction operation according to an embodiment of this application is shown;
[0013] Figure 5 A schematic diagram of a color component merging operation provided according to an embodiment of this application is shown;
[0014] Figure 6 A schematic diagram of another image preprocessing apparatus provided according to an embodiment of this application is shown;
[0015] Figure 7 A schematic diagram illustrating the principle of a hardware format conversion according to an embodiment of this application is shown;
[0016] Figure 8 A schematic diagram illustrating the principle of another hardware format conversion provided according to an embodiment of this application is shown;
[0017] Figure 9AA schematic diagram of the structure of an image processing system according to an embodiment of this application is shown;
[0018] Figure 9B A schematic diagram of another image processing system according to an embodiment of this application is shown;
[0019] Figure 10 An application scenario diagram for ultra-large scene image processing according to an embodiment of this application is shown. Detailed Implementation
[0020] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0021] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0022] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0023] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0024] In the technical solution of this application, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical terms in the embodiments of this application will be explained below.
[0026] Ultra-large scene images refer to images or video frames with spatial resolution higher than the conventional HD or 4K standards, captured using high-resolution camera equipment (e.g., megapixel array cameras), and the data volume of a single frame can be in the megapixel range.
[0027] An accelerator can refer to an integrated circuit whose logic functions can be configured through programming. This allows the accelerator's inputs, outputs, and internal logic relationships to be configured according to business needs without redesigning the hardware. For example, an accelerator can be a programmable device. Programmable devices can include at least one of the following: programmable logic devices, field-programmable gate arrays (FPGAs), complex programmable logic devices, programmable logic arrays, programmable array logic, or general-purpose array logic, etc.
[0028] Streaming image slicing refers to the real-time segmentation of video frames into multiple independent sub-image regions during the transmission and processing of video frames in a streaming manner, without first storing the complete video frame in memory before segmentation. In the embodiments of this application, it can refer to the process of synchronously capturing sub-image regions while the video pixel stream flows through the bus.
[0029] Streaming image slicing module: A hardware logic unit that implements the above-mentioned streaming image slicing function. During the transmission of the video pixel stream carrying the video frame on the bus in the form of a parallel bus pixel stream, the module captures, reassembles, and outputs a sub-image pixel stream in real time according to predetermined slicing parameters. Its processing delay can be determined by a predetermined number of clock cycles and is independent of the spatial resolution of the video frame.
[0030] Parallel bus pixel stream: This refers to the form in which image pixel data is transmitted via an internal chip bus or an inter-board bus. In this form, a data packet with a fixed bit width is transmitted per clock cycle. This data packet may contain partial or complete color component data of one or more pixels. The data is presented as a continuous stream synchronized with the pixel clock.
[0031] Predefined cropping parameters: A set of parameters used to define the sub-image region to be cropped from the parallel bus pixel stream. These must include at least the sub-image's starting horizontal coordinates, starting vertical coordinates, width, and height within the source video frame.
[0032] Processing latency: refers to the number of clock cycles elapsed from the capture of a data segment to the output sub-image pixel stream. This processing latency does not change with the spatial resolution of the input video frames.
[0033] Video timing signals: A set of control signals accompanying the video pixel stream, used to define the timing structure of video frames. They typically include: a line synchronization signal, indicating the start boundary of the scan line of the video frame; a frame synchronization signal, indicating the start boundary of the video frame; and a data enable signal, indicating whether a pixel is a valid pixel.
[0034] Hardware format conversion module: A dedicated module implemented in hardware logic for changing the format of image data. In this application, it is mainly used to convert data in the YUV color space to data in the RGB color space. Highly pipelined and parallelized designs are typically employed to achieve high-speed conversion.
[0035] Lookup table computation: a hardware-accelerated computing method. It pre-calculates and stores in memory all possible input-output correspondences for complex mathematical operations (such as matrix multiplication in YUV-to-RGB conversion). During actual operation, the input data is directly used as the address index to look up the table, and the calculation result can be obtained in a single or very few clock cycles, offering the advantage of trading storage resources for computing speed and power consumption.
[0036] Half-plane format: A storage format for YUV image data. The luminance component (Y) is stored in a continuous plane, while the two chrominance components (U / V) are interleaved and stored in the same plane, and the spatial resolution of the chrominance components is typically half that of the luminance components.
[0037] Heterogeneous computing refers to an architecture within the same computer system that uses two or more different types of processing units. These processing units are each adept at different types of computing tasks, and they work together to improve the overall system performance and efficiency.
[0038] Image fusion device: mainly responsible for receiving source video streams, performing software decoding, multi-channel video splicing, superimposing second position information, and finally outputting fused video frames for display.
[0039] The system architecture for ultra-large scene image processing includes a computational camera fusion unit and an artificial intelligence (AI) analyzer, where the computational camera fusion unit comprises multiple graphics processing units (GPUs). In the processing flow, the network interface card (NIC) reads the camera array stream data, which is then H.264 decoded by the central processing unit (CPU). Multiple video streams are then stitched together by stitching software, and finally, the images are scaled or cropped by the GPUs before being sent to the GPUs for multi-target tracking and detection. Because ultra-large scene images have a large overall data volume, reaching hundreds of millions of pixels, the backend AI model processing platform experiences significant latency during image loading and processing, failing to achieve real-time performance and typically only reaching 2 frames per second. Therefore, improving the processing speed of the aforementioned system to achieve higher real-time performance is a pressing technical problem that needs to be solved in this field.
[0040] This application aims to provide a high-efficiency, low-latency image processing technology solution to address the throughput limitations and high latency issues encountered in real-time processing of high-resolution video streams. The core of this application lies in the hardware-level architectural reshaping of the image preprocessing workflow and the construction of a complete processing system with precise synchronization capabilities.
[0041] First, this application proposes an image preprocessing apparatus. This image preprocessing apparatus employs a streaming processing architecture. Its core is a streaming image slicing module, configured to synchronously and pipelinedly extract and reassemble sub-image data based on predetermined slicing parameters during the transmission of the video pixel stream along a parallel bus pixel manifold, outputting a sub-image pixel stream corresponding to the predetermined slicing parameters. This solidifies the delay in sub-image extraction to a fixed clock cycle independent of the input video frame resolution, thereby eliminating the performance bottleneck caused by high-resolution data access.
[0042] To construct a complete preprocessing pipeline from the source video stream to the sub-image pixel stream and further reduce system latency through hardware acceleration of format conversion, the device further integrates a hardware format conversion module. This module is configured to convert the source video pixel stream and output a video pixel stream. Optimized for common formats such as YUV half-plane, through the collaborative design of the luma component buffer unit, chroma component buffer unit, and format conversion unit, high-throughput, pipelined color space conversion matching the chroma subsampling characteristics is achieved. To form an end-to-end hardware acceleration path, the front end of the image preprocessing device also deploys a hardware decoding module and a data transmission control module. The hardware decoding module is used to hardware decode the compressed source video stream, and the data transmission control module is used to read the source video frame stream from the video frame buffer module via the direct memory access transmission channel and output it.
[0043] The above modules can be flexibly deployed on a single or collaborative hardware accelerator and can be replicated to form multiple parallel processing links to support the synchronous processing of multiple video streams.
[0044] Based on the aforementioned image preprocessing device, this application further proposes an image processing system. This system integrates an image preprocessing device, a target detection device, an image fusion device, and a timing device. The timing device obtains a time offset based on the processing delays of the image preprocessing device and the target detection device, as well as the output delay of the image fusion device for caching source video frames. It then generates a frame data extraction instruction based on the time offset and the current time. This frame data extraction instruction instructs the image fusion device to extract source video frames from the frame buffer queue that match the current AI analysis result time, and to perform superposition and fusion. This ensures that the background video and foreground analysis annotations in the final output image are synchronized in the time dimension, solving the problem of frame misalignment between heterogeneous processing flows.
[0045] Figure 1 A schematic diagram of the structure of an image preprocessing apparatus provided according to an embodiment of this application is shown.
[0046] like Figure 1 As shown, the image preprocessing apparatus 110 may include a streaming slicer module 111. The streaming slicer module 111 may include at least one. The image preprocessing apparatus 110 can be used to output sub-image pixel streams corresponding to at least one source video stream. Each of the at least one video frame stream may have a corresponding streaming slicer module 111 for processing the video frame stream. The video frame stream may be obtained by format conversion of the source video frame stream. The source video frame stream may be obtained by decoding the source video stream. For example, the source video frame stream may be obtained by hardware decoding of the source video stream. Any streaming slicer module 111 can be used to process at least one video frame stream.
[0047] The streaming image slicing module 111 can be deployed on an accelerator. At least one streaming image slicing module 111 may each have a deployed accelerator. Multiple streaming image slicing modules 111 may have the same or different deployed accelerators. The deployment method of the streaming image slicing module 111 and the number of video frame streams it can process can be determined according to actual business needs and accelerator performance; this embodiment does not limit this. The accelerator may include a field-programmable gate array (FPGA), a complex programmable logic device (FPGA), a programmable logic array (PLA), a programmable array logic, or a general-purpose array logic.
[0048] For any video frame stream from at least one video frame stream, the corresponding streaming slicing module 111 can synchronously perform streaming slicing operations in a pipelined manner while the video pixel stream carrying the video frame is transmitted on the bus in the form of a parallel bus pixel stream. The parallel bus pixel stream can refer to pixels being continuously transmitted via the bus in the form of data packets of a predetermined bit width, driven by a clock. The predetermined bit width can be configured according to actual business needs and is not limited here. For example, the predetermined bit width can be 512 bits.
[0049] Synchronization in a pipelined manner refers to the simultaneous execution of streaming image slicing and data transmission, rather than waiting for the entire video frame to arrive before processing. In other words, the initiation of the slicing operation does not depend on the availability of a complete video frame. Specifically, when the video pixel stream flows through the bus as a parallel bus pixel stream, it simultaneously flows through the streaming image slicing module 111. The streaming image slicing module 111 can be driven by a pixel clock signal originating from the same source as the bus clock, so that the flowing pixel data, under the synchronous control of the pixel clock signal, sequentially passes through each part of the streaming image slicing module 111, thereby completing the streaming image slicing operation during data transmission. This streaming image slicing operation reduces the access latency required to temporarily store pixel data in external memory, achieving hardware-level synchronization between data transmission and processing. For example, the streaming image slicing module can be configured to perform the slicing operation after receiving the first valid pixel of the video frame pixel stream.
[0050] Data segments are extracted from the parallel bus pixel stream according to predetermined truncation parameters. These predetermined truncation parameters can be used to define the sub-image region to be extracted within the video frames carried by the parallel bus pixel stream. For example, the predetermined truncation parameters may include predetermined start coordinates and predetermined size parameters for the sub-image. The predetermined start coordinates may be used to characterize predetermined vertex coordinates of the sub-image region. Vertices may include the top-left vertex.
[0051] Once a data segment is obtained, the truncated data segment can be reassembled into a continuous data block that matches the bus's bit width format. The bit width format can be configured according to actual business requirements and is not limited here. For example, the truncated data segments can be assembled into a continuous data segment that conforms to the bus's bit width format. Since pixel data flows through the bus in row scan order, the streaming tiling module 111 determines whether the pixel data on the current bus belongs to the sub-image region through internal logic during the clock cycle and performs truncation. The streaming tiling module 111 reassembles the truncated data segment into a continuous data block that matches the bus's bit width format. This process involves data repackaging and bit width alignment to ensure that the output continuous data block format is standardized, facilitating subsequent module processing. Thus, based on the continuous data block, a sub-image pixel stream corresponding to predetermined truncation parameters can be output.
[0052] The streaming slice module 111 can be configured such that the processing latency from the captured data segment to the output sub-image pixel stream can be deterministic. For example, the streaming slice module 111 can be configured such that the processing latency from the captured data segment to the output sub-image pixel stream can be determined based on a predetermined number of clock cycles. The predetermined number of clock cycles can be determined based on the configurable logic structure of the streaming slice module 111. For example, the processing latency can be determined based on the pipeline depth, excluding the buffering time of video frames. The pipeline depth can refer to the number of hardware stages (e.g., register stages) that the process from the captured data segment to the output sub-image pixel stream needs to pass sequentially. The time required for the number of hardware stages can be determined based on clock cycles. For example, each register stage introduces a one-clock-cycle latency.
[0053] Furthermore, the predetermined number of clock cycles is independent of the spatial resolution of the video frame, meaning that the processing latency does not increase with the increase of the spatial resolution of the video frame. Since the streaming tiling operation operates in a flow-through-processing manner, its processing latency can be determined by the pipeline depth. Regardless of the spatial resolution of the input video frame, for any sub-image, the streaming tiling operation completes processing at the moment the pixel data corresponding to that sub-image flows through. Therefore, the processing latency is independent of the spatial resolution of the complete video frame (i.e., the total number of pixels) and is related to the number of clock cycles required by the streaming tiling module 111 to process the video frame pixel stream of a single sub-image.
[0054] In one implementation, the streaming image slicing module 111 may include an on-chip storage module. The streaming image slicing module can be used to write consecutive blocks of data to the on-chip storage module. For example, the on-chip storage module may include the block memory of an accelerator (e.g., a programmable logic device). The total capacity of the on-chip storage module may be less than the capacity required to store a complete video frame. For example, the total capacity of the on-chip storage module may be less than K times the capacity required to store a line of pixel data, where K can be an integer greater than 1 and K is less than the vertical resolution of the video frame. Embodiments of this application may eliminate the need to store the complete video frame in on-board memory.
[0055] According to embodiments of this application, the streaming image slicing module is designed such that the entire processing latency from data segment extraction to the output sub-image pixel stream is a fixed, predetermined number of clock cycles, for example, tens of clock cycles. Since the processing latency primarily depends on the pipeline stages and reassembly logic latency within the module, and does not require waiting for the entire frame of high-resolution data to load, this predetermined number of clock cycles is independent of the spatial resolution of the processed video frame. Compared to the memory access latency in traditional solutions, which is directly related to image size, this application significantly reduces preprocessing latency, achieving near real-time sub-image extraction.
[0056] According to embodiments of this application, by configuring a streaming image slicing module to synchronously perform data segment truncation and reassembly operations during the bus transmission of the video pixel stream, the image preprocessing device achieves the fusion of the sub-image extraction process and the data flow. Since the truncation and reassembly operations directly act on the data packets flowing through the bus and proceed strictly in a pipelined manner, the entire processing flow is determined only by a fixed number of logic stages and clock cycles. This design solidifies the processing latency from truncation to output to a predetermined, finite number of clock cycles, and this latency is completely independent of the size or resolution of the input video frame. Therefore, the image preprocessing device can output the sub-image pixel stream with constant and extremely low latency, thereby achieving low-latency, high-throughput image region truncation.
[0057] The following will combine Figures 2-7 This section introduces streaming image slicing operations. Streaming image slicing operations can include invalid pixel filtering (i.e., bus bubble reconstruction), color component separation (i.e., image channel splitting), sub-image cropping, and color component merging (i.e., image channel merging).
[0058] As one implementation, the streaming tiling module 111 can be configured to separate and extract corresponding data segments of multiple color components constituting a pixel from a parallel bus pixel stream according to predetermined tiling parameters. For example, the predetermined tiling parameters may include predetermined starting coordinates and predetermined size parameters of the sub-image.
[0059] The streaming image slicing module 111 can separate and extract the corresponding data segments of multiple color components constituting a pixel from the parallel bus pixel stream according to predetermined starting coordinates and predetermined size parameters. For example, in RGB888 format, pixels are represented by 24 bits of data and may be transmitted in an interleaved manner on the bus. By using predetermined starting coordinates and predetermined size parameters, flexible and accurate extraction of regions of interest in video frames can be achieved, meeting diverse business needs and enhancing application flexibility.
[0060] According to embodiments of this application, by specifying the predetermined cropping parameters as the predetermined starting coordinates and predetermined size parameters of the sub-image, the streaming image cropping module obtains a precise basis for geometric analysis and positioning of the original image. Based on these parameters, the module can accurately identify and separate all color component data belonging to the target region from the video pixel stream, enabling the device to support the extraction of regions of arbitrary specified shapes and positions in the original image, achieving flexibility and spatial accuracy in sub-image cropping.
[0061] As one implementation, a video frame pixel stream can include a valid pixel range, a line blanking range, and a field blanking range. The streaming slicing module can be configured to filter out invalid pixels from the parallel bus pixel stream to obtain a valid pixel stream, and then, according to predetermined truncation parameters, separate and extract the corresponding data segments of multiple color components constituting a pixel from the valid pixel stream. For ease of understanding, the following section combines... Figure 2 Please provide an explanation.
[0062] Figure 2 A schematic diagram is shown illustrating the process of obtaining a parallel bus effective pixel stream from a parallel bus pixel stream according to an embodiment of this application.
[0063] like Figure 2 As shown, the streaming slicing module 111 can filter out invalid pixels from the parallel bus pixel stream based on the video timing signal or an internally generated equivalent control signal, obtaining a parallel bus valid pixel stream that includes valid pixels. In this parallel bus valid pixel stream, there are no gaps between valid pixels in adjacent rows, and there are also no gaps between valid pixels within the same row, thus forming a compact and continuous valid pixel stream. Then, from this compact parallel bus valid pixel stream, according to predetermined truncation parameters, the corresponding data segments of multiple color components constituting a pixel are separated and extracted.
[0064] According to embodiments of this application, by filtering out invalid pixels from the parallel bus pixel stream, the streaming tiling module reassembles the input video pixel stream into a parallel bus valid pixel stream containing only consecutive valid pixels. This preprocessing step eliminates the interference of blanking period data on the tiling logic, providing a clean and continuous data view for subsequent coordinate-based tiling operations. This simplifies the logical complexity of subsequent tiling of corresponding data segments according to predetermined tiling parameters, improving the tiling efficiency and accuracy.
[0065] As one implementation, the streaming slicing module 111 can be configured to filter out invalid pixels from the parallel bus pixel stream based on the video timing signal to obtain the valid pixel stream of the parallel bus.
[0066] Video timing signals may include a data enable signal. The data enable signal can be used to indicate whether a pixel is a valid pixel. As one implementation, a high level on the data enable signal indicates that pixel data on the bus is a valid pixel. Alternatively, video timing signals may include a frame synchronization signal, a line synchronization signal, and a data enable signal. The frame synchronization signal can be used to indicate the start boundary of a video frame. The line synchronization signal can be used to indicate the start boundary of a scan line of a video frame.
[0067] According to embodiments of this application, by relying on video timing signals to filter out invalid pixels, it is ensured that the input interface of the streaming image slicing module is consistent with common video transmission specifications. By utilizing the signal combination in frame synchronization, line synchronization, and data enable signals, the valid range of image data can be reliably identified, thereby seamlessly interfacing with various standard-compliant video sources or pre-processing modules. This improves the versatility and integrability of the image preprocessing device, allowing it to be easily integrated into other video processing systems.
[0068] As one implementation, the streaming image slicing module 111 can be configured to perform color component separation on the effective pixel stream of the parallel bus to obtain multiple pixel component streams. Data segments are then extracted from each of the multiple pixel component streams according to predetermined truncation parameters. These extracted data segments are reassembled into continuous data blocks matching the bit width format of the bus. Finally, at least one continuous data block corresponding to each of the multiple color components is merged to output a sub-image pixel stream corresponding to the predetermined truncation parameters. For ease of understanding, the following section combines... Figures 3-5 Please provide an explanation.
[0069] Figure 3 A schematic diagram of a color component separation operation provided according to an embodiment of this application is shown.
[0070] like Figure 3 As shown, an image channel splitting operation is performed on the effective pixel stream of the parallel bus. If the input is a pixel-interleaved format, such as RGBRGB..., it is split into independent pixel component streams, for example, independent R (Red) component stream, G (Green) component stream, and B (Blue) component stream. Multiple pixel component streams can be arranged in spatial order according to the video frame.
[0071] According to embodiments of this application, color component separation enables independent and parallel extraction of each color channel, fully utilizing the hardware's parallel computing potential and improving throughput. Data segments extracted from multiple pixel component streams are reassembled into continuous data blocks matching the bus's bit width format, ensuring output data conforms to the format specification. An optional merging step can adapt to different needs of the back-end processing unit, flexibly providing interleaved or planar output formats, enhancing the system compatibility of the image preprocessing device.
[0072] Figure 4 A schematic diagram of a data segment extraction operation provided according to an embodiment of this application is shown.
[0073] like Figure 4 As shown, the streaming image slicing module 111 can, according to predetermined slicing parameters, extract sub-image regions (e.g., sub-image regions) from independent pixel component streams in parallel. Figure 1 or child Figure 2The corresponding data segment for that color component. Parallel truncation fully utilizes hardware parallelism, significantly improving the processing throughput of multi-channel data. The streaming truncation module 111 can reassemble the data segments truncated from multiple pixel component streams into continuous data blocks that match the bus bit width format. For example, it can pack the truncated, discrete R component data into a continuous R component data block stream. The reassembly operation ensures the regularity and high-efficiency transmission of the output data stream, conforming to standard bus interface requirements.
[0074] Figure 5 A schematic diagram of a color component merging operation provided according to an embodiment of this application is shown.
[0075] like Figure 5 As shown, the streaming slicing module 111 can, as needed, re-merge the continuous data block streams corresponding to multiple color components into a sub-image pixel stream output in a pixel-interleaved format. It should be understood that if the back-end processing unit supports planar data input, the color component merging operation can be omitted, and the sub-image pixel streams of each color component can be directly output. This "separate first, then independently slice, then optionally merge" process facilitates hardware parallelization and, through the optional merging step, provides flexibility to adapt to different back-end processing units, such as image processor models that support different data formats.
[0076] As one implementation, the streaming image slicing module 111 can be configured to generate data gating information based on predetermined slicing parameters and the interleaving format of the effective pixel stream of the parallel bus. Based on the data gating information, corresponding pixel component data is extracted from the effective pixel stream of the parallel bus, and a data segment corresponding to the predetermined slicing parameters is obtained based on the pixel component data. This implementation method has a short implementation path and can be applied to latency-sensitive scenarios.
[0077] The streaming image slicing module 111 can generate data gating information based on predetermined truncation parameters and the data interleaving format of the parallel bus effective pixel stream. This data gating information can be used to identify pixel component data in the flowing parallel bus pixel stream that corresponds to the sub-image corresponding to the predetermined truncation parameters. For example, the data interleaving format of the parallel bus effective pixel stream can be used to determine which pixels and which color components are included in the clock cycle data packet. As one implementation, the data gating information can be a gating signal or mask synchronized with the parallel bus effective pixel stream, used to identify the pixel component data belonging to the sub-image in the current bus during the clock cycle of the pixel data flow. In actual operation, the streaming image slicing module 111 can directly truncate the corresponding pixel component data from the parallel bus effective pixel stream based on the data gating information. Then, the pixel component data can be accumulated and bit-width aligned to obtain a data segment corresponding to the predetermined truncation parameters.
[0078] According to embodiments of this application, by pre-generating data gating information based on truncation parameters and stream format, the streaming image tiling module provides a more direct bus data filtering mechanism. This mechanism allows the module to directly extract effective pixel components from the original interleaved data stream in real time based on the gating signal without completely separating the color components. This approach reduces processing steps on the data path, lowers logical complexity and potential latency, and provides an efficient hardware implementation path for achieving ultra-low latency sub-image tiling.
[0079] As one implementation, the streaming slicing module 111 can be configured to perform color component separation on the parallel bus pixel stream, resulting in multiple parallel pixel component streams. According to predetermined truncation parameters, data segments are extracted from each of the multiple parallel pixel component streams, and these extracted data segments are reassembled into continuous data blocks that match the bit width format of the bus. Thus, at least one continuous data block corresponding to each of the multiple color components is merged to obtain a sub-image pixel stream corresponding to the predetermined truncation parameters.
[0080] As another implementation, the streaming tiling module 111 can be configured to generate data gating information based on predetermined truncation parameters and the interleaving format of the parallel bus pixel stream. Based on the data gating information, it can extract corresponding pixel component data from the parallel bus pixel stream and obtain a data segment corresponding to the predetermined truncation parameters based on the pixel component data. The data gating information can be used to identify the pixel component data in the flowing parallel bus pixel stream that corresponds to the sub-image with the predetermined truncation parameters.
[0081] Based on this, the image preprocessing device 110 may further include a hardware format conversion module, a hardware decoding module, and a data transmission control module. For ease of understanding, the following will be combined with... Figures 6-8 Please provide an explanation.
[0082] Figure 6 A schematic diagram of another image preprocessing apparatus provided according to an embodiment of this application is shown.
[0083] like Figure 6 As shown, the image preprocessing device 110 may include a hardware decoding module 112, a data transmission control module 113, a hardware format conversion module 114, and a streaming image slicing module 111.
[0084] The hardware format conversion module 114 can be configured to convert the format of the source video pixel stream and output a video pixel stream.
[0085] The source video pixel stream can perform format conversion on the source video pixel stream from the hardware decoding module 112 and output the converted video pixel stream, which can be used as input to the streaming image slicing module. Format conversion may include color space conversion, for example, from a first color space to a second color space. The first color space can be the YUV color space. The second color space can be the RGB color space.
[0086] According to embodiments of this application, by integrating a hardware format conversion module at the front end of the streaming image slicing module, the image preprocessing apparatus constructs a preprocessing pipeline from the source video pixel stream to the sub-image pixel stream. The hardware format conversion module can directly receive and process video pixel streams in common formats such as YUV, and efficiently perform color space conversion at the hardware level, thereby outputting the pixel format required by the streaming image slicing module. This integrated design eliminates the latency and overhead of external format conversion, achieving a high degree of consistency in preprocessing tasks and overall performance optimization.
[0087] As one implementation, the hardware format conversion module 114 may include a luminance component buffer unit, a chrominance component buffer unit, and a format conversion unit.
[0088] The luma component buffer unit can be configured to receive and buffer the luma component data of one frame of the source video pixel stream, and output the luma component row data in row order. The luma component data includes multiple rows of luma component row data arranged in rows.
[0089] The format conversion unit can be configured to perform a first format conversion between the real-time version of the same row of chroma component row data and the first row of luminance component row data in two consecutive rows, and a second format conversion between the delayed version of the same row of chroma component row data and the second row of luminance component row data, according to a predetermined color space conversion relationship, and output two consecutive target row data in the video pixel stream.
[0090] The chroma component buffer unit can be configured to buffer real-time and delayed versions of the same line of chroma component row data from a source video pixel stream received line by line. Predefined color space conversion relationships are used to convert the source video pixel stream from a first color space to a second color space.
[0091] To facilitate understanding, the following will be combined with... Figure 7 Please provide an explanation.
[0092] Figure 7 A schematic diagram illustrating the principle of a hardware format conversion according to an embodiment of this application is shown.
[0093] like Figure 7 As shown, optimizations have been made for the half-plane format (i.e., the chroma component interleaving format), achieving pipelined and high throughput.
[0094] The luminance component buffer unit receives and buffers the complete luminance component data of a single video frame from the source video pixel stream (e.g., YUV420sp). The luminance component row data is output in row order.
[0095] The chroma component buffer unit buffers the chroma component row data received line by line from the source video pixel stream. The chroma component buffer unit can simultaneously buffer two versions of the same row of chroma component row data: a real-time version and a delayed version. The real-time version represents the chroma component row data of the current row. The delayed version represents the chroma component row data input in the previous cycle, i.e., the delayed version of the current row of chroma component row data.
[0096] The format conversion unit performs the actual color space conversion. For example, in the first format conversion, according to a predetermined color space conversion relationship, the real-time version of the same row of chroma component data is combined with the first row of two consecutive rows of luminance data from the luminance buffer unit to output the first row of target data in the video pixel stream. In the second format conversion, the delayed version of the same row of chroma component data is combined with the second row of two consecutive rows of luminance data to output the second row of target data in the video pixel stream.
[0097] According to embodiments of this application, a specific architecture including a luma component buffer unit, a chroma component buffer unit, and a format conversion unit is employed. Optimized pipeline processing is implemented for format conversion of YUV-type half-plane formats. The luma component buffer unit ensures that Y data can be accessed repeatedly, while the chroma component buffer unit is configured to buffer both real-time and delayed versions of the same row of chroma component data from the source video pixel stream received line by line. This allows the same row of UV data to be simultaneously supplied to two consecutive rows of Y data for two format conversions. This architecture precisely matches the data characteristics of chroma vertical subsampling, enabling the module to continuously output RGB data without interruption at a rhythm of "two rows of Y corresponding to one row of UV," achieving maximum throughput in the conversion process.
[0098] As one implementation, the luminance component buffer unit may include a frame buffer and a first line buffer.
[0099] A frame buffer can be configured to buffer the luminance component data of one frame of a source video pixel stream. A first-line buffer can be configured to buffer luminance component line data from the frame buffer in line order.
[0100] A frame buffer can be used to cache the luminance component data of a complete frame. For example, a frame buffer may include on-chip memory or off-chip memory. A first-line buffer can be used to read and temporarily store luminance component line data line by line from the frame buffer for use by the format conversion unit. For example, the first-line buffer may include a first first-in, first-out (FIFO) queue buffer, i.e., first-line buffer FIFO1. The first-line buffer solves the problem that luminance components need to be reused twice, while the frame buffer provides a smooth data feed.
[0101] As one implementation, the chroma component buffer unit may include at least one second real-time line buffer and a second delayed line buffer FIFO2-2 corresponding to each of the at least one second real-time line buffer FIFO2-1.
[0102] The second real-time line buffer can be configured to buffer the chroma component line data of the source video pixel stream received line by line to provide a real-time version of the chroma component line data. For example, the second real-time line buffer may include a second first-in-first-out queue FIFO-2-1.
[0103] A second delayed line buffer corresponding to the second real-time line buffer can be configured to buffer a row of chroma component line data output from the corresponding second real-time line buffer to provide a delayed version of the chroma component line data. For example, the second delayed line buffer may include a third first-in-first-out queue FIFO-2-2.
[0104] As one implementation method, the hardware format conversion module may also include a timing control unit.
[0105] The timing control unit can be configured to perform the following operations in parallel in response to completing the luminance component data buffering of a video frame: controlling a first line buffer to read luminance component line data from the frame buffer line by line, and controlling at least one second real-time line buffer to receive and buffer the chrominance component line data input line by line.
[0106] When the timing control unit detects that the luminance component data of a video frame has been buffered in the frame buffer, it issues a control signal and performs the following two operations in parallel: 1) It controls the first line buffer to start reading luminance data line by line from the frame buffer. 2) It controls the second real-time line buffer to start receiving and buffering the subsequently input chroma component line data.
[0107] According to embodiments of this application, by further employing a specific structure comprising a frame buffer, a first line buffer, a second real-time line buffer, a second delayed line buffer, and a timing control unit, the hardware format conversion module achieves fine-grained management of the source video pixel stream. The frame buffer addresses the global temporary storage requirement for luminance component data within a frame, while the line buffer provides a smooth source video pixel stream for the conversion unit. The timing control unit coordinates the timing of luminance component data reading and chrominance component data reception, ensuring precise synchronization of the two data streams at the conversion point. These components work together to form a stable, efficient, and precisely controlled pipeline system, ensuring reliable operation of high-throughput conversion.
[0108] As one implementation, the format conversion unit can be configured to, in response to the first line buffer and at least one second real-time line buffer being non-empty, perform a first format conversion on the 2m-1th line of luminance component data from the first line buffer and the mth line of chrominance component data from at least one second real-time line buffer according to a predetermined color space conversion relationship, and output the 2m-1th line of target line data in the video pixel stream.
[0109] The format conversion unit can be configured to, in response to the first line buffer and at least one second delayed line buffer being non-empty, perform a second format conversion on the 2m-th line of luminance component data from the first line buffer and the m-th line of chroma component data from at least one second delayed line buffer, according to a predetermined color space conversion relationship, and output the 2m-th line of target data in the video pixel stream. m can be an integer greater than or equal to 1.
[0110] When the timing control unit determines that both the first line buffer and the second real-time line buffer are not empty, for example, the first line buffer outputs the 2m-1th line of Y data, and the second real-time line buffer outputs the mth line of real-time chroma data, the format conversion unit is triggered. The format conversion unit can read the 2m-1th line of luminance component data and the mth line of chroma component data, perform the first format conversion, and output the 2m-1th line of target line data, for example, the 2m-1th line of RGB data. The mth line of chroma component data after the first format conversion is buffered in the second delayed line buffer as a delayed version of the mth line of chroma component data.
[0111] Once the timing control unit determines that both the first line buffer and the second delayed line buffer are not empty, the format conversion unit is triggered again. The format conversion unit can read the 2m-th line of luminance component data and the m-th line of chrominance component data, perform a second format conversion, and output the 2m-th line of target data.
[0112] According to embodiments of this application, a high-throughput, low-latency pipelined conversion of "two rows of Y data and one row of UV data" is achieved in this manner. The process proceeds with highly regular clock cycles, has deterministic processing latency, sequential memory access, and requires no address, making it suitable for integration into larger-scale real-time processing systems.
[0113] As one implementation method, taking the source video pixel stream as a chroma component interleaving format as an example, to further improve the efficiency of format conversion, the timing control unit initiates the reading of luma component line data and the receiving of chroma component line data in parallel. Luma component line data is read line by line from the frame buffer into the first line buffer. At the same time, chroma component line data is sent to the second real-time line buffer. After participating in the first format conversion, this chroma component line data is buffered into its corresponding second delayed line buffer, thus obtaining the real-time version and the delayed version of the same line of chroma component line data.
[0114] The lookup table method can be applied to the first format conversion. For example, a lookup table is pre-stored to convert 8-bit Y, U, V values into corresponding R, G, B values. This lookup table replaces real-time floating-point multiplication operations. For instance, when the first line buffer outputs the 2m-1th line of luminance component data and the second real-time line buffer outputs the mth line of chrominance component data, the format conversion unit uses the lookup table to combine the luminance and chrominance values of the corresponding pixels as addresses to obtain the converted RGB pixel values, and outputs them as the 2m-1th line of target data.
[0115] When the first buffer outputs the 2mth line of luminance data and the second delayed buffer outputs the same mth line of chrominance data, the format conversion unit again uses a lookup table to combine the luminance and chrominance values of the corresponding pixels as addresses to obtain the converted RGB pixel values, which are then output as the 2mth line of target data. This process is repeated, achieving efficient pipelined conversion on a line-by-line basis, where every two lines of luminance data reuse one line of chrominance data. The overall processing latency is fixed and much lower than that of software implementation, ultimately outputting a continuous stream of video frame pixels for use by subsequent streaming image slicing modules.
[0116] As one implementation, when the source video pixel stream is in chroma component interleaving format, the chroma component buffer unit may include a second real-time line buffer and a second delayed line buffer. Chroma component line data is stored alternately within the same plane, and the chroma component buffer unit only requires a single buffer channel. That is, a second real-time line buffer can be configured to buffer the interleaved chroma component line data received line by line from the source video pixel stream. A second delayed line buffer can be configured to buffer a single line of interleaved chroma component line data output from the second real-time line buffer. This saves hardware resources.
[0117] According to the embodiments of this application, the above-mentioned dual FIFO structure achieves real-time and delayed dual-path supply of chroma data with minimal hardware cost. This allows each line of chroma component data to be received to be combined with two lines of luminance data to generate two lines of RGB data, perfectly matching the characteristic of halving the resolution of chroma components in the vertical direction in the YUV420 sampling format. This achieves high-throughput, pipelined processing with "two lines of Y to one line of UV", which can improve speed by tens of times compared to the software solution of pixel-by-pixel conversion.
[0118] As another implementation, when the source video pixel stream is in chroma component plane format, the chroma component row data can include N independent chroma component row data, and the chroma component buffer unit can include a second component buffer subunit corresponding to each of the N independent chroma components. N can be an integer greater than 1.
[0119] Any of the N second component cache sub-units may include a second real-time line cache and a second delayed line cache.
[0120] The second real-time line buffer can be configured to receive and buffer the independent chroma component line data corresponding to the second real-time line buffer line by line from the source video pixel stream.
[0121] The second delayed line buffer can be configured to buffer a single line of independent chroma component data from the output of the second real-time line buffer.
[0122] To facilitate understanding, the following will be combined with... Figure 8 Please provide an explanation.
[0123] Figure 8 A schematic diagram illustrating the principle of another hardware format conversion provided according to an embodiment of this application is shown.
[0124] like Figure 8 As shown, the source video pixel stream is in chroma component plane format, for example, YUV420p, where the Y, U, and V color components are stored independently. The chroma component data includes N independent chroma component line data streams; for example, N can be 2. The second component buffer subunit includes a second real-time line buffer and a second delayed line buffer corresponding to the second real-time line buffer.
[0125] According to embodiments of this application, this multi-channel parallel structure can efficiently process planar formats. Through modular design, different numbers of independent components can be adapted simply by copying sub-units, enhancing the module's scalability. The module's architecture allows it to be flexibly adapted to different input formats through configuration, improving the applicability of the entire preprocessing device.
[0126] As one implementation, when the source video pixel stream is in chroma component plane format, the chroma component buffer unit may include a second real-time line buffer and a second delayed line buffer.
[0127] The chroma plane format chroma component row data can include N independent chroma component row data. These N independent chroma component row data can be processed to obtain the chroma component row data in the chroma component interleaving format. Therefore, the chroma component buffer unit described above, applicable when the source video pixel stream is in chroma component interleaving format, can be utilized. This saves hardware resources.
[0128] According to embodiments of this application, by designing different buffer unit configurations adaptable to chroma interleaving and chroma plane formats, the hardware architecture flexibility and configurability of the hardware format conversion module are improved. For interleaving formats, a single-path buffer is used; for chroma plane formats, multi-path parallel buffer sub-units are used. This modular design enables the same core conversion logic to efficiently process various input video streams of different formats through changes in the configuration of the front-end buffer structure, expanding the applicability of the image preprocessing device.
[0129] As one implementation, the hardware decoding module 112 can be configured to perform hardware decoding on the source video stream and output the source video frame stream.
[0130] The data transmission control module 113 can be configured to configure a direct memory access transmission channel for transmitting source video frame streams from the video frame buffer module to the outside according to transmission control parameters, read source video frame streams from the video frame buffer module via the direct memory access transmission channel, convert the read source video pixel streams into streaming data format, and output them to the outside.
[0131] This embodiment further extends the front-end data processing chain of the image preprocessing device, constructing an end-to-end hardware acceleration path from compressed bitstream to high-speed serial output.
[0132] The hardware decoding module can be a hard-core decoding IP deployed in a programmable logic device, used to perform hardware decoding on the input compressed source video stream, output the decompressed source video frame stream, improve decoding speed, and free up central processing unit resources.
[0133] The data transmission control module may include a direct memory access engine and a high-speed serial interface controller. Based on configured transmission control parameters (e.g., frame size or destination address), this module establishes a direct memory access transmission channel from the onboard video frame buffer module storing decoded video frames to an external direct memory access (DMI) transmission channel. The DMI engine efficiently reads the source video frame stream through this DMI transmission channel without central processing unit intervention, converts it into a continuous streaming data format, i.e., a source video pixel stream, and sends it to the hardware format conversion module via a high-speed serial interface (e.g., a fiber optic interface based on a high-speed serial point-to-point link protocol). This constitutes an efficient, low-latency data path from decoding to forwarding, enabling high-speed direct data transmission between chips and avoiding the additional latency and bandwidth bottlenecks caused by traversing host memory.
[0134] The hardware decoding module, data transmission control module, hardware format conversion module, and streaming image slicing module have different deployment strategies on different hardware acceleration platforms. The deployment strategy is related to the hardware acceleration platform resources and the user's actual needs.
[0135] As one implementation approach, the streaming image slicing module, hardware format conversion module, hardware decoding module, and data transmission control module can each be deployed on different accelerators. This highly decoupled approach facilitates phased system upgrades and maintenance.
[0136] Optionally, the streaming image slicing module and the hardware format transcoding module can be deployed on the same accelerator. The hardware decoding module and the data transmission control module can be deployed on a different accelerator.
[0137] Optionally, the streaming image slicing module, hardware format conversion module, hardware decoding module, and data transmission control module can each be deployed on the same accelerator. This highly integrated solution is suitable for scenarios with strict requirements on cost and size. Preferably, the streaming image slicing module and hardware format conversion module can be integrated and deployed on the same accelerator to form the core preprocessing acceleration card; while the hardware decoding module and data transmission control module are deployed on another accelerator to form the video access and forwarding card. The two are interconnected via a high-speed link. This dual-card architecture with separate access and processing functions has a clear division of labor, which is conducive to functional specialization and performance optimization.
[0138] As one implementation, at least one streaming image slicing module, at least one hardware format conversion module, at least one hardware decoding module, and at least one data transmission control module are configured to form a processing link corresponding to at least one source video stream. The processing link is used to process the source video stream to obtain a sub-image pixel stream corresponding to the source video stream.
[0139] Each processing link can include the aforementioned streaming image slicing module, hardware format conversion module, hardware decoding module, and data transmission control module, used to independently process one or more source video streams and output the corresponding sub-image pixel stream. This parallel design enables the system to process multiple video inputs simultaneously, such as 16 channels of 4K video, significantly improving overall throughput and system processing capabilities, and meeting the needs of multi-view monitoring in ultra-large scenes.
[0140] According to embodiments of this application, by explicitly defining multiple deployment strategies between modules, such as card splitting or integration, and supporting the design of multi-processing link replication, system-level flexibility and scalability are demonstrated. Different deployment methods adapt to different application scenarios ranging from high integration to high performance. Furthermore, the parallel replication capability of the processing links allows the system to achieve synchronous and parallel processing of multiple video streams by linearly increasing hardware resources, enabling the overall system throughput to be effectively expanded according to demand, meeting the needs of large-scale video analytics applications.
[0141] Figure 9A A schematic diagram of the structure of an image processing system according to an embodiment of this application is shown.
[0142] like Figure 9A As shown, the image processing system may include an image preprocessing unit 110 and a timing unit 120. Additionally, it may include a target detection unit 130 and an image fusion unit 140.
[0143] The image preprocessing unit 110 can be used to output sub-image pixel streams corresponding to at least one source video stream. For example, for any one of the at least one source video streams, the image preprocessing unit 110 can perform the following operations:
[0144] During the transmission of the video pixel stream carrying video frames on the bus in the form of a parallel bus pixel stream, the following operations are performed synchronously in a pipeline manner: according to predetermined truncation parameters, data segments are truncated from the parallel bus pixel stream, the truncated data segments are reassembled into continuous data blocks that match the bit width format of the bus, and according to the continuous data blocks, a sub-image pixel stream corresponding to the predetermined truncation parameters is output.
[0145] The target detection device 130 can be used to obtain first position information of an object in a sub-image, and to convert the first position information into second position information of the object in a source video frame.
[0146] For example, the target detection device 130 can obtain the first position information of the object in the sub-image based on the sub-image pixel stream, and convert the first position information into the second position information of the object in the source video frame according to a predetermined position mapping relationship that characterizes the position mapping relationship between the pixels of the sub-image and the pixels of the source video frame. The predetermined position mapping relationship describes the position and scaling ratio of the sub-image in the source video frame, converting the first position information into the second position information of the object in the coordinate system of the source video frame. By combining the fast preprocessing of the front-end hardware with the powerful artificial intelligence analysis capabilities of the back-end, the extraction and coordinate mapping from pixels to semantic information are completed.
[0147] The timing device 120 can be configured to, for at least one of the source video streams, obtain a time offset based on the respective processing delays of the image preprocessing device and the target detection device, as well as the output delay of the image fusion device for buffering source video frames, and generate a frame data extraction instruction based on the time offset and the current time, instructing the image fusion device to extract source video frames that are earlier than the time offset than the current time.
[0148] The timing device 120, acting as a synchronization controller for target display, is the core component for achieving synchronized "video-analysis" display. It continuously monitors or acquires the following parameters: the total processing latency of the image preprocessing unit, the processing latency of the target detection unit, and the output latency of the image fusion unit for buffering source video frames and preparing them for output.
[0149] The time offset ΔT = total processing delay + processing delay - output delay. Based on this time offset and the current time T, a frame data extraction instruction indicating the time offset is generated. This instruction directs the image fusion device to extract historical video frames with timestamps of (T-ΔT). Through these steps, precise compensation for the total delay of the image processing pipeline is achieved, enabling synchronized display of video analysis.
[0150] The image fusion apparatus 140 can be configured to extract the source video frame corresponding to the frame data extraction time from the frame buffer queue according to the frame data extraction time indicated by the frame data extraction instruction. Based on the source video frame and second position information corresponding to at least one source video stream, it outputs a fused video frame.
[0151] The image fusion apparatus 140 may include a software decoder, a stitcher, and a frame buffer queue. It can extract the source video frame, i.e., the original high-resolution frame, from its internally maintained frame buffer queue according to the frame data extraction time indicated by the frame data extraction instruction from the timing device 120.
[0152] The image fusion device 140 can stitch together at least one source video frame to obtain a stitched video frame. The stitched video frame is then labeled according to at least one second location information to obtain a fused video frame.
[0153] The video frames preceding the video frames obtained by the image fusion device 140 from the video frame buffer queue will be deleted. That is, depending on the time of image preprocessing and target detection processing, the full image display will perform corresponding frame extraction processing.
[0154] Precise control via a timing device reduces display asynchrony issues such as frame misalignment caused by different processing path delays, ensuring the authenticity and accuracy of the large-screen display.
[0155] According to embodiments of this application, the image processing system integrates an image preprocessing unit, a target detection unit, an image fusion unit, and a timing unit. The timing unit calculates a time offset based on the processing delays of the image preprocessing unit and the target detection unit, as well as the output delay of the image fusion unit for buffering source video frames. It then generates a frame data extraction instruction based on the time offset and the current time. This frame data extraction instruction instructs the image fusion unit to extract a source video frame from the frame buffer queue that matches the time of the current image analysis result and to perform superposition and fusion. This ensures that the background video and foreground analysis annotations in the final output image are synchronized in the time dimension, solving the problem of frame misalignment between heterogeneous processing flows.
[0156] In one implementation, the processing latency of the image preprocessing unit can be determined based on the decoding time of the hardware decoding module, the transmission time of the source video pixel stream, the processing time of the hardware format conversion module, and the processing time of the streaming image slicing module. The frame data extraction time can characterize the difference between the total processing latency determined based on the respective processing latencies of the image preprocessing unit and the target detection unit, and the output latency.
[0157] Total processing delay T of the image preprocessing unit preprocess It is the sum of the processing latency and transmission latency of each module, which may include: the decoding time T of the hardware decoding module. decode The transmission time T of the source video pixel stream from the decoding card to the preprocessing card. transfer Processing time T of the hardware format conversion module convert And the fixed processing time T of the streaming image slicing module crop That is, T. preprocess =T decode +T transfer +T convert +T crop .
[0158] In order to ensure that the stitched video frames and the superimposed second position information in the displayed fused video frames correspond in time, the source video frames to be extracted should be earlier than the current time by an amount equivalent to the total time consumed by the image processing pipeline (T). preprocess +T detectThat frame. detect This indicates the processing delay of the target detection device.
[0159] In addition, the cache latency T of the display path itself can also be considered. fusion_buffer To fine-tune the extraction points.
[0160] The timing device achieves millisecond-level synchronization across heterogeneous processing units by precisely calculating and compensating for these delays. This dynamic frame extraction strategy based on precise delay calculation can adaptively handle image processing time fluctuations T caused by scenarios of varying complexity. detect And the transmission time T that may be caused by network jitter transfer It maintains high-quality synchronization throughout the process, demonstrating strong robustness. It resolves the technical issue of "screen and frame misalignment" caused by asynchronous processing paths.
[0161] exist Figure 9A On this basis, Figure 9B A schematic diagram of the structure of another image processing system according to an embodiment of this application is shown.
[0162] The image processing system may include accelerators corresponding to each of the multiple source video streams. A single source video stream may have two accelerators. One accelerator deploys a hardware decoding module and a data transmission control module, as well as a receiving module. The two accelerators can communicate with each other via optical fiber. Figure 9B Other modules can be found in the description above, and will not be repeated here.
[0163] Figure 10 An application scenario diagram for ultra-large scene image processing according to an embodiment of this application is shown.
[0164] like Figure 10 As shown, the system architecture for ultra-large scene image processing includes a computational camera fusion unit and an artificial intelligence analyzer. The artificial intelligence analyzer may include the image preprocessing device, timing device, and target detection device described in the embodiments of this application. The artificial intelligence analyzer can be used to send second location information and frame data extraction time to the image fusion unit.
[0165] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0166] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0167] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.
[0168] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.
Claims
1. An image preprocessing apparatus, characterized in that, include: The streaming slice module is configured to perform the following operations synchronously in a pipelined manner while the video pixel stream carrying video frames is transmitted on the bus as a parallel bus pixel stream: Data segments are extracted from the parallel bus pixel stream according to predetermined extraction parameters; The captured data segments are reassembled into continuous data blocks that match the bit width format of the bus; Based on the continuous data blocks, output a sub-image pixel stream corresponding to the predetermined truncation parameters; The streaming image slicing module is configured such that the processing delay from slicing the data segment to outputting the sub-image pixel stream is a predetermined number of clock cycles, and the predetermined number of clock cycles is independent of the spatial resolution of the video frame.
2. The apparatus according to claim 1, characterized in that, The streaming image slicing module is configured as follows: According to the predetermined truncation parameters, the corresponding data segments of multiple color components constituting a pixel are separated and extracted from the parallel bus pixel stream, wherein the predetermined truncation parameters include the predetermined starting coordinates and predetermined size parameters of the sub-image.
3. The apparatus according to claim 2, characterized in that, The streaming image slicing module is configured as follows: Filtering out invalid pixels from the parallel bus pixel stream yields a parallel bus valid pixel stream, wherein there are no gaps between valid pixels in adjacent rows and no gaps between valid pixels in the same row; and According to the predetermined truncation parameters, the corresponding data segments of multiple color components constituting a pixel are separated and extracted from the effective pixel stream of the parallel bus.
4. The apparatus according to claim 3, characterized in that, The streaming image slicing module is configured as follows: The effective pixel stream of the parallel bus is subjected to color component separation to obtain multiple pixel component streams; According to the predetermined truncation parameters, data segments are extracted from each of the multiple pixel component streams; The data segments extracted from each of the multiple pixel component streams are reassembled into a continuous data block that matches the bit width format of the bus; The color components of at least one consecutive data block corresponding to each of the multiple color components are merged, and a sub-image pixel stream corresponding to the predetermined truncation parameters is output.
5. The apparatus according to claim 3, characterized in that, The streaming image slicing module is configured as follows: Data gating information is generated based on the predetermined truncation parameters and the interleaving format of the parallel bus effective pixel stream, wherein the data gating information is used to identify pixel component data in the parallel bus effective pixel stream that corresponds to the sub-image corresponding to the predetermined truncation parameters; Based on the data gating information, the corresponding pixel component data is extracted from the effective pixel stream of the parallel bus; and Based on the pixel component data, a data segment corresponding to the predetermined truncation parameter is obtained.
6. The apparatus according to claim 3, characterized in that, The streaming image slicing module is configured as follows: Based on the video timing signal, invalid pixels are filtered from the parallel bus pixel stream to obtain the parallel bus valid pixel stream. The video timing signal includes a data enable signal, or the video timing signal includes a frame synchronization signal, a line synchronization signal, and the data enable signal. The data enable signal is used to indicate whether a pixel is a valid pixel. The frame synchronization signal is used to indicate the start boundary of the video frame. The line synchronization signal is used to indicate the start boundary of the scan line of the video frame.
7. The apparatus according to any one of claims 1 to 6, characterized in that, Also includes: The hardware format conversion module is configured to convert the format of the source video pixel stream and output the video pixel stream.
8. The apparatus according to claim 7, characterized in that, The hardware format conversion module includes: A luminance component buffer unit is configured to receive and buffer the luminance component data of a video frame of the source video pixel stream, and output the luminance component row data in row order, wherein the luminance component data includes multiple rows of luminance component row data arranged in rows. The chroma component buffer unit is configured to buffer both real-time and delayed versions of the same line of chroma component line data from the source video pixel stream received line by line; and The format conversion unit is configured to perform a first format conversion between the real-time version of the same row of chroma component data and the first row of luminance component data in two consecutive rows, and a second format conversion between the delayed version of the same row of chroma component data and the second row of luminance component data, according to a predetermined color space conversion relationship, and output two consecutive target rows of data in the video pixel stream. The predetermined color space conversion relationship is used to convert the source video pixel stream from a first color space to a second color space.
9. The apparatus according to claim 8, characterized in that, The luminance component buffer unit includes: A frame buffer is configured to buffer the luminance component data of one video frame from the source video pixel stream; The first line buffer is configured to buffer luminance component line data sequentially from the frame buffer; and / or The chroma component buffer unit includes: At least one second real-time line buffer is configured to buffer the chroma component line data of the source video pixel stream received line by line to provide a real-time version of the chroma component line data; A second delayed line buffer corresponding to at least one of the second real-time line buffers, the second delayed line buffer being configured to buffer a line of chroma component line data output from the corresponding second real-time line buffer to provide a delayed version of the chroma component line data; and / or The hardware format conversion module further includes: The timing control unit is configured to perform the following operations in parallel in response to completing the luminance component data buffering of a video frame: controlling the first line buffer to read luminance component line data line by line from the frame buffer, and controlling at least one second real-time line buffer to receive and buffer the chrominance component line data input line by line.
10. The apparatus according to claim 9, characterized in that, The format conversion unit is configured as follows: In response to the first line buffer and at least one second real-time line buffer being non-empty, according to the predetermined color space conversion relationship, the 2m-1th line of luminance component data from the first line buffer and the mth line of chrominance component data from at least one second real-time line buffer are subjected to a first format conversion, and the 2m-1th line of target line data in the video pixel stream is output. In response to the first line buffer and at least one second delayed line buffer being in a non-empty state, according to the predetermined color space conversion relationship, the 2mth line of luminance component data from the first line buffer and the mth line of chrominance component data from at least one second delayed line buffer are subjected to a second format conversion, and the 2mth line of target data in the video pixel stream is output. Where m is an integer greater than or equal to 1.
11. The apparatus according to claim 9, characterized in that, When the source video pixel stream is in chroma component interleaving format, the chroma component buffer unit includes: A second real-time line buffer is configured to buffer the chroma component line data of the source video pixel stream received line by line and interleaved; and The second delayed line buffer is configured to buffer one line of interleaved chroma component line data from the output of the second real-time line buffer; or When the source video pixel stream is in chroma component plane format, the chroma component row data includes N independent chroma component row data, and the chroma component buffer unit includes: Each of the N independent chroma components corresponds to a second component cache subunit, and any one of the N second component cache subunits includes: A second real-time line buffer is configured to receive and buffer independent chroma component line data corresponding to the second real-time line buffer line by line from the source video pixel stream; and The second delayed line buffer is configured to buffer one line of independent chroma component line data from the output of the second real-time line buffer; Where N is an integer greater than 1.
12. The apparatus according to claim 8, characterized in that, Also includes: The hardware decoding module is configured to perform hardware decoding on the source video stream and output the source video frame stream; The data transmission control module is configured to configure a direct memory access transmission channel for transmitting source video frame streams from the video frame buffer module to the outside according to transmission control parameters, read the source video frame stream from the video frame buffer module through the direct memory access transmission channel, convert the read source video pixel stream into a streaming data format, and output it to the outside.
13. The apparatus according to claim 12, characterized in that, The streaming image slicing module, the hardware format conversion module, the hardware decoding module, and the data transmission control module are each deployed on different accelerators; or The streaming image slicing module and the hardware format transcoding module are deployed on the same accelerator, while the hardware decoding module and the data transmission control module are deployed on another accelerator; or The streaming image slicing module, the hardware format conversion module, the hardware decoding module, and the data transmission control module are each deployed on the same accelerator; or At least one of the streaming image slicing modules, at least one of the hardware format conversion modules, at least one of the hardware decoding modules, and at least one of the data transmission control modules are configured to form a processing link corresponding to at least one source video stream, the processing link being used to process the source video stream to obtain a sub-image pixel stream corresponding to the source video stream.
14. An image processing system, characterized in that, include: An image preprocessing apparatus, wherein the image preprocessing apparatus is the apparatus according to any one of claims 1 to 13, is configured to output a sub-image pixel stream corresponding to at least one source video stream; as well as A timing device is configured to, for at least one of the source video streams, obtain a time offset based on the processing delays of the image preprocessing device and the target detection device, and the output delay of the image fusion device for buffering source video frames; generate a frame data extraction instruction based on the time offset and the current time, instructing the image fusion device to extract source video frames earlier than the current time than the time offset, so that the image fusion device extracts the source video frame corresponding to the frame data extraction time from the frame buffer queue according to the frame data extraction time indicated by the frame data extraction instruction, and outputs a fused video frame based on the source video frames and second position information corresponding to each of the at least one source video stream, wherein the second position information is obtained by the target detection device by converting first position information into the position information of the object in the source video frame according to a predetermined position mapping relationship, and the first position information is the position information of the object in the sub-image obtained by the target detection device according to the sub-image pixel stream, wherein the predetermined position mapping relationship characterizes the position mapping relationship between the pixels of the sub-image and the pixels of the source video frame.
15. The system according to claim 14, characterized in that, The timing device is configured as follows: The time offset is obtained by the difference between the total processing delay determined by the processing delays of the image preprocessing device and the target detection device and the output delay of the image fusion device for caching source video frames. The processing delay of the image preprocessing device is determined based on the decoding time of the hardware decoding module, the transmission time of the source video pixel stream, the processing time of the hardware format conversion module, and the processing time of the streaming image slicing module.
Citation Information
Patent Citations
Streaming character recognition method and device, electronic equipment and storage medium
CN113971808A
Graphics display agent device based on PCIE bus
CN121742787A