Electronic device, image processing method, computer storage medium, and program product
Patent Information
- Application Number
- CN202610748171.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-27
- Publication Date
- 2026-09-01
AI Technical Summary
[0003]然而,在实际应用中,若采用经图像信号处理流程处理后的图像数据进行目标感知,在环境光照复杂多变的场景下,所得到的图像较容易出现动态范围不足、亮区过曝、暗区信息丢失以及色彩失真等问题
[0010] The electronic device, image processing method, computer storage medium, and program product of this disclosure, by directly acquiring the original image, retain the high dynamic range and original color information of the sensor output, enabling the complete preservation of the brightness levels and color differences of the target area under complex lighting conditions, thus providing a reliable information foundation for accurate perception. Based on this, the target object is detected from the original image to obtain a target detection box, and then a region image covering only the target object is cropped according to the detection box. This ensures that the data to be processed and transmitted subsequently focuses on the effective pixels carrying the perception task, rather than all the data of the entire image. Subsequently, the region image is synthesized with a zero-background image to determine the target image, so that only the location of the target object in the target image is filled with the true original pixel value, while the remaining large area of background is zero. These zero-value regions can be efficiently compressed to a very small data volume during encoding compression, while the original information of the target area is not lost. This effectively reduces the data transmission bandwidth and storage space occupation while ensuring the quality of the image information required for perception.
Smart Images

Figure CN122675698A_ABST
Abstract
Description
Technical Field
[0001] This disclosure belongs to the field of image processing technology, and in particular relates to an electronic device, an image processing method, a computer storage medium, and a program product. Background Technology
[0002] Currently, in various visual perception systems for target detection, recognition, or state determination, the image acquisition and transmission methods directly impact the system's final performance. One common approach, to save on data transmission bandwidth and storage pressure, involves converting the raw image data acquired by the sensor into processed image data of a specific format through image signal processing, and then transmitting the processed image data to the perception system for target perception. Another approach is to directly transmit the raw image data to the perception system for target perception, thus preserving more complete original light intensity information and a higher dynamic range.
[0003] However, in practical applications, if image data processed through image signal processing is used for target perception, in scenarios with complex and variable ambient lighting, the resulting images are prone to problems such as insufficient dynamic range, overexposure in bright areas, loss of information in dark areas, and color distortion. This leads to a decrease in the accuracy of target perception, especially for small, distant, or color-sensitive targets. While directly transmitting raw image data to the perception system for target perception largely preserves the image's true dynamic response and color information, it significantly increases transmission bandwidth and storage costs, posing a significant challenge to the overall system resources and real-time performance.
[0004] Therefore, how to balance the accuracy of image perception with the system bandwidth and storage overhead caused by data transmission is a technical problem that the perception system urgently needs to solve. Summary of the Invention
[0005] This disclosure provides an electronic device, image processing method, computer storage medium, and program product that can reduce data transmission bandwidth and storage overhead while improving image perception accuracy.
[0006] On one hand, this disclosure provides an electronic device, which includes: a processor and a memory storing computer program instructions; the processor executes the computer program instructions to implement: Obtain the original image; Detect target objects in the original image and obtain at least one target detection box; The original image is cropped based on at least one target detection box to obtain the region image corresponding to the target object; The target image is determined based on the region image and the zero-background image.
[0007] On the other hand, embodiments of this disclosure provide an image processing method, including: Obtain the original image; Detect target objects in the original image and obtain at least one target detection box; The original image is cropped based on at least one target detection box to obtain the region image corresponding to the target object; The target image is determined based on the region image and the zero-background image.
[0008] In another aspect, embodiments of this disclosure provide a computer storage medium on which computer program instructions are stored, and which implement an image processing method when executed by a processor.
[0009] In another aspect, embodiments of this disclosure provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform an image processing method.
[0010] The electronic device, image processing method, computer storage medium, and program product of this disclosure, by directly acquiring the original image, retain the high dynamic range and original color information of the sensor output, enabling the complete preservation of the brightness levels and color differences of the target area under complex lighting conditions, thus providing a reliable information foundation for accurate perception. Based on this, the target object is detected from the original image to obtain a target detection box, and then a region image covering only the target object is cropped according to the detection box. This ensures that the data to be processed and transmitted subsequently focuses on the effective pixels carrying the perception task, rather than all the data of the entire image. Subsequently, the region image is synthesized with a zero-background image to determine the target image, so that only the location of the target object in the target image is filled with the true original pixel value, while the remaining large area of background is zero. These zero-value regions can be efficiently compressed to a very small data volume during encoding compression, while the original information of the target area is not lost. This effectively reduces the data transmission bandwidth and storage space occupation while ensuring the quality of the image information required for perception. Attached Figure Description
[0011] Figure 1 This is a schematic diagram of the structure of an image processing system provided in one embodiment of the present disclosure; Figure 2 This is a schematic diagram of a processor executing instructions to implement image processing according to another embodiment of the present disclosure; Figure 3 This is a schematic diagram of a processor executing instructions to perform image processing through a target detection box, according to yet another embodiment of this disclosure. Figure 4 This is a schematic diagram of a processor executing instructions to obtain a region image through overlay graphics, according to yet another embodiment of this disclosure; Figure 5 This is a schematic diagram of a processor executing instructions to obtain a regional image from a local image, according to yet another embodiment of this disclosure; Figure 6 This is a schematic diagram illustrating the process of a processor executing instructions to obtain a target image by filling a background image with zeros, according to yet another embodiment of this disclosure. Figure 7 This is a schematic diagram of the process of a processor executing instructions to determine a target detection box through extended processing, according to another embodiment of this disclosure. Figure 8 This is a schematic diagram of a processor executing instructions to perform image processing by augmenting the image through confidence, according to yet another embodiment of this disclosure. Figure 9 This is a schematic diagram of a processor executing instructions to perform image processing through boundary extension, according to another embodiment of this disclosure. Figure 10 This is a schematic diagram illustrating the process of a processor executing instructions to achieve image encoding compression, according to yet another embodiment of this disclosure; Figure 11 This is a schematic diagram of the process of a processor executing instructions to determine a target detection box through alignment processing, according to another embodiment of this disclosure. Figure 12 This is a schematic diagram illustrating the process of a processor executing instructions to perform image processing by expanding the processing based on confidence, according to yet another embodiment of this disclosure. Detailed Implementation
[0012] The features and exemplary embodiments of various aspects of this disclosure will now be described in detail. To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description, in conjunction with the accompanying drawings and specific embodiments, will provide a further detailed description. It should be understood that the specific embodiments described herein are intended only to explain this disclosure and not to limit it. For those skilled in the art, this disclosure can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this disclosure by illustrating examples.
[0013] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0014] It should be noted that the acquisition, storage, use, and processing of data in this embodiment comply with the relevant provisions of national laws and regulations.
[0015] It should be noted that in the embodiments disclosed herein, certain software, components, models, and other existing solutions in the industry may be mentioned. These should be considered as exemplary and are intended only to illustrate the feasibility of implementing the technical solutions disclosed herein. However, they do not mean that the applicant has used or necessarily used such solutions.
[0016] Application Overview Currently, using YUV format image data processed by an ISP (Image Signal Processor) as input to an image perception system is problematic. YUV format images are prone to losing information under complex lighting conditions (such as backlighting, strong light, cloudy days, and nighttime), resulting in overexposure in bright areas or loss of information in dark areas, or color shifts leading to color distortion, thus affecting the accuracy of target detection and recognition. While using RAW (Raw Image Format) data acquired by the sensor can preserve the complete dynamic range and true color information of the scene, thereby improving perception accuracy, RAW image data is significantly larger than YUV format image data, placing a heavy burden on transmission bandwidth and storage resources. Therefore, how to reduce system bandwidth and storage overhead while improving image perception accuracy has become an urgent problem to be solved.
[0017] To address the aforementioned problems, this disclosure provides an image processing system, an electronic device, an image processing method, a computer storage medium, and a computer program product. The image processing system provided in this disclosure will be described first.
[0018] Exemplary System To address the aforementioned technical problems, this disclosure provides an image processing system. The image processing system may include an electronic device, an image sensor module, and a hardware encoder. The electronic device may include a processor and a memory storing computer program instructions, and may also include a hardware encoder. In this system, the image sensor module generates raw image data and sends it to the electronic device; the electronic device acquires the raw image; detects a target object in the raw image to obtain at least one target detection box; crops the raw image based on the at least one target detection box to obtain a region image corresponding to the target object; determines the target image based on the region image and a zero-background image; the hardware encoder acquires the target image data, performs encoding compression, and outputs the target image data, which can be stored or transmitted externally.
[0019] As an example, in the field of intelligent driving, the image sensor module can be an in-vehicle forward-facing camera module, including an optical lens and a CMOS image sensor, responsible for acquiring raw image data of the scene in front of the vehicle and transmitting it to electronic devices. The electronic devices can be hardware such as an in-vehicle domain controller, image processing chip, or in-vehicle computing platform. Specifically, the electronic device receives the raw image data from the forward-facing camera and runs a target detection model, identifying target objects such as traffic lights, pedestrians, and vehicles in the raw image and outputting target detection boxes. Based on the detection boxes, it crops the region image, constructs a zero-background image, and fills the region image into it to generate the target image. A hardware encoder can acquire the target image data and perform encoding compression, outputting compressed target image data. The compressed target image data can be sent to the object recognition module via in-vehicle Ethernet for real-time perception. The target recognition module, based on the target image data, performs accurate category determination and state recognition for various target objects (such as traffic lights, vehicles, pedestrians, traffic signs, etc.). After the target recognition module completes the recognition, it outputs the structured perception results (such as target category, location, speed, traffic light color status, etc.) to the downstream planning and decision-making module, which then performs path planning and vehicle control based on these results.
[0020] As another example, in the field of robotics, the image sensor module can be an industrial camera. An industrial camera may include a global shutter CMOS sensor and an optical lens, responsible for capturing raw image data of equipment or workpiece surfaces and transmitting it to electronic equipment. The electronic equipment can be a robot control unit or a computing chip, etc. Specifically, the electronic equipment receives the raw image data from the industrial camera and runs a target detection model. It detects physical entities such as tables, chairs, and boxes in the raw image and outputs target detection boxes. It then crops the corresponding area of the image and fills it with a zero-background image to obtain the target image. A hardware encoder reads the target image data and performs compression. The compressed target image data is transmitted to the robot's perception and recognition module. The perception and recognition module uses the target image data to accurately identify and determine the state of the detected physical entities, such as determining the category and pose of items on a table. The recognition results are directly input into the robot's decision-making and planning module, driving the robot to perform corresponding operational tasks.
[0021] Exemplary electronic devices like Figure 1 As shown, this disclosure provides an electronic device, which includes: at least one processor 101; and a memory 102 coupled to the processor 101 and storing computer-executable instructions; Memory 102 may include a large-capacity memory for data or instructions. For example, and not limitingly, memory 102 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 102 may include removable or non-removable (or fixed) media. Where appropriate, memory 102 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 102 is non-volatile solid-state memory. Memory 102 may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. Thus, typically, memory 102 includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it can perform the operations described in the object identification method of the embodiments below.
[0022] The processor 101 can be a graphics processing unit (GPU), a neural processing unit (NPU), or other general-purpose accelerators with matrix operation capabilities.
[0023] The processor 101 can be integrated into edge computing devices or embedded systems for scenarios such as real-time video analytics and multimodal interaction. A processor with approximately 9GB of dedicated video memory is sufficient to support long-context and streaming applications.
[0024] The processor 101 reads and executes computer program instructions stored in the memory 102 to perform image processing in the following embodiments.
[0025] In one possible implementation, the electronic device may further include an input interface 103 electrically connected to the processor 101 for receiving raw data from the outside; an output interface 104 electrically connected to the processor 101 for outputting processing results; and a communication bus 105 through which the processor 101, memory 102, input interface 103 and output interface 104 communicate with each other.
[0026] The processor 101 is configured to perform the steps described in the "Exemplary Method" section below by executing instructions stored in the memory 102. Specifically: the input interface 103 receives input data and transmits it to the processor 101 via a bus; the processor 101 executes instructions to perform corresponding processing; and the processor 101 sends output data to a downstream device via the output interface 104.
[0027] like Figure 1 As shown, in some embodiments, the aforementioned exemplary system can be implemented, in whole or in part, in a single electronic device. For example, this electronic device may be an in-vehicle domain controller, an in-vehicle computing chip, a robot main control unit, or a smart camera processing board.
[0028] As an example, in the field of intelligent driving, the processor is a system-on-a-chip (SoC) within the vehicle's domain controller, and the memory includes RAM and an onboard solid-state drive (SSD). After power-on, the SoC loads computer program instructions and detection model weight parameters from non-volatile memory into the RAM. Raw image data captured by the forward-looking camera is directly written to the frame buffer in the RAM. The processor reads the raw image data from the RAM and performs object detection, region cropping, and image filling. During this process, the generated region images and zero-background images are temporarily stored in the RAM. The processor retains the filled target image data in the RAM, which is then read and compressed by the hardware encoder. The compressed target image data is written from the RAM to the onboard SSD for persistent storage. Simultaneously, some data is sent from the RAM to the target recognition module via the vehicle's Ethernet. The target recognition module uses the target image data to accurately classify and identify the status of various target objects (e.g., traffic lights, vehicles, pedestrians, traffic signs, etc.). After the target recognition module completes the recognition, it outputs the structured perception results (such as target category, location, speed, traffic light color status, etc.) to the downstream planning and decision-making module, which then performs path planning and vehicle control based on these results.
[0029] As another example, in the field of robotics, the processor is a system-on-a-chip (SoC) within the robot's main control unit, and the memory includes RAM and local non-volatile memory. After the SoC powers on, it loads program instructions and detection model parameters from the non-volatile memory into the RAM. Raw image data captured by an industrial camera is stored in the RAM. The processor reads the raw image data from the RAM and performs target detection, cropping, and filling operations on the physical entities present (such as tables, chairs, and boxes). The resulting region images and zero-background images are temporarily stored in the RAM. After being compressed by a hardware encoder, the target image data is written from the RAM to the local non-volatile memory for temporary storage, and simultaneously uploaded from the RAM to the robot's perception module. The perception and recognition module identifies and determines the state of the physical entities based on the target image data, and the recognition results are input into the robot's decision-making and planning module to drive the robot to perform corresponding operational tasks.
[0030] By embedding the exemplary method in the memory of the electronic device and executing it by the processor, the device achieves the technical effect of reducing the consumption of transmission bandwidth and storage resources while preserving the high dynamic range and color fidelity of the original image of the target area.
[0031] Based on the above exemplary system and exemplary electronic device, such as Figure 1 As shown, the processor executes computer program instructions to perform image processing, such as... Figure 2 As shown, the processor executes computer program instructions to implement steps S11-S14: S11, Obtain the original image.
[0032] As an example, the raw image can refer to the image data output by the image sensor. This raw image can be a Bayer array RAW image directly from the sensor, or it can be an image in another format (e.g., a YUV or RGB image) that has only undergone linear preprocessing such as black level correction or bad pixel removal. The specific image format can be selected according to the application scenario. For example, in applications requiring high image quality, the raw image can be a high-bit-depth RAW format image to fully preserve the light intensity levels and color distinctions of the target area under complex lighting conditions.
[0033] Specifically, one implementation method is to directly acquire the data from the image sensor. In vehicle-mounted or robotic perception systems, the image sensor captures scene light through the optical lens of the camera module. Each photosensitive unit in the sensor array converts the light signal into an electrical signal, which is then output as a digital signal via an analog-to-digital converter. This digital signal undergoes no image signal processing steps such as white balance or color correction; it is directly output in the Bayer array format (a color filter arrangement pattern covering the surface of the image sensor). Each pixel location records only the raw response value of a single color channel (a digital value directly output after quantizing the analog voltage generated by the sensor's photosensitive unit, without any processing). The data bit depth is typically 10 bits or 12 bits. This set of continuous frame data can be received through a mobile industrial processor (a high-speed serial communication protocol used to connect the image sensor module and the main processor or system chip) interface or a standard interface such as a parallel port, and cached in a designated address space in memory. At this point, the image stored in memory is the RAW image, which can be directly retrieved and used in subsequent steps. This method has a short data link, acquires RAW images with strong real-time performance, and completely preserves the original dynamic range of the sensor.
[0034] Specifically, as another implementation method, data can be read from existing storage media or data streams. In practical engineering, for offline algorithm verification, model training, or data refeeding testing, physical cameras are often not directly connected. Instead, pre-captured and saved RAW format files are read from solid-state drives, embedded multimedia cards, or cloud storage. The system can call the file system interface to parse the file header information according to the agreed data encapsulation format to obtain metadata such as image resolution, bit depth, and Bayer arrangement mode. Then, the subsequent pixel binary payload is read into memory to reconstruct a RAW image consistent with the original shooting. Alternatively, RAW data streams from remote sensor nodes can be received in real time via communication buses such as vehicle Ethernet. The receiving end then assembles the RAW image frame by frame into a complete RAW image in memory for use by downstream tasks. This method does not rely on local camera hardware, offers flexible deployment, and is suitable for large-scale data training and simulation testing scenarios.
[0035] S12, detect the target object in the original image and obtain at least one target detection box.
[0036] As an example, a target object can refer to any physical entity that a visual perception task needs to identify, such as traffic lights, vehicles, pedestrians, or buildings. A bounding box can refer to the graphical coordinate range used to define the location of a target object in the original image. The bounding box can be of any shape or form; for example, it can be a rectangular, polygonal, or elliptical box depending on the shape of the target object. It can also be a set of ordered or unordered coordinate points that describe the trajectory of the target object's edges. Contour point sets more closely resemble the actual shape of the target object and are often used in scenarios where high shape accuracy is required, such as area delineation in defect detection of industrial parts.
[0037] As an example, an image detection model can be used to identify target objects in the original image, resulting in at least one target detection box.
[0038] Specifically, as an example, such as Figure 3 As shown, when the processor executes computer program instructions to implement S12, it also implements the following steps S121-S122.
[0039] S121, Identify the target object in the original image using an image detection model to obtain at least one initial detection box.
[0040] As an example, an image detection model can refer to an algorithmic module used to locate target objects in a raw image, such as a convolutional neural network or a Transformer detection model. This model takes a RAW image as input and outputs detection bounding boxes containing coordinates and confidence scores to define the location of the target object within the image. The initial detection boxes can refer to the original rectangular boundary coordinates and size information directly output by the image detection model from the RAW image; these boxes are directly generated from the model's inference results.
[0041] Specifically, as an example, a convolutional neural network model is built for RAW image input. The backbone of this model contains several cascaded convolutional and pooling layers. The convolutional layers are configured with filters of different sizes to extract the local texture and spatial structure features of the target object in the RAW image. During the training phase, using a set of RAW image samples with labeled target bounding boxes, the network weights are iteratively optimized through the backpropagation algorithm, enabling the model to distinguish the target category and location.
[0042] It should be noted that ordinary low dynamic range images such as YUV, RGB, and JPEG can only record a limited range of brightness levels, while RAW images, with their high dynamic range characteristics and higher bit depth, can cover a wider brightness range from extremely dark to extremely bright. Compared to general convolutional neural networks that typically process YUV / RGB multi-channel images, the convolutional neural network model in this embodiment is adapted to single-channel RAW image input. By preserving high-resolution feature layers and a lightweight, depthwise separable convolutional structure, it reduces computational overhead while improving the feature extraction capability for distant, small targets, thus adapting to the characteristics of high dynamic range raw RAW image data.
[0043] Specifically, as another example, a Transformer detection model can be built. The core of this model includes an image block embedding module and a sequence feature encoding module. During the training phase, a set of RAW image samples labeled with the bounding boxes of the intersecting targets is used. The AdamW optimization algorithm is applied to minimize the classification loss and the bounding box regression loss, enabling the model to recognize the target objects.
[0044] It should be noted that general Transformer detection models typically employ 12 or more layers of encoder stacking. The computational complexity of global self-attention increases quadratically with sequence length, making it difficult to meet real-time inference requirements. This embodiment of the Transformer detection model reduces the number of encoder layers and introduces a sliding window self-attention mechanism, restricting global attention computation to a local window while preserving cross-window connections at key feature layers. This reduces computational cost while maintaining the ability to model the context of small targets.
[0045] S122, adjust the size of at least one initial detection box to obtain at least one target detection box.
[0046] Specifically, as an example, after object detection is performed on the original image, the initial detection boxes output by the image detection model may have positioning errors or sizes that do not match the requirements of subsequent encoding. This step adjusts the size of the initial detection boxes to make the adjusted target detection boxes more suitable for subsequent cropping and / or encoding compression processing requirements. The rectangles obtained after size adjustment are the target detection boxes. The specific adjustment method is described below and will not be repeated here.
[0047] In the process of obtaining target detection boxes, the electronic device of this disclosure first acquires initial detection boxes through an image detection model, then actively adjusts the size of these initial detection boxes, and finally uses the adjusted results as the final target detection boxes for cropping. This makes the cropping range not completely dependent on the original output of the model, but can be further optimized according to the engineering requirements of subsequent encoding or acquisition. While retaining the model's detection capabilities, it improves the adaptability of the cropping region to actual application scenarios.
[0048] S13, crop the original image based on at least one target detection box to obtain the region image corresponding to the target object.
[0049] As an example, a region image can refer to a local pixel data block containing the target object, extracted from the complete original image according to the boundaries of the target detection box, and its content is the original sensor value.
[0050] As an example, cropping an original image based on at least one object detection box includes: in response to the presence of multiple object detection boxes, determining a coverage pattern that covers the multiple object detection boxes based on the multiple object detection boxes, and cropping the original image based on the coverage pattern to obtain a region image including multiple target objects; or, determining a local pattern of each object detection box based on each object detection box, and cropping the original image based on each local pattern to obtain a region image including multiple target objects.
[0051] As one implementation of this disclosure, such as Figure 4 As shown, when the processor executes computer program instructions to implement S13, it also implements S13A1-S13A2.
[0052] S13A1, determines the coverage pattern that covers multiple target detection boxes.
[0053] As an example, an overlay graph can refer to a geometry that completely encompasses all object detection boxes in the RAW image. For instance, it could be the smallest bounding rectangle of all detection boxes, or the smallest convex polygon that can enclose all the vertices of the detection boxes. This graph defines the closed area for subsequent one-time cropping.
[0054] Specifically, as an example, after obtaining all object detection boxes in the current RAW image, the minimum bounding rectangle is used as the coverage shape. In practice, the vertex coordinates of all object detection boxes are traversed, and the minimum and maximum values of the x-coordinate and y-coordinate are found for each vertex. These four boundary values form an axis-aligned rectangle with its top-left corner at (minimum x-coordinate, minimum y-coordinate), its width equal to the difference between the maximum and minimum x-coordinate, and its height equal to the difference between the maximum and minimum y-coordinate. This rectangle is the minimum bounding rectangle that covers all object detection boxes.
[0055] Specifically, as another example, when the bounding boxes are irregularly distributed in space and using a bounding rectangle would introduce too much irrelevant background, a minimum convex polygon can be used as the covering shape. In practice, all vertices of all bounding boxes are first extracted to form a point set. Then, convex hull calculation is performed on this point set, for example using the Graham scan method or the Jarvis step method, to find the minimum convex polygon vertex sequence that can contain all points in the point set, which serves as the minimum convex polygon covering shape.
[0056] S13A2, based on the overlay pattern, crops the original image to obtain the region image corresponding to the target object.
[0057] Specifically, as an example, for the smallest bounding rectangle, the pixel data within the calculated width and height range can be copied from the RAW image, starting from the top left corner of the rectangle, to generate an image of the region corresponding to the overlay graphic.
[0058] For example, consider two traffic light detection boxes in the image. Box A has coordinates ranging from (1000, 300) to (1060, 360), while box B has coordinates ranging from (1020, 400) to (1080, 460). After traversing the vertices of both boxes, the minimum x-coordinate is 1000 and the maximum is 1080, while the minimum y-coordinate is 300 and the maximum is 460. Therefore, the top-left corner of the minimum bounding rectangle is determined to be (1000, 300), with a width of 80 and a height of 160. The cropped area completely contains the RAW pixels of both traffic lights and the area between them.
[0059] Specifically, as another example, for a minimum convex polygon overlay pattern, a mask can be constructed that only covers the internal region of the convex polygon. Pixels are then extracted from the RAW image using this mask as the cropping boundary. In the cropping result, pixels outside the convex polygon are set to zero or null values, retaining only the original sensor data of the pixels inside the convex polygon.
[0060] For example, the three detection boxes in the image are distributed in a triangle. A hexagonal vertex sequence is obtained by convex hull calculation. The cropped area image only retains RAW pixels within the hexagonal range, which reduces the introduction of invalid background and ensures that the target objects corresponding to the three detection boxes are completely covered.
[0061] Specifically, such as Figure 5 As shown, as another implementation of S13, the processor implements the following steps S13B1-S13B2 when implementing S13.
[0062] S13B1, based on each target detection box, the original image is cropped to obtain the local image corresponding to each target detection box.
[0063] When multiple object detection boxes exist, the independence of each box can be maintained, and cropping can be performed on each box separately to generate multiple independent region images. Specifically, for each object detection box, based on its coordinate information (e.g., top-left corner coordinates and width / height information), the starting offset and row stride of the box at the starting address of the RAW image data can be calculated. The original pixel values at the corresponding positions within the object detection box are copied row by row and stored as a separate data block.
[0064] S13B2 treats each local image as a region image.
[0065] Specifically, as an example, after obtaining the various local images through the above cropping steps, no further merging or stitching processing is performed on these local images. Instead, each local image is directly treated as an independent region image; that is, each data block is a region image corresponding to a target object. This means that the number of region images corresponds one-to-one with the number of target detection boxes, and each region image contains the original pixel data corresponding to a target object.
[0066] The electronic device of this disclosure performs a cropping operation on each target detection box independently. The local images corresponding to each target object are extracted separately and directly used as region images, ensuring that the original pixel data of each target region are independent and do not interfere with each other. This method reduces redundant background pixels that may be introduced when merging multiple targets into a single overlay image. Each region image contains only a single target and its necessary neighborhood information, and the region boundaries more closely match the actual size of the target. In subsequent processing, each region image can be operated independently, facilitating differentiated adjustments for different target categories or sizes.
[0067] S14, determine the target image based on the region image and the zero background image.
[0068] As an example, a zero-background image can refer to a blank image where all pixels are initially set to zero. This image can serve as a base for the target image, and its size can be flexibly set according to subsequent encoding requirements and transmission bandwidth constraints. The target image can refer to the image generated by filling a region image with a zero-background image, where only the filled regions contain the original data of the target object, while the remaining regions have zero values.
[0069] Specifically, as an example, a zero-background image can be created first. The size of this zero-background image can be set according to actual needs. For example, based on the size and relative position of all regions to be filled, a minimum bounding rectangle that can just accommodate all regions is calculated, and this is used as the preset size of the zero-background image. Then, the pixel values of the regions within the minimum bounding rectangle are filled into the zero-background image, thus obtaining the target image. This target image only contains a portion of the invalid coded blocks within the minimum bounding rectangle, further reducing the number of zero values that the subsequent encoder needs to process.
[0070] The image processing method of this disclosure, by directly acquiring the original image, retains the high dynamic range and original color information of the sensor output, enabling the complete preservation of the brightness levels and color differences of the target area under complex lighting conditions, thus providing a reliable information foundation for accurate perception. Based on this, the target object is detected from the original image to obtain a target detection box. Then, based on the detection box, a region image covering only the target object is cropped, ensuring that the data to be processed and transmitted subsequently focuses on the effective pixels carrying the perception task, rather than all the data of the entire image. Subsequently, the region image is synthesized with a zero-background image to determine the target image, so that only the location of the target object in the target image is filled with the true original pixel value, while the remaining large areas of background are zero values. These zero-value regions can be efficiently compressed to a very small data volume during encoding compression, while the original information of the target area is not lost. This effectively reduces the data transmission bandwidth and storage space usage while ensuring the quality of the image information required for perception.
[0071] As another implementation of this disclosure, such as Figure 6 As shown, when the processor executes computer program instructions to implement S14, it is also used to implement the following steps S141-S142.
[0072] S141, construct a zero-background image with the same dimensions as the original image.
[0073] As an example, a zero-background image refers to an image where all pixel values are initialized to zero. The width and height of this image are exactly the same as the pixel resolution of the original image. Using the same size as the original image allows the input resolution to remain constant during subsequent calls to the hardware encoder, eliminating the initialization overhead of reconfiguring encoding parameters and the block grid for each different image size received.
[0074] Specifically, as an example, the width and height values of the original image are obtained. Based on these two parameters, a contiguous storage space of the corresponding capacity is allocated in memory, and all bytes in this storage space are uniformly set to zero, thus completing the construction of an all-zero background image.
[0075] It's important to note that setting the zero-background image to the same size as the original image has clear engineering significance in practical deployments. For hardware encoders, their encoding pipelines typically operate continuously in a fixed-resolution mode, where each input frame is expected to have the same width and height. If the size of the target image generated after each padding changes frequently, the encoder needs to repeatedly exit the current encoding session, reallocate internal buffers, and re-divide the encoding block grid according to the new size. This frequent mode switching introduces additional control latency and computational resource consumption. Maintaining the same fixed size as the original image ensures that regardless of how many targets are detected or how many regions are cropped in the current frame, the encoder always receives an image with a constant resolution, allowing for continuous and stable operation and reducing the overall control complexity of the encoding process.
[0076] S142, fill the region image with a background image of zero to obtain the target image.
[0077] Specifically, as an example, all region images can be iterated over and placed sequentially on the zero-background image according to the order in which they were output in S32. The first region image is placed at the origin of the zero-background image. Each subsequent region image is placed immediately adjacent to the right boundary of the previously placed region image. When the remaining width of a row is insufficient to accommodate the next region image, it automatically wraps to the left boundary of the next row and continues placement. When each region image is filled, its pixel values are copied row by row into the currently allocated range in the zero-background image. After all region images have been filled, the zero-background image is transformed into the target image.
[0078] The electronic device of this disclosure constructs a zero-background image with the same size as the original image, keeping the image resolution fed into the encoder constant. In a continuous acquisition and encoding workflow, the encoder does not need to repeatedly exit the current encoding session, reallocate internal buffers, and re-divide the encoding block grid due to frequent changes in the target image size. This reduces the initialization overhead and control latency introduced by mode switching, allowing the encoding pipeline to maintain a steady-state operating mode. Based on this, the target image is obtained by filling the zero-background image with the region image. Only the positions of the region image are filled in the target image to carry effective pixel data, while the remaining large areas remain zero. The zero-value regions are efficiently compressed into a very small bitstream during encoding compression, so that the final output target image data occupies only a low transmission bandwidth and storage space while retaining the original information of the target region. This maintains the stability and low latency of the encoder operation while effectively compressing the total amount of data.
[0079] Furthermore, when implementing S142, the processor can also specifically implement the following steps: based on the position of the region image in the original image, fill the corresponding position of the zero-background image with the same size as the original image to obtain the target image.
[0080] Specifically, as an example, all regions to be filled can be iterated through. For each region image, its associated original spatial coordinates are read, i.e., the coordinates of the target detection box when the region image was cropped in the original image, including the x and y coordinates of the top-left corner of the box, as well as the width and height of the box. Based on these coordinates, a corresponding write range is determined in the all-zero background image, whose starting row, column, and size are exactly the same as the original coordinates. Then, the pixel data of the region image is written row by row into the write range in the all-zero background image. During the writing process, the pixel arrangement order within the region image remains unchanged, and the original value of each pixel is copied to the corresponding position, overwriting the zero value there.
[0081] For example, in a RAW image containing two traffic light targets, two region images are obtained after detection and cropping. Region image A has coordinates (1000, 300) in the original image, a width of 72 pixels, and a height of 60 pixels; region image B has coordinates (2200, 450), a width of 80 pixels, and a height of 64 pixels. The rectangular region from row 300 to row 359 and column 1000 to column 1071 in the zero-background image is covered with the pixel data of region image A. Similarly, the rectangular region from row 450 to row 513 and column 2200 to column 2279 is covered with the pixel data of region image B.
[0082] After padding, the zero-background image is transformed into the target image. In this target image, region image A is located in the upper left region, and region image B is located in the middle right region. The horizontal and vertical spacing between them perfectly matches their relative positions in the original image. Therefore, the target image not only carries the high-fidelity original pixel information of each of the two traffic light targets, but also implicitly preserves the spatial topological relationship between them, that is, the layout information of one light group in the left front and one light group in the right front of the image. This spatial relationship is crucial for downstream perception tasks to understand the relative orientation of multiple target objects in the physical environment. At the same time, all regions in the target image except for these two rectangular blocks are zero values, which are efficiently processed during encoding and compression, ensuring that the output bitstream still maintains a small data volume.
[0083] The electronic device of this disclosure fills a region image with a zero-value background image of the same size as the original image according to the original spatial coordinates. This ensures that the relative orientations and distances between multiple target objects in the target image remain consistent with their layout in the original image, providing a structured input for the downstream perception module to understand the spatial topological relationships between targets. Simultaneously, the zero-value background image size is consistent with the original image, allowing the encoder to operate continuously at a fixed resolution and avoiding repeated initialization due to changes in input size. The large area of zero-value background outside the target region is efficiently compressed during encoding, effectively controlling the output data volume while preserving the target spatial relationships and original information.
[0084] As another implementation of this disclosure, such as Figure 7 As shown, when the processor executes the computer program instructions to implement S122, it is also used to implement the following step S1221.
[0085] S1221, expand at least one initial detection box to obtain at least one target detection box.
[0086] As an example, expansion processing can refer to the operation of extending the boundaries of the initial detection box outward to increase its coverage. Its purpose is to leave a safety margin for subsequent clipping and increase tolerance to deviations in target position.
[0087] As an example, expanding at least one initial detection box includes: expanding the initial detection box based on the confidence level of the initial detection box; or, expanding the boundary of the initial detection box according to a preset expansion parameter; wherein the expansion parameter can be a preset expansion ratio value or a preset fixed expansion pixel value.
[0088] Specifically, as an example, a mapping relationship between different target object categories and expansion parameters can be established in advance. After obtaining the initial detection box and its corresponding category label, this mapping relationship is queried to extract the preset expansion ratio or fixed expansion pixel value corresponding to that target category. For example, for traffic light targets, since their detection boxes are relatively compact and their size is relatively fixed, a higher expansion ratio can be assigned to fully expand the acquisition range; for vehicle targets, since the initial detection box has already completely enveloped the target body, a lower expansion ratio or no expansion can be assigned.
[0089] For each initial detection box, boundary expansion is performed according to the expansion parameters corresponding to its category, and the adjusted coordinates are calculated. All the expanded bounding boxes differentiated by category are the target detection boxes. This ensures that the decision on the expansion magnitude matches the actual shape and detection features of the target object, reducing the problem of over-expansion or under-expansion that may occur when a single expansion strategy is used to handle different types of targets.
[0090] The electronic device of this disclosure specifically employs an expansion process to adjust the size of the initial detection box, ensuring that the coverage area of each target detection box is larger than the initial range of the original model output. The expanded detection box can include more neighboring pixels around the target object during cropping. Even if there is a slight deviation between the initial detection box and the true boundary of the target object, the expanded box still has a high probability of completely encompassing the target object, thereby reducing the risk of target edges being cut off due to detection box positioning deviations.
[0091] As another implementation of this disclosure, such as Figure 8 As shown, when the processor executes the computer program instructions to implement S1221, it is also used to implement the following steps S1221A1-S1221A2.
[0092] S1221A1 identifies target objects in the original image using an image detection model and obtains the confidence scores corresponding to each initial detection box.
[0093] As an example, an image detection model refers to an algorithmic module used to locate target objects in a raw image, such as a convolutional neural network or a Transformer detection network. This model receives the raw image as input, performs forward inference, and outputs a set of candidate detection results, each containing the target's location information and a confidence score. The initial detection box refers to the original rectangular boundary directly output by the image detection model, without subsequent filtering or adjustment. The confidence score can refer to a numerical value simultaneously provided by the image detection model when outputting each initial detection box, indicating the probability that a real target object exists within that initial detection box; the higher the confidence score, the higher the probability that a real target object exists within that initial detection box.
[0094] Specifically, as an example, the original image is fed as an input tensor into a pre-trained convolutional neural network detection model. This model extracts multi-scale feature maps of the original image through its backbone network, predicting multiple candidate detection boxes and the corresponding target class probability for each candidate box at each spatial location on the feature maps. The highest predicted target class probability is used as the confidence score of that candidate box. After one forward propagation, the model outputs the coordinates of all candidate detection boxes in that frame of the original image and the confidence score for each box. These candidate boxes and their confidence scores constitute the initial detection boxes and their corresponding confidence scores.
[0095] S1221A2, the initial detection box with a confidence level greater than the first preset confidence level is determined as the target detection box, wherein the first preset confidence level is less than the second preset confidence level, and the second preset confidence level is determined based on the detection accuracy of the image detection model.
[0096] As an example, the first pre-set reliability can refer to a low confidence threshold set during the data acquisition phase. All initial detection boxes with a confidence level higher than this threshold are accepted to retain candidate boxes that the image detection model is not entirely certain about but may actually be true, thereby improving the target recall rate. The second pre-set reliability can refer to a higher, conventional confidence threshold determined based on the image detection model's evaluation results on a standard test set, balancing precision and recall. This value is the default threshold used in routine inference for non-data acquisition tasks to ensure the reliability of the detection results. Detection accuracy can refer to a comprehensive measure of the image detection model's performance, typically evaluated on a validation set using metrics such as precision and recall. It reflects the correctness with which the model determines the existence of the target object and is the primary basis for determining the second pre-set reliability.
[0097] Specifically, as an example, two sets of confidence parameters can be pre-configured. For instance, a first pre-set confidence level is set to 0.3, dedicated to the data acquisition mode; a second pre-set confidence level is set to 0.7, dedicated to the regular perceptual inference mode. When performing the data acquisition task, the output filtering threshold of the image detection model is set to the first pre-set confidence level of 0.3. The image detection model performs forward inference on the original image, outputting a confidence value between 0 and 1 for each detected candidate box. When traversing all candidate boxes, its confidence value is read and compared with the currently effective first pre-set confidence level of 0.3. If the confidence level is greater than 0.3, the candidate box is marked as valid and retained as an initial detection box; otherwise, it is discarded.
[0098] The electronic device of this disclosure, by setting a first preset confidence level lower than the threshold corresponding to conventional detection accuracy, allows candidate boxes whose confidence level output by the model does not reach the conventional confirmation standard but still has a certain probability of existence to be retained. These low-confidence boxes, which would be discarded in the conventional process, can participate in subsequent cropping and encoding in the acquisition scenario, improving the probability of capturing targets under difficult conditions such as long distance, small size, or low light, and reducing missed detections.
[0099] As another implementation of this disclosure, such as Figure 9 As shown, when the processor executes computer program instructions to implement S51, it is also used to implement the following step S1221B1.
[0100] S1221B1, perform boundary expansion processing on the initial detection box to obtain the target detection box; wherein, the boundary expansion processing includes expanding the width and height of the detection box according to a preset ratio, and / or expanding the boundary of the detection box by a fixed pixel value.
[0101] As an example, boundary expansion processing refers to pushing the rectangular boundaries of the initial detection box outwards to increase coverage. Its purpose is to provide margin for error in subsequent cropping, compensating for positioning deviations in the initial detection box or minor displacements of the target between frames. Expanding by a preset percentage means multiplying the current width and height of the initial detection box by a set percentage to calculate the number of pixels expanded in each direction. For example, expanding by 20% will increase the width and height accordingly, with the expansion amount proportional to the target box size. Expanding by a fixed pixel value means uniformly pushing a fixed number of pixels outwards on all four boundaries of the initial detection box, such as expanding each by 8 pixels. The absolute number of pixels expanded remains consistent across all target boxes and does not change with the target size.
[0102] Specifically, as an example, a fixed expansion ratio applicable to all target categories can be preset, such as 20%. During execution, each initial detection box retained after confidence filtering is iterated over, and its top-left corner coordinates and width and height values are obtained. For each initial detection box, the expanded width is calculated as the original width multiplied by 1.2, and the expanded height is calculated as the original height multiplied by 1.2. The increased width and height are evenly distributed to the left and right sides and top and bottom sides, i.e., expanding the original width by 10% on the left and right sides and expanding the original height by 10% on the top and bottom sides. Based on the expanded width and height, the top-left corner coordinates are recalculated to obtain the expanded rectangular boundary. All rectangles expanded according to this ratio are the target detection boxes.
[0103] For example, such as Figure 6 As shown, the initial detection bounding box 502 of a RAW image 501 has top-left corner coordinates of (500, 400), a width of 100 pixels, and a height of 50 pixels. After being expanded by 20%, the expanded image 503 is obtained, with a width of 120 pixels and a height of 60 pixels. The top-left corner coordinates are adjusted accordingly to (490, 395). The expanded box has an additional 10% margin in each of the four directions.
[0104] Furthermore, as an example, the expansion ratio can be determined by statistically analyzing multiple samples. Specifically, an original image sample set containing labeled ground truth bounding boxes can be constructed, covering different target distances, lighting conditions, and angles. An image detection model is used to infer from the sample set, outputting a predicted detection box for each target object, and matching it with the corresponding ground truth bounding box according to a rule that the intersection-union ratio (IU) is greater than a preset threshold. For each successfully matched pair of predicted and ground truth bounding boxes, the ratio of the width of the predicted detection box to the width of the ground truth bounding box, and the ratio of the height of the predicted box to the height of the ground truth bounding box, are calculated. The width and height ratios of all matching pairs are collected, and their distributions are statistically analyzed. Based on these distributions, the expansion ratios for width and height are determined.
[0105] Specifically, as another example, different fixed pixel expansion values can be pre-set for the four expansion directions to accommodate the potential inconsistency in the detection deviation of the target object in the horizontal and vertical directions. In practice, the horizontal outer expansion value is set to 10 pixels, the horizontal inner expansion value to 6 pixels, the vertical outer expansion value to 12 pixels, and the vertical inner expansion value to 8 pixels. Each initial detection box can be traversed, and its center point position within the entire image can be read. Based on the relative position of the center point to the image's central axis, the spatial orientation of these four expansion directions is determined. For example, for the target box on the left side of the image, the direction towards the image center is considered inner, and the direction away from the image center is considered outer. According to the preset parameters, the corresponding number of pixels are advanced in each of the four directions to obtain the expanded graphic boundaries. All the graphic boxes processed in this way are the target detection boxes.
[0106] Further, as an example, the extended value can be determined based on the regression bias statistics. Specifically, an original image sample set of labeled ground truth bounding boxes of target objects can be obtained. The image detection model is used to infer from the sample set to obtain the predicted detection box for each target object. The predicted detection box is matched with the corresponding ground truth bounding box, and successful matches are used for bias statistics. For each pair of matched boxes, the coordinate bias of the four boundaries is calculated. The left boundary bias is obtained by subtracting the left boundary x-coordinate of the predicted box from the left boundary x-coordinate of the ground truth box; similarly, the right boundary bias, top boundary bias, and bottom boundary bias are obtained. Positive values indicate that the predicted box boundary is offset outward relative to the ground truth box, while negative values indicate inward contraction. The bias data of all paired boxes in the four directions are collected, and the bias distribution in each direction is plotted. Based on the bias distribution, an extended reference value is determined for each direction. This extended reference value should be able to cover the boundary error of the predicted detection box relative to the ground truth bounding box in most cases.
[0107] The electronic device of this disclosure expands the detection boxes by a preset ratio to provide a safety margin proportional to their size, while expanding by a fixed pixel value provides uniform boundary compensation for all detection boxes. The use of these two methods, individually or in combination, enables the expanded target detection boxes to effectively compensate for positioning errors in the model output and minute displacements of the target at the moment of acquisition, thus improving the robustness of the clipping range.
[0108] As another implementation of this disclosure, such as Figure 10 As shown, the processor executes computer program instructions to further implement the following step S15.
[0109] S15, the target image is encoded and compressed to obtain the target image data.
[0110] As an example, encoding compression can refer to the process of converting a target image into a bitstream that takes up less data using an encoder (such as a JPEG hardware encoder). The target image data can refer to the storable or transmittable binary bit sequence output after encoding compression, i.e., the compressed bitstream.
[0111] Specifically, as an example, a hardware encoder can be invoked to encode and compress the target image. In automotive computing platforms or embedded vision systems, system chips typically integrate hardware JPEG encoders or video encoders. During operation, information such as the target image's starting address in memory, image width, image height, and input pixel format can be configured into the hardware encoder's input register via a driver program. Encoding parameters such as the target bitrate or compression quality factor are then configured, and a hardware encoding session is initiated. The hardware encoder reads the target image's pixel data in a pipelined manner, performing discrete cosine transform, quantization, and entropy encoding sequentially in preset pixel blocks, ultimately outputting a continuous binary bit sequence. This bit sequence is the target image data, which can be directly written to a specified output buffer via direct memory access for subsequent transmission or storage.
[0112] In this target image, large areas are zero-value backgrounds. When the encoder processes these all-zero coding blocks, all coefficients become zero after transformation and quantization, and they are efficiently compressed into an extremely short bitstream during the run-length encoding stage, contributing very little to the overall output data volume. The real data overhead comes from the coding blocks occupied by each region of the image. Since the encoder only needs to be called once to complete the compression of the entire image, the overhead of multiple initializations and session management caused by encoding each region of the image independently is avoided, making the encoding process more streamlined.
[0113] Furthermore, as an example, in various visual sensing devices, such as smart cameras, mobile robots, or drones, after encoding and compressing target image data, this data can be temporarily stored in the device's local flash memory or removable memory card, awaiting transmission conditions. When the device connects to the management platform via a wireless network or wired interface, the accumulated target image data is uploaded in batches to a cloud-based object storage service or a private file server. The background analysis program can retrieve this data from storage as needed, decode and restore RAW information containing only the target area, and use it for model iteration training, data annotation verification, or event backtracking analysis. This method transmits only the valid target area, and the background zero values are efficiently compressed, significantly reducing upload bandwidth consumption and remote storage usage while preserving the original information.
[0114] Furthermore, as an example, during encoding and compression, metadata for the target image of that frame is generated synchronously, recording the top-left corner coordinates and width and height information of each region image within the target image. This metadata is written as a structured field into the header of the encoded bitstream or a custom data segment. After receiving the target image data, the decoding end first parses the metadata from the bitstream header to obtain the spatial location and extent of all region images. Then, it performs decoding on the target image, temporarily storing the decoded complete image in memory. Based on the coordinates and dimensions of each region recorded in the metadata, it directly locates and extracts pixel blocks from the decoded image, outputting these pixel blocks to the downstream perception module for analysis. The remaining regions in the decoded image that are not labeled by the metadata are directly discarded and do not participate in subsequent calculations. This implementation allows the decoding end to quickly obtain effective regions without traversing the entire image, and it completes the explicit transfer of spatial location information from the encoding end to the decoding end at the bitstream level.
[0115] In the electronic device of this disclosure, only the areas occupied by the regional images in the target image carry valid pixel data, while the remaining large areas of the background are all zero values. When the encoder processes the all-zero coding block, all coefficients are zero after transformation and quantization, and it is efficiently compressed into an extremely short bitstream during the encoding stage, with minimal impact on the final data volume. The coding block containing the regional images that actually generate data overhead can completely retain the original information of the target object after encoding. At the same time, the entire target image only needs to be compressed by calling the encoder once, avoiding the multiple initialization overhead caused by encoding multiple regional images independently, making the encoding process more streamlined. Thus, while retaining high-fidelity original data of the target area, the total amount of output bitstream is effectively controlled, reducing the occupation of transmission bandwidth and storage resources.
[0116] As another implementation of this disclosure, such as Figure 11 As shown, when the processor executes the computer program instructions to implement S122, it is also used to implement the following steps 1222.
[0117] S1222, Align at least one initial detection box according to the preset encoding unit size used during encoding compression to obtain at least one target detection box.
[0118] As an example, a preset coding unit can refer to the basic processing block that the encoder relies on when performing compression on an image. For example, a JPEG encoder uses 8×8 pixel blocks as a unit, performing transformation and quantization operations block by block; each 8×8 region in the image is a coding unit. The preset coding unit size can refer to the pixel values corresponding to the width and height of the coding unit. For example, the preset coding unit size in the JPEG standard is 8 pixels wide and 8 pixels high. Alignment processing can refer to adjusting the boundary coordinates of the initial detection box so that its width and height become integer multiples of the preset coding unit size. This adjustment only changes the outer boundary of the region, without modifying the original RAW values of the pixels within the region, aiming to eliminate unnecessary padding overhead caused by size mismatch in the encoder.
[0119] Specifically, as an example, the original image is logically divided into grids. Using the top-left pixel coordinates of the original image as the origin, the image is divided into multiple non-overlapping rectangular grid units with a step size equal to the width of a preset coding unit in the horizontal direction and a step size equal to the height of that preset coding unit in the vertical direction. Each grid unit has the same size as the preset coding unit. Then, the grid set covered by the initial detection box is determined. The top-left and bottom-right coordinates of the initial detection box are read, and the horizontal and vertical coordinates of the top-left corner are rounded down to the nearest grid boundary to obtain the position of the starting grid. The horizontal and vertical coordinates of the bottom-right corner are rounded up to the nearest grid boundary to obtain the position of the ending grid. All grid units within the rectangular area defined by the starting and ending grids constitute the grid set occupied by the initial detection box. Finally, the rectangular area defined by the outer boundary of the above grid set is taken as the target detection box. The boundary of this target detection box completely coincides with the grid boundary, and its width and height are integer multiples of the preset coding unit size.
[0120] The electronic device of this disclosure further introduces alignment processing based on a preset coding unit size during the initial detection box size adjustment process. This means that regardless of previous confidence filtering or boundary expansion, the final target detection box boundary used for cropping will be fine-tuned to an integer multiple of the encoder's basic processing unit. This alignment ensures that the boundaries of each region image block in the subsequently filled target image coincide with the block grid inside the encoder, avoiding internal padding due to size mismatch by the encoder, thereby improving the processing efficiency of encoding compression and the adaptability to the encoder.
[0121] As another implementation of this disclosure, such as Figure 12 As shown, when the processor executes the computer program instructions to implement S1222, it is also used to implement the following step S12221.
[0122] S12221, Adjust the boundary of the region corresponding to the initial detection box so that the width and height of the region corresponding to the adjusted initial detection box are both integer multiples of the preset encoding unit size used during encoding compression.
[0123] As an example, the region corresponding to the initial detection box can refer to the local pixel range defined by the graphic boundary of the initial detection box in the original image. This region is the original input object for subsequent alignment operations, containing the raw sensor data of the target object and its adjacent background. The boundary of the region can refer to the spatial boundaries of the graphic region in the original image; for example, the left and right boundaries determine the width range, and the top and bottom boundaries determine the height range. Adjusting the boundary of the region can refer to fine-tuning the boundary coordinates of the rectangular region so that its width and height values become integer multiples of the target values. This operation only changes the boundary position and does not modify the pixel values within the region.
[0124] Specifically, as an example, let's take a preset encoding unit size of 8×8 pixels. For each initial detection box, first extract its width and height values. Divide the width and height values by 8 respectively. If they are evenly divisible, no adjustment is needed; if not, round the width and height up to the next integer multiple of 8. Calculate the amount of pixels to be expanded outwards. The width is calculated as 8 minus the remainder modulo 8, and the height is calculated as 8 minus the remainder modulo 8. Divide the expansion amount into left and right halves and top and bottom halves, and adjust the coordinates of the left, right, top, and bottom boundaries of the initial detection box accordingly. This implementation method always expands the region boundaries without cropping pixels within the original detection range, ensuring the integrity of the target information.
[0125] For example, if an initial detection box is 73 pixels wide, and 73 divided by 8 leaves a remainder of 1, it needs to be expanded outwards by 7 pixels to make the width 80. Expand the left and right sides by 3.5 pixels each, which can be rounded up to 3 pixels for the left edge and 4 pixels for the right edge. Adjust the height in the same way. If the adjusted area's width is 80 and its height is also an integer multiple of 8, then the alignment process is complete.
[0126] Specifically, as another embodiment, when there are higher requirements for the compactness of the region size in the project, a rounding method can be used. For each initial detection box, the width value is divided by 8. If it is not divisible by an integer, it is rounded down to an integer multiple of the previous 8. The number of pixels that need to be shrunk in width is the remainder of the width modulo 8, and half of this remainder is shrunk on both sides. The region image produced by this implementation is slightly smaller than the original target detection range, but in some processes where sufficient margin has been provided by the edge expansion step, the slight boundary shrinkage will not destroy the integrity of the core target pixels, while obtaining a more compact region and reducing the amount of redundant data in the encoding process.
[0127] For example, the initial detection box width is 106 pixels. 106 divided by 8 leaves a remainder of 2, which is rounded down to 104. The left and right sides are shrunk by 1 pixel. The height direction is handled similarly. The adjusted region boundaries fall exactly as a multiple of 8, while maintaining minimal loss of valid data within the region.
[0128] The image processing method of this disclosure involves aligning the regions by directly adjusting the boundary coordinates of the corresponding areas of the initial detection box, ensuring that the adjusted width and height are integer multiples of the coding unit size. This boundary adjustment only changes the outer geometry of the region and does not modify the pixel values within the region. Therefore, the aligned region still retains the original RAW data. This ensures that while achieving the coding efficiency improvement brought about by size normalization, the integrity and authenticity of the original information of the target region remain unaffected.
[0129] Exemplary methods Based on the electronic device provided in the above embodiments, this disclosure also provides specific implementations of the image processing method. Please refer to the following embodiments.
[0130] Obtain the original image; Detect target objects in the original image and obtain at least one target detection box; The original image is cropped based on at least one target detection box to obtain the region image corresponding to the target object; The target image is determined based on the region image and the zero-background image.
[0131] The image processing method of this disclosure, by directly acquiring the original image, retains the high dynamic range and original color information of the sensor output, enabling the complete preservation of the brightness levels and color differences of the target area under complex lighting conditions, thus providing a reliable information foundation for accurate perception. Based on this, the target object is detected from the original image to obtain a target detection box. Then, based on the detection box, a region image covering only the target object is cropped, ensuring that the data to be processed and transmitted subsequently focuses on the effective pixels carrying the perception task, rather than all the data of the entire image. Subsequently, the region image is synthesized with a zero-background image to determine the target image, so that only the location of the target object in the target image is filled with the true original pixel value, while the remaining large areas of background are zero values. These zero-value regions can be efficiently compressed to a very small data volume during encoding compression, while the original information of the target area is not lost. This effectively reduces the data transmission bandwidth and storage space usage while ensuring the quality of the image information required for perception.
[0132] As another implementation of this disclosure, the target image is determined based on the region image and the zero-background image, including: Construct an image with zero background and the same dimensions as the original image; The target image is obtained by filling the region image with a background image of zero.
[0133] As another implementation of this disclosure, the target image is obtained by filling the region image with a background image of zero, including: Based on the position of the region image in the original image, the region image is filled into the corresponding position of a zero-background image of the same size as the original image to obtain the target image.
[0134] As another implementation of this disclosure, in response to the presence of multiple target detection boxes, the original image is cropped based on at least one target detection box to obtain a region image corresponding to the target object, including: Determine a coverage pattern that covers multiple target detection boxes; crop the original image based on the coverage pattern to obtain the region image corresponding to the target object; or... The original image is cropped based on each target detection box to obtain the local image corresponding to each target detection box; each local image is used as a region image.
[0135] As another implementation of this disclosure, detecting a target object in the original image to obtain at least one target detection box includes: The target object is identified in the original image using an image detection model, and at least one initial detection box is obtained. The size of at least one initial detection box is adjusted to obtain at least one target detection box.
[0136] As another implementation of this disclosure, the size of at least one initial detection box is adjusted to obtain at least one target detection box, including: At least one initial detection box is expanded to obtain at least one target detection box.
[0137] As another implementation of this disclosure, at least one initial detection box is expanded to obtain at least one target detection box, including: The image detection model identifies target objects in the original image, and the confidence scores of each initial detection box are obtained. Initial detection boxes with confidence scores greater than a first preset confidence score are identified as target detection boxes; or... The initial detection box is expanded to obtain the target detection box; The boundary expansion process includes expanding the width and height of the detection box by a preset ratio, and / or expanding the boundary of the detection box by a fixed pixel value.
[0138] As another implementation of this disclosure, the method further includes: The target image is encoded and compressed to obtain the target image data.
[0139] As another implementation of this disclosure, at least one initial detection box is expanded to obtain at least one target detection box, including: At least one initial detection box is aligned according to the preset encoding unit size used during encoding compression to obtain at least one target detection box.
[0140] As another implementation of this disclosure, at least one initial detection box is aligned according to the preset encoding unit size used during encoding compression to obtain at least one target detection box, including: Adjust the boundaries of the region corresponding to the initial detection box so that the width and height of the region corresponding to the adjusted initial detection box are integer multiples of the preset encoding unit size used during encoding compression.
[0141] Exemplary computer program products and computer-readable storage media In addition to the methods and apparatus described above, embodiments of this disclosure may also provide a computer program product, including computer program instructions that, when executed by a processor, cause the processor to perform the steps of the image processing methods of the various embodiments of this disclosure described in the "Exemplary Methods" section above.
[0142] Computer program products can be written in any combination of one or more programming languages to perform the operations of embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0143] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps of the image processing methods of the various embodiments of this disclosure described in the "Exemplary Methods" section above.
[0144] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may include, but is not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0145] The basic principles of this disclosure have been described above with reference to specific embodiments. However, the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0146] Various modifications and variations can be made to this disclosure without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, this disclosure is also intended to include such modifications and variations.
[0147] It should also be noted that the exemplary embodiments mentioned in this disclosure describe methods or systems based on a series of steps or apparatus. However, this disclosure is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
Claims
1. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; the processor executes the computer program instructions to implement: Obtain the original image; Detect target objects in the original image to obtain at least one target detection box; The original image is cropped based on the at least one target detection box to obtain the region image corresponding to the target object; The target image is determined based on the region image and the zero-background image.
2. The device according to claim 1, characterized in that, The processor is specifically used to implement: Construct the zero-background image with the same dimensions as the original image; The target image is obtained by filling the region image with the zero background image.
3. The device according to claim 2, characterized in that, The processor is specifically used to implement: Based on the position of the region image in the original image, the region image is filled into the corresponding position of the all-zero background image with the same size as the original image to obtain the target image.
4. The device according to any one of claims 1-3, characterized in that, In response to the presence of multiple target detection boxes, the processor is specifically configured to implement: Determine a coverage pattern that covers multiple target detection boxes; crop the original image based on the coverage pattern to obtain a region image corresponding to the target object; or... The original image is cropped based on each of the target detection boxes to obtain a local image corresponding to each target detection box; each of the local images is used as the region image.
5. The device according to any one of claims 1-4, characterized in that, The processor is specifically used to implement: The target object is identified in the original image using an image detection model to obtain at least one initial detection box. The size of the at least one initial detection box is adjusted to obtain at least one target detection box.
6. The device according to claim 5, characterized in that, The processor is specifically used to implement: The at least one initial detection box is expanded to obtain at least one target detection box.
7. The device according to claim 6, characterized in that, The processor is specifically used to implement: The image detection model identifies the target object in the original image and obtains the confidence score corresponding to each initial detection box; the initial detection boxes with a confidence score greater than a first preset confidence score are determined as the target detection boxes; or, The initial detection box is subjected to boundary expansion processing to obtain the target detection box; wherein, the boundary expansion processing includes expanding the width and height of the detection box by a preset ratio, and / or expanding the boundary of the detection box by a fixed pixel value.
8. The device according to any one of claims 5-7, characterized in that, The processor is also used to implement: The target image is encoded and compressed to obtain target image data.
9. The device according to claim 8, characterized in that, The processor is specifically used to implement: The at least one initial detection box is aligned according to the preset encoding unit size used during the encoding compression to obtain at least one target detection box.
10. The device according to claim 9, characterized in that, The processor is specifically used to implement: Adjust the boundary of the region corresponding to the initial detection box so that the width and height of the region corresponding to the adjusted initial detection box are both integer multiples of the preset encoding unit size used in the encoding compression.
11. An image processing method, characterized in that, include: Obtain the original image; Detect target objects in the original image to obtain at least one target detection box; The original image is cropped based on the at least one target detection box to obtain the region image corresponding to the target object; The target image is determined based on the region image and the zero-background image.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the image processing method as described in claim 11.
13. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the image processing method as described in claim 11.