Image processing method and device, vehicle and storage medium
Patent Information
- Application Number
- CN202610658966.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-13
- Publication Date
- 2026-09-01
AI Technical Summary
[0002]目前智能驾驶系统主要依赖可见光摄像头、激光雷达和毫米波雷达的融合感知,对车辆进行控制,然而,在低光照条件下,可见光摄像头图像质量会较严重退化,激光雷达对低反射率物体敏感度下降,毫米波雷达较难提供准确语义信息,导致车辆基于采集到的图像进行视觉检测与识别任务的准确性较低
[0033] In this embodiment, firstly, based on initial or target control parameters, the vehicle's image acquisition device is controlled to acquire images of the vehicle's surrounding environment, obtaining raw images. Then, based on an image enhancement model, the raw images are enhanced to obtain the target image. This application controls the vehicle's image acquisition device based on initial or target control parameters, resolving the contradiction between the lack of data during cold starts and the need for adaptive dynamic environments in nighttime perception through the synergy of prior presets and feedback improvements. This ensures stable acquisition of effective images even when first entering complex scenes, avoiding perception interruptions. The acquisition strategy is dynamically fine-tuned based on real-world task performance, achieving dynamic improvement in acquisition quality with varying environments and tasks without increasing hardware costs. The original image is fed into the image enhancement model for processing. The image quality of the target image can be used to determine the device adjustment parameters of the original image. This allows for the adjustment of the initial control parameters used by the image acquisition device when acquiring the original image for the next time, thus achieving closed-loop adjustment of the control parameters of the image acquisition device. This enables the acquisition strategy to be guided in reverse according to actual perception needs, thereby reducing false detections and missed detections caused by image degradation. It also improves the accuracy and stability of visual detection and recognition tasks in nighttime and low-light environments, thereby solving the technical problem of low accuracy in vehicle visual detection and recognition tasks based on acquired images in related technologies.
Smart Images

Figure CN122676142A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle technology, and more specifically, to an image processing method, apparatus, vehicle, and storage medium. Background Technology
[0002] Currently, intelligent driving systems mainly rely on the fusion perception of visible light cameras, lidar, and millimeter-wave radar to control vehicles. However, under low light conditions, the image quality of visible light cameras degrades significantly, lidar becomes less sensitive to objects with low reflectivity, and millimeter-wave radar struggles to provide accurate semantic information. This results in low accuracy for vehicles performing visual detection and recognition tasks based on the acquired images.
[0003] There is currently no good solution to the above problems. Summary of the Invention
[0004] This application provides an image processing method, apparatus, vehicle, and storage medium to at least solve the technical problem of low accuracy in visual detection and recognition tasks of vehicles based on acquired images in related technologies.
[0005] According to one aspect of the embodiments of this application, an image processing method is provided, comprising: controlling an image acquisition device of a vehicle to acquire images of the vehicle's surrounding environment based on initial control parameters or target control parameters to obtain an original image, wherein the target control parameters are obtained by adjusting the initial control parameters based on device adjustment parameters corresponding to historical images; and performing image enhancement processing on the original image based on an image enhancement model to obtain a target image.
[0006] In the above embodiments of this application, the method further includes: determining the image quality index of the target image; and determining the device adjustment parameters corresponding to the original image based on the image quality index, wherein the device adjustment parameters corresponding to the original image are used to determine the target control parameters used when acquiring the next original image.
[0007] In the above embodiments of this application, determining the device adjustment parameters corresponding to the original image based on the image quality index includes: performing an image application task based on the target image to obtain the task execution result, wherein the image application task is used to represent a visual detection and recognition task performed based on the target image; and determining the device adjustment parameters corresponding to the original image based on the image quality index and the result confidence of the task execution result.
[0008] In the above embodiments of this application, the image acquisition device includes at least: an image sensor and a vehicle lighting unit; controlling the vehicle's image acquisition device to acquire images of the vehicle's surrounding environment based on initial control parameters or target control parameters to obtain an original image includes: when no historical image exists, controlling the image acquisition device based on the initial control parameters to obtain an original image; when historical image exists, adjusting the initial control parameters based on the device adjustment parameters corresponding to the historical image to obtain target control parameters, and controlling the image acquisition device based on the target control parameters to obtain an original image.
[0009] In the above embodiments of this application, the method further includes: acquiring vehicle driving status information and environmental status information of the environment in which the vehicle is located; and determining initial control parameters based on the driving status information and / or environmental status information.
[0010] In the above embodiments of this application, the image enhancement model includes a target encoder and a target decoder; the image enhancement processing of the original image based on the image enhancement model to obtain the target image includes: using the target encoder and multiple image enhancement tasks to perform image enhancement processing on the original image to obtain an intermediate feature representation, wherein the multiple image enhancement tasks include at least two of the following: image denoising, image illumination compensation and image dehazing; and using the target decoder to perform feature reconstruction on the intermediate feature representation to obtain the target image.
[0011] In the above embodiments of this application, the target encoder includes a self-attention module, a node learning module, and a multi-scale feature extraction module; using the target encoder and multiple image enhancement tasks, image enhancement processing is performed on the original image to obtain an intermediate feature representation, including: using the self-attention module to perform target-dimensional attention processing on the original image to obtain a weighted feature map; using the node learning module and multiple image enhancement tasks to perform feature modulation processing on the weighted feature map to obtain a composite feature representation; and using the multi-scale feature extraction module to perform multi-scale feature extraction on the composite feature representation to obtain an intermediate feature representation.
[0012] In the above embodiments of this application, a self-attention module is used to perform target-dimensional attention processing on the original image to obtain a weighted feature map, including: performing spatial-dimensional attention processing on the degraded region in the original image to obtain a first feature map; performing channel-dimensional attention processing on the target image channel of the original image to obtain a second feature map, wherein the target image channel is determined based on multiple image enhancement tasks; and obtaining a weighted feature map based on the first feature map and / or the second feature map.
[0013] In the above embodiments of this application, the composite feature representation includes task-shared features and / or task-specific features for multiple image enhancement tasks; using a node learning module and multiple image enhancement tasks, feature modulation processing is performed on the weighted feature map to obtain the composite feature representation, including: determining an inter-task collaboration index between any two image enhancement tasks in the multiple image enhancement tasks to obtain at least one inter-task collaboration index; if, among the at least one inter-task collaboration index, there is an inter-task collaboration index greater than a preset collaboration threshold, a shared network is invoked to perform feature modulation processing on the weighted feature map to obtain inter-task shared features; if, among the at least one inter-task collaboration index, there is an inter-task collaboration index less than or equal to the preset collaboration threshold, task adapters corresponding to the two image enhancement tasks are invoked to perform feature modulation processing on the weighted feature map to obtain task-specific features.
[0014] According to another aspect of the embodiments of this application, an image processing apparatus is also provided, comprising: an acquisition module, configured to control an image acquisition device of a vehicle to acquire images of the vehicle's surrounding environment based on initial control parameters or target control parameters, thereby obtaining an original image, wherein the target control parameters are obtained by adjusting the initial control parameters based on device adjustment parameters corresponding to historical images; and an image enhancement processing module, configured to perform image enhancement processing on the original image based on an image enhancement model, thereby obtaining a target image.
[0015] The acquisition module is also used to determine the image quality index of the target image; and to determine the device adjustment parameters corresponding to the original image based on the image quality index. The device adjustment parameters corresponding to the original image are used to determine the target control parameters to be used when acquiring the next original image.
[0016] The acquisition module is also used to perform image application tasks based on the target image and obtain task execution results. The image application task is used to represent visual detection and recognition tasks performed based on the target image. Based on the image quality index and the confidence level of the task execution results, the device adjustment parameters corresponding to the original image are determined.
[0017] The image acquisition device includes at least: an image sensor and a vehicle lighting unit; the acquisition module is also used to control the image acquisition device based on initial control parameters to obtain the original image when no historical image exists; and to adjust the initial control parameters based on the device adjustment parameters corresponding to the historical image to obtain target control parameters, and to control the image acquisition device based on the target control parameters to obtain the original image.
[0018] The acquisition module is also used to acquire vehicle driving status information and environmental status information of the vehicle's environment; and to determine initial control parameters based on the driving status information and / or environmental status information.
[0019] The image enhancement model includes a target encoder and a target decoder; the image enhancement processing module is also used to perform image enhancement processing on the original image using the target encoder and multiple image enhancement tasks to obtain intermediate feature representations. The multiple image enhancement tasks include at least two of the following: image denoising, image illumination compensation, and image dehazing; the target decoder is used to perform feature reconstruction on the intermediate feature representations to obtain the target image.
[0020] The target encoder includes a self-attention module, a node learning module, and a multi-scale feature extraction module. The image enhancement processing module is also used to apply attention processing to the target dimension of the original image using the self-attention module to obtain a weighted feature map. The node learning module and multiple image enhancement tasks are used to perform feature modulation processing on the weighted feature map to obtain a composite feature representation. The multi-scale feature extraction module is used to extract multi-scale features from the composite feature representation to obtain an intermediate feature representation.
[0021] The image enhancement processing module is further used to perform spatial dimension attention processing on the degraded regions in the original image to obtain a first feature map; to perform channel dimension attention processing on the target image channels of the original image to obtain a second feature map, wherein the target image channels are determined based on multiple image enhancement tasks; and to obtain a weighted feature map based on the first feature map and / or the second feature map.
[0022] The composite feature representation includes task-shared features and / or task-specific features for multiple image enhancement tasks. The image enhancement processing module is further configured to determine the task-to-task coordination index between any two image enhancement tasks to obtain at least one task-to-task coordination index. If, among the at least one task-to-task coordination index, there exists a task-to-task coordination index greater than a preset coordination threshold, a shared network is invoked to perform feature modulation processing on the weighted feature map to obtain task-shared features. If, among the at least one task-to-task coordination index, there exists a task-to-task coordination index less than or equal to the preset coordination threshold, task adapters corresponding to the two image enhancement tasks are invoked to perform feature modulation processing on the weighted feature map to obtain task-specific features.
[0023] According to another aspect of the embodiments of this application, a vehicle is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.
[0024] The aforementioned memory can refer to devices inside a computer used to store data and programs, including RAM, hard disks, etc. RAM can be used to temporarily store running programs and data, while hard disks can be used to store programs and data long-term. Memory enables the computer to read and write data and execute programs. The aforementioned processor is responsible for executing instructions in computer programs and performing data processing. It can also be responsible for controlling and executing various operations, including arithmetic operations, logical operations, and data transmission.
[0025] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.
[0026] The aforementioned computer storage media can refer to the media used in computer memory to store certain discontinuous physical quantities. Computer storage media mainly include semiconductors, magnetic cores, magnetic drums, magnetic tapes, laser discs, etc. Computer-readable storage media include stored programs, which can be a set of instructions that a computer can recognize and execute, running on an electronic computer to meet certain information needs.
[0027] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0028] The aforementioned computer program products can refer to software programs that have been written, tested, and released, and can run on computers or other devices. Computer program products can include application programs, operating systems, utility software, etc., used to achieve specific functions or solve specific problems.
[0029] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the methods in various embodiments of this application.
[0030] The aforementioned non-volatile computer-readable storage medium can refer to a medium for storing data. Non-volatile computer-readable storage media can retain data without loss when power is off and can be used to store long-term data, such as operating systems, applications, and user files. Non-volatile storage media can include hard disk drives, solid-state drives, optical disks, and flash memory storage devices, etc.
[0031] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.
[0032] The aforementioned computer program can refer to a set of instructions used to tell the computer to perform specific tasks or operations. Computer programs can be written by programmers using specific programming languages and can include algorithms, data structures, logic, and control flow. Computer programs can be used for a variety of purposes, including application software, operating systems, etc.
[0033] In this embodiment, firstly, based on initial or target control parameters, the vehicle's image acquisition device is controlled to acquire images of the vehicle's surrounding environment, obtaining raw images. Then, based on an image enhancement model, the raw images are enhanced to obtain the target image. This application controls the vehicle's image acquisition device based on initial or target control parameters, resolving the contradiction between the lack of data during cold starts and the need for adaptive dynamic environments in nighttime perception through the synergy of prior presets and feedback improvements. This ensures stable acquisition of effective images even when first entering complex scenes, avoiding perception interruptions. The acquisition strategy is dynamically fine-tuned based on real-world task performance, achieving dynamic improvement in acquisition quality with varying environments and tasks without increasing hardware costs. The original image is fed into the image enhancement model for processing. The image quality of the target image can be used to determine the device adjustment parameters of the original image. This allows for the adjustment of the initial control parameters used by the image acquisition device when acquiring the original image for the next time, thus achieving closed-loop adjustment of the control parameters of the image acquisition device. This enables the acquisition strategy to be guided in reverse according to actual perception needs, thereby reducing false detections and missed detections caused by image degradation. It also improves the accuracy and stability of visual detection and recognition tasks in nighttime and low-light environments, thereby solving the technical problem of low accuracy in vehicle visual detection and recognition tasks based on acquired images in related technologies. Attached Figure Description
[0034] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0035] Figure 1 This is a flowchart of an image processing method according to an embodiment of the present invention;
[0036] Figure 2 This is a schematic diagram of an image processing system architecture according to an embodiment of the present invention;
[0037] Figure 3 This is a schematic diagram of an image enhancement process for an original image according to an embodiment of the present invention;
[0038] Figure 4 This is a schematic diagram of a process for generating a composite feature representation according to an embodiment of the present invention;
[0039] Figure 5 This is a schematic diagram of an image processing procedure according to an embodiment of the present invention;
[0040] Figure 6 This is a schematic diagram of an image processing apparatus according to an embodiment of the present invention. Detailed Implementation
[0041] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0042] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0043] According to an embodiment of this application, an embodiment of an image processing method is provided. The steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowcharts, in some cases the steps shown or described may be performed in a different order than that shown here.
[0044] This embodiment provides an image processing method. Figure 1 This is a flowchart of an image processing method according to an embodiment of this application, such as... Figure 1 As shown, the process includes the following steps:
[0045] Step S102: Based on the initial control parameters or target control parameters, control the vehicle's image acquisition device to acquire images of the vehicle's surrounding environment to obtain the original image.
[0046] The target control parameters are obtained by adjusting the initial control parameters based on the device adjustment parameters corresponding to historical images.
[0047] The aforementioned initial control parameters refer to the set of control commands for the image acquisition device determined based on current vehicle driving information, such as vehicle speed, gear, and steering angle, and environmental information, such as ambient light intensity, weather conditions, and time, without referencing historical image feedback. Initial control parameters can be used to initiate the image acquisition process for the first time or when there is no historical data. These parameters may include sensor configuration parameters such as exposure time, analog gain, digital gain, and frame rate, as well as lighting control commands such as whether to enable high dynamic range mode and whether to turn on the supplementary lighting.
[0048] The aforementioned target control parameters refer to the acquisition control commands obtained by dynamically improving the initial control parameters based on the corresponding device adjustment parameters of historical images, assuming the existence of historical images and corresponding device adjustment parameters. The target control parameters are the output of a closed-loop feedback mechanism, aiming to improve image acquisition quality. The adjustment is based on the enhancement effect evaluation of the previous frame or multiple frames and the confidence analysis of the upper-level perception task. For example, if the signal-to-noise ratio is low in the dark area of the previous frame, the target control parameters will appropriately increase exposure or decrease gain to balance noise and brightness.
[0049] The aforementioned image acquisition equipment refers to a hardware system installed on a vehicle to acquire visual information about the surrounding environment. It may include at least an image sensor, such as an automotive-grade image sensor, and vehicle lighting units, such as headlights, taillights, and adjustable matrix headlights. The image acquisition equipment possesses programmable control capabilities, enabling it to dynamically adjust parameters such as sensor exposure, gain, and shutter speed according to control parameters. Simultaneously, it controls the brightness, color temperature, and beam distribution mode of the supplementary lighting, such as low beam, high beam, and area illumination, thereby improving the signal acquisition quality of the original image at a physical level.
[0050] The aforementioned raw image can refer to the original visual data frame actually acquired by the image acquisition device under the drive of initial control parameters or target control parameters. The raw image may contain degradation phenomena such as low light, noise, fog, overexposure of highlights, and motion blur, and serves as input data for subsequent image enhancement models.
[0051] The aforementioned historical images may refer to image data acquired by an image acquisition device at a previous point in time, such as the previous frame or several previous frames.
[0052] The aforementioned equipment adjustment parameters can refer to parameters determined based on historical images and used to adjust the initial control parameters. These parameters may include the sensor's exposure time, gain, frame rate, etc., as well as the brightness, color temperature, and on / off status of the fill light.
[0053] As an optional implementation, initial control parameters can be used to drive image acquisition based on prior environmental information. This approach is suitable for scenarios without historical feedback. When a vehicle first enters a low-light environment, such as a road section without streetlights at night, and there are no historical image records, initial control parameters are generated based on real-time perceived environmental and driving information. For example, image sensor exposure time, analog gain, and digital gain can be set, and the low beam mode of the headlights and near-infrared supplementary lights can be activated to enhance infrared reflection signals. At this time, the image acquisition device performs acquisition according to a preset scene and parameter mapping table without feedback, and outputs the first frame of raw image. This process relies on a static rule base and sensor input to ensure the perception capability of image acquisition during cold starts and avoid image acquisition failure due to missing parameters.
[0054] As an alternative implementation, target control parameter-driven acquisition based on closed-loop feedback can be performed. This approach is applicable to scenarios with historical feedback. When a previous frame of the original image has been acquired and processed, and the perception quality analyzer reports that the signal-to-noise ratio of the original image is below a threshold and the target detection confidence has decreased, the device adjustment parameters corresponding to the historical image can be determined. Combined with the current performance degradation level, the target control parameters are dynamically adjusted, such as increasing exposure and increasing the brightness of the fill light. This adjustment strategy prioritizes improving signal strength and suppressing gain noise without increasing motion blur. Then, the image acquisition device performs acquisition according to the target control parameters to obtain an improved original image. This process achieves a self-evolving closed loop of perception result feedback adjustment acquisition strategy.
[0055] By employing the above settings, the vehicle's image acquisition equipment is controlled based on initial or target control parameters. Through the synergy of prior presets and feedback improvements, the contradiction between the lack of data during cold starts in nighttime perception and the need for adaptive adaptation to dynamic environments is resolved. This ensures stable acquisition of effective images even when first entering complex scenes, avoiding perception interruptions. Furthermore, the acquisition strategy is dynamically fine-tuned based on real-world task performance, achieving dynamic improvement in acquisition quality with changing environments and tasks without increasing hardware costs. This enhances the robustness and continuity of perception in dynamic scenarios such as tunnel entrances / exits, rain, fog, and nighttime conditions, providing a reliable data foundation for end-to-end intelligent enhancement.
[0056] Step S104: Perform image enhancement processing on the original image based on the image enhancement model to obtain the target image.
[0057] The aforementioned image enhancement model refers to a deep neural network model deployed in the backend algorithm improvement layer, used for multi-task collaborative visual enhancement processing of the original image. The image enhancement model can adopt an encoder-decoder structure and may include components such as self-attention modules, task-oriented node learning modules, and multi-scale feature extraction modules. It can handle various image degradation problems such as image denoising, low-light enhancement, and nighttime defogging. Through a dynamic routing mechanism that shares task-specific features, the image enhancement model can achieve multi-scene adaptive enhancement while ensuring computational efficiency, outputting a target image with improved visual quality.
[0058] The target image mentioned above can refer to the visual image output by the image enhancement model after processing the original image, which has undergone multi-dimensional improvements. The target image has higher brightness, contrast, detail preservation, and lower noise level, making it suitable for subsequent autonomous driving perception tasks, such as object detection, semantic segmentation, and lane recognition. The quality of the target image is quantitatively evaluated through the confidence level of the perception results and image quality indicators, and can be used as a feedback signal to drive the adjustment of initial control parameters.
[0059] As an optional implementation, when it is determined that the current original image is mainly degraded by low illumination and noise superposition, such as a road without streetlights at night, and the synergy index between low illumination enhancement and image denoising tasks, such as gradient cosine similarity, exceeds a corresponding preset threshold, the image enhancement model activates the synergistic enhancement mode. After the original image is focused on the dark and textured regions by the self-attention module, the feature stream enters the shared network. In the encoder, structural and brightness features are extracted through a unified convolutional path. Subsequently, the shared parameters in the task-oriented node learning module globally modulate the features to complete brightness stretching and noise suppression. The multi-scale feature extraction module fuses local edges and global illumination distribution, and the decoder reconstructs a target image with uniform brightness, clear details, and suppressed noise.
[0060] As an alternative implementation, when the original image was captured in a dense foggy environment at night, the coordination between nighttime defogging and image denoising tasks is low. Furthermore, defogging requires enhancing color channels while denoising requires suppressing high-frequency textures, leading to feature conflicts. In such cases, the image enhancement model can switch to an independent enhancement mode. After the self-attention module in the image enhancement model locates the fog area, the node learning module activates both the defogging and denoising task adapters. The defogging task adapter reconstructs color saturation through conditional normalization and an atmospheric scattering model, while the denoising task adapter uses lightweight residual blocks to locally filter out particle noise. The features output from the defogging and denoising task adapters retain task-specific information. After independent fusion by multi-scale modules, the decoder performs weighted merging in the spatial domain to generate a target image that restores the contours of distant objects while preserving nearby texture details. The image enhancement model avoids negative transfer through task isolation, ensuring the accuracy of key perception tasks.
[0061] With the above settings, the image enhancement model has the ability to dynamically adapt to tasks, improves computational efficiency and resource utilization in collaborative scenarios, prevents semantic distortion caused by feature interference in conflict scenarios, and enhances the confidence and stability of target detection and semantic understanding in complex nighttime scenarios, providing highly reliable perception input for autonomous driving.
[0062] In this embodiment, firstly, based on initial or target control parameters, the vehicle's image acquisition device is controlled to acquire images of the vehicle's surrounding environment, obtaining raw images. Then, based on an image enhancement model, the raw images are enhanced to obtain the target image. This application controls the vehicle's image acquisition device based on initial or target control parameters, resolving the contradiction between the lack of data during cold starts and the need for adaptive dynamic environments in nighttime perception through the synergy of prior presets and feedback improvements. This ensures stable acquisition of effective images even when first entering complex scenes, avoiding perception interruptions. The acquisition strategy is dynamically fine-tuned based on real-world task performance, achieving dynamic improvement in acquisition quality with varying environments and tasks without increasing hardware costs. The original image is fed into the image enhancement model for processing. The image quality of the target image can be used to determine the device adjustment parameters of the original image. This allows for the adjustment of the initial control parameters used by the image acquisition device when acquiring the original image for the next time, thus achieving closed-loop adjustment of the control parameters of the image acquisition device. This enables the acquisition strategy to be guided in reverse according to actual perception needs, thereby reducing false detections and missed detections caused by image degradation. It also improves the accuracy and stability of visual detection and recognition tasks in nighttime and low-light environments, thereby solving the technical problem of low accuracy in vehicle visual detection and recognition tasks based on acquired images in related technologies.
[0063] In the above embodiments of this application, the method further includes: determining the image quality index of the target image; and determining the device adjustment parameters corresponding to the original image based on the image quality index, wherein the device adjustment parameters corresponding to the original image are used to determine the target control parameters used when acquiring the next original image.
[0064] The aforementioned image quality indicators can refer to evaluation parameters used to quantitatively assess the visual quality and perceived usability of target images, serving as the basis for improving the initial control parameters in the feedback loop.
[0065] As an optional implementation, in low-complexity scenarios, such as an empty nighttime highway, where the confidence level of upper-layer perception tasks, such as target detection, is stable, closed-loop improvement is prioritized based on low-level image quality metrics. The perception quality analyzer calculates the signal-to-noise ratio (SNR), average gradient, and contrast of the current target image. If the SNR is below the corresponding threshold and the average gradient is low, it indicates that noise in the dark areas of the target image is not sufficiently suppressed and details are blurred. Reverse deduction can be performed, revealing that the original image suffers from noise amplification due to insufficient exposure and excessive gain. This allows for the generation of device adjustment parameters for the original image, such as extending the exposure of the next frame, reducing gain, and fine-tuning the brightness of the fill light, prioritizing signal enhancement and reducing gain noise. This process, through image physical quality self-correction, is fast-responding, computationally lightweight, and suitable for continuous improvement in stable environments.
[0066] As an alternative implementation, when the confidence level for detecting a pedestrian ahead drops sharply in rainy or foggy nighttime conditions, even if the image signal-to-noise ratio is high, the target edges are blurred and the semantics are unclear. This is determined to be a perception failure rather than poor image quality. In this case, the perception quality analyzer uses the task confidence level as a feedback signal to infer that the original image has high-frequency detail loss or contrast distortion at the feature level. Then, the initial control parameters are adjusted: instead of increasing exposure to avoid motion blur, the color temperature of the fill light is increased to enhance contour contrast, and the high dynamic range multi-frame fusion mode of the image sensor is activated to improve edge sharpness. This adjustment does not change the overall brightness but specifically improves semantic discriminability, forming an accurate reverse control from the perception result to the acquisition strategy.
[0067] Through the above settings, the feedback mechanism enables improvements driven by image quality indicators, ensuring the physical fidelity of images. It also enables accurate and adaptive iteration of acquisition strategies in complex and dynamic environments, improving the robustness and usability of the perception system in real-world scenarios, and building a closed-loop, intelligent visual improvement engine.
[0068] In the above embodiments of this application, determining the device adjustment parameters corresponding to the original image based on the image quality index includes: performing an image application task based on the target image to obtain the task execution result, wherein the image application task is used to represent a visual detection and recognition task performed based on the target image; and determining the device adjustment parameters corresponding to the original image based on the image quality index and the result confidence of the task execution result.
[0069] The aforementioned image application tasks refer to visual recognition and detection tasks performed by upper-level modules of an autonomous driving system for environmental perception and decision-making, based on target images processed by image enhancement models. Image application tasks may include, but are not limited to, object detection, lane recognition, semantic segmentation, traffic light recognition, and drivable area estimation.
[0070] The aforementioned task execution results can refer to the specific execution results output by the corresponding perception algorithm running on the target image in an image-based application task. The task execution results can include task-related entity information and spatial localization. For example, in an object detection task, the output is a set of bounding boxes, category labels, and corresponding spatial coordinates; in a lane line recognition task, the output is multiple curve parameters or a sequence of pixel-level line segments; in a semantic segmentation task, the output is a category label map with the same size as the input image.
[0071] The aforementioned result confidence level can refer to the quantitative evaluation value of the reliability of each recognition result generated internally by the perception algorithm when the image application task outputs the task execution result, which is used to characterize the credibility of the task execution result.
[0072] As an optional implementation, when the target image is processed by the target detection network and the confidence level for detecting pedestrians ahead is low, while the signal-to-noise ratio and contrast of the target image are within the normal range, it can be determined that the semantic information of the original image is insufficient. Analysis shows that this low confidence level stems from blurred target edges and lack of texture, suggesting that the original image suffers from underexposure or insufficient supplementary lighting, resulting in the loss of high-frequency details. Based on this, device adjustment parameters corresponding to the original image are generated. For example, without increasing the exposure time, the power of the automotive-grade near-infrared supplementary light is increased and switched to a narrow beam focusing mode to enhance the infrared reflection intensity of the pedestrian's reflective material. At the same time, the sensor's analog gain is slightly increased to strengthen the edge response. This adjustment achieves an accurate closed loop of inferring the acquisition strategy from the results.
[0073] As another optional implementation, when the target image is used for lane recognition, and the confidence level of the task execution result remains consistently high, and the clarity and noise level of the target image are better than the historical average, it is determined that the current acquisition and enhancement process is in an optimal state. In this case, the current acquisition strategy of the original image can be maintained to avoid blind intervention, prevent artifacts or power waste caused by excessive improvement, and achieve minimization of resource usage and maximization of stability.
[0074] By integrating the confidence level of the task execution results with image quality indicators, the system can accurately distinguish between two types of problems: recognition failure due to poor visual quality and problems with good visual quality but missing semantic information, enabling differentiated control. This improves the reliability of perception in complex nighttime conditions, avoids invalid parameter disturbances, enhances the system's robustness, energy efficiency, and decision-making safety in real traffic scenarios, and helps to build an intelligent data acquisition closed loop guided by perception needs.
[0075] In the above embodiments of this application, the image acquisition device includes at least: an image sensor and a vehicle lighting unit; controlling the vehicle's image acquisition device to acquire images of the vehicle's surrounding environment based on initial control parameters or target control parameters to obtain an original image includes: when no historical image exists, controlling the image acquisition device based on the initial control parameters to obtain an original image; when historical image exists, adjusting the initial control parameters based on the device adjustment parameters corresponding to the historical image to obtain target control parameters, and controlling the image acquisition device based on the target control parameters to obtain an original image.
[0076] The aforementioned image sensor can refer to a hardware device installed externally on a vehicle to capture visible light images of the surrounding environment. It can be an automotive-grade image sensor used to convert sensed optical signals into digital image data. The image sensor has a programmable control interface, supporting dynamic adjustment of parameters such as exposure time, analog gain, digital gain, shutter speed, and automatic white balance. In this application, the image sensor serves as the direct acquisition source of the raw image. Its parameter configuration is adjusted in real-time by the intelligent imaging controller based on initial or target control parameters to maximize effective signal capture capability in complex scenes such as low light, high dynamic range, or motion blur, suppress noise and overexposure, and provide a high-quality raw data foundation for subsequent image enhancement.
[0077] The aforementioned vehicle lighting unit can refer to a programmable automotive-grade light source system used for active lighting environments on a vehicle, including but not limited to headlights, matrix LED headlights, adaptive high beams, taillights, and ambient lighting, such as dedicated near-infrared or white light supplementary lighting modules. The vehicle lighting unit is controlled by an intelligent imaging controller, which can dynamically adjust brightness, color temperature, and beam distribution patterns according to control parameters, such as zoned illumination, blocking oncoming vehicle areas, on / off states, and emission wavelengths, such as visible light or near-infrared. In this application, the vehicle lighting unit works in conjunction with an image sensor to enhance scene illumination through active supplementary lighting. In environments without natural light or with low illumination, such as tunnels without streetlights or rural roads at night, it can effectively enhance the signal-to-noise ratio of the image sensor, achieving physical layer image quality improvement.
[0078] As an optional implementation, when a vehicle enters a tunnel without streetlights or a rural road at night for the first time, and there are no historical image records, the image acquisition device generates initial control parameters based on a preset environment and parameter mapping table. For example, based on the illuminance and vehicle speed detected by the ambient light sensor, it automatically sets the image sensor mode and turns on the low beam headlights and near-infrared supplementary lights. The image acquisition device relies on real-time input from environmental and driving sensors, without requiring historical feedback, and directly performs acquisition, outputting the first frame of raw image. This ensures stable acquisition of effective visual input even without historical data, avoiding perceptual gaps.
[0079] As an alternative implementation, when acquiring and processing the original image, if the perception quality analyzer reports low pedestrian detection confidence and low signal-to-noise ratio in the current target image, it determines the corresponding device adjustment parameters for the original image. Based on the current performance degradation trend, it dynamically generates target control parameters, such as extending exposure to increase signal strength, reducing gain to suppress noise amplification, and increasing the brightness of the near-infrared illuminator. The image acquisition device executes acquisition according to the target control parameters to obtain the next frame of the original image, achieving a self-learning closed loop where the acquisition strategy is derived from the perception results, thus continuously approaching a better imaging state.
[0080] Through the above settings, a seamless connection between reliable startup and intelligent evolution is achieved. The cold start mode ensures perception capabilities in unknown environments, while the closed-loop optimization mode continuously improves the acquisition strategy based on real-world task performance, preventing static parameters from failing in dynamic scenarios. This enables the system to cope with complex environments upon first entry and to adaptively adjust based on actual perception results, improving the stability, continuity, and energy efficiency of nighttime perception. This provides a practical and dynamic response foundation for end-to-end intelligent vision enhancement.
[0081] In the above embodiments of this application, the method further includes: acquiring vehicle driving status information and environmental status information of the environment in which the vehicle is located; and determining initial control parameters based on the driving status information and / or environmental status information.
[0082] The aforementioned driving status information refers to real-time data on the vehicle's own operating status, sourced from the vehicle's controller local area network bus and sensor system. This data is used to determine the dynamic demands of the current driving scenario, thereby assisting in decision-making regarding image acquisition strategies. Driving status information may include, but is not limited to, vehicle speed, gear information, steering angle and yaw rate, braking status and acceleration, and autonomous driving mode status.
[0083] The aforementioned environmental condition information refers to real-time perceived data of the vehicle's external environment, used to assess the degree of degradation of current imaging conditions, thereby guiding the parameter configuration of the image acquisition equipment. Environmental condition information may include, but is not limited to, ambient light intensity, weather conditions, time information, location and map information, humidity, and visibility data.
[0084] As an optional implementation, when a vehicle enters a suburban road at night, the ambient light sensor detects the illuminance, and the weather sensor determines that there is no precipitation, no fog, and high air transparency. Based on a preset environment and parameter mapping table, initial control parameters are determined. For example, the image sensor is set to medium exposure and medium gain, the active fill light is turned off (to avoid light scattering interference due to the dark but fog-free environment), and near-infrared mode is enabled to enhance road marking reflection. Because the vehicle speed is low, high dynamic range multi-frame fusion is not triggered to reduce latency. This process relies on real-time readings from the environmental sensor and does not require vehicle speed or path information, enabling a rapid response where the light environment determines the acquisition strategy. It is suitable for stable, low-interference nighttime driving scenarios.
[0085] As an alternative implementation, when a vehicle enters a nighttime section of an urban expressway, the ambient light sensor detects the illuminance, identifies the current location as a tunnel exit using a high-precision map, and obtains the vehicle speed via the controller area network bus. Determining that the scenario presents a dual risk of sudden changes in strong light and high-speed movement, a joint decision-making logic is initiated. To suppress glare and motion blur at the exit, the initial control parameters are set to short exposure and high gain. The headlight matrix intelligent dimming is activated, illuminating only the carriageway area to avoid dazzling oncoming vehicles, and the near-infrared supplemental lights are activated to high brightness mode to ensure clear imaging of high-speed moving targets in low-light conditions. This process integrates environmental and driving information to achieve scenario-risk-oriented parameter improvements.
[0086] Through the above settings, collaborative modeling of environmental perception and driving intent is achieved. Environmental perception ensures reasonable basic imaging, while driving intent enhances the ability to predict high-risk dynamic scenarios. This application upgrades the acquisition parameters to active adaptation through multi-source information fusion, improving the usability of images in high-dynamic and high-risk scenarios without increasing hardware costs. This provides more stable and information-rich raw input for subsequent enhancement and perception tasks, contributing to the construction of an all-weather intelligent visual perception system.
[0087] In the above embodiments of this application, the image enhancement model includes a target encoder and a target decoder; the image enhancement processing of the original image based on the image enhancement model to obtain the target image includes: using the target encoder and multiple image enhancement tasks to perform image enhancement processing on the original image to obtain an intermediate feature representation, wherein the multiple image enhancement tasks include at least two of the following: image denoising, image illumination compensation and image dehazing; and using the target decoder to perform feature reconstruction on the intermediate feature representation to obtain the target image.
[0088] The aforementioned target encoder can refer to the backbone of a neural network in an image enhancement model, used to extract and abstract semantic features layer by layer from the original image. Its function is to map the low-quality input original image into an intermediate feature representation with high discriminative power. The target encoder adapts to different degradation types through a dynamic mechanism, achieving multi-knowledge collaborative enhancement.
[0089] The aforementioned target decoder can refer to the network part of an image enhancement model used to progressively restore the intermediate feature representation output by the encoder to the target image. It can employ upsampling techniques, such as transposed convolution, pixel shuffling, and skip connection structures, forming a symmetrical encoder-decoder architecture with the encoder. Its function is to fuse multi-level feature information, restore the spatial details and texture structure of the image, and output a target image improved in terms of brightness, contrast, sharpness, and color fidelity. The output of the target decoder and the original image can undergo loss calculation at the pixel level, supporting end-to-end training.
[0090] The aforementioned image enhancement tasks refer to at least two image degradation repair tasks that the image enhancement model needs to handle. These may include, but are not limited to, image denoising, i.e., reducing granular or speckled artifacts introduced by high gain or sensor thermal noise in the image; image illumination compensation, i.e., low-light enhancement, improving overall brightness and contrast, restoring visible details in dark areas, and correcting non-uniform illumination; and image dehazing, i.e., eliminating scattering effects caused by fog, water vapor, or dust at night, and improving image clarity and color saturation. These image enhancement tasks often coexist and intertwine in nighttime driving scenarios. This application uses a unified network architecture for collaborative processing to avoid information fragmentation and computational redundancy.
[0091] The aforementioned intermediate feature representation refers to a high-dimensional feature set rich in semantics and spatial details, generated after multi-level, multi-scale fusion of composite feature representations by a multi-scale feature extraction module. The intermediate feature representation integrates attention focusing, task modulation, and multi-scale perceptual information, serving as a semantic compression code for image reconstruction. The quality of the intermediate feature representation affects the sharpness, detail preservation, and noise suppression of the target image. Based on the intermediate feature representation, the target decoder can progressively reconstruct a high-quality target image through upsampling and residual connections.
[0092] As an optional implementation, when the original image exhibits typical nighttime low-light and high-gain noise characteristics, such as road sections without streetlights, image illumination compensation and image denoising become the primary tasks, and the synergistic index of image illumination compensation and image denoising exceeds a threshold. The target encoder first focuses on dark and textured areas through a self-attention module to extract shared basic features; subsequently, the task-oriented node learning module calls the shared network to uniformly modulate the features, achieving brightness stretching and noise suppression, enhancing overall brightness in the low-frequency part, and preserving edges and suppressing speckle noise in the high-frequency part. The multi-scale feature extraction module captures local details and global light distribution in parallel, forming a high-dimensional intermediate feature representation. The target decoder fuses the multi-layer features of the encoder through skip connections, gradually reconstructing the image, and outputting a target image with uniform brightness, clear texture, and suppressed noise while preserving structural integrity.
[0093] As an alternative implementation, when the original image was taken at night in rain and fog, the coordination between image dehazing and image denoising is low. This is because dehazing requires improving color saturation and global contrast, while denoising requires suppressing high-frequency artifacts. After the encoder extracts basic features, the node learning module activates two independent task adapters. The dehazing adapter enhances the cyan and blue channels and reconstructs the atmospheric scattering model through conditional normalization; the denoising adapter uses a lightweight residual structure to locally filter out particle noise. The two branches generate task-specific intermediate feature representations, which are independently enhanced by multi-scale modules and then weighted and fused by the decoder in the spatial domain. This prioritizes preserving the fog area contour and nearby texture details, avoiding feature interference and ensuring that dehazing does not blur edges and denoising does not lose structure.
[0094] Through the above settings, intelligent adaptation to complex degradation scenarios can be achieved, collaborative tasks share computing resources to improve efficiency, and conflicting tasks are handled independently to ensure accuracy. This enables the output of high-fidelity target images even in low-light, foggy, and noisy environments. It also improves the input quality of target detection and semantic segmentation, achieving a better balance between performance and efficiency on computationally limited in-vehicle platforms, and realizing intelligent image enhancement processing with multi-knowledge collaboration, low overhead, and high robustness.
[0095] In the above embodiments of this application, the target encoder includes a self-attention module, a node learning module, and a multi-scale feature extraction module; using the target encoder and multiple image enhancement tasks, image enhancement processing is performed on the original image to obtain an intermediate feature representation, including: using the self-attention module to perform target-dimensional attention processing on the original image to obtain a weighted feature map; using the node learning module and multiple image enhancement tasks to perform feature modulation processing on the weighted feature map to obtain a composite feature representation; and using the multi-scale feature extraction module to perform multi-scale feature extraction on the composite feature representation to obtain an intermediate feature representation.
[0096] The aforementioned self-attention module can refer to a neural network component embedded in the front end of the target encoder, used to dynamically allocate the importance of different spatial and channel regions in an image. The self-attention module automatically assigns higher weights to degraded regions by performing joint attention calculations on the original image in both spatial and channel dimensions, outputting a weighted feature map. Essentially, this is a reweighted initial feature response, enabling subsequent modules to focus on more valuable regions and channels.
[0097] The aforementioned node learning module refers to a processing unit located after the self-attention module and before the multi-scale feature extraction module, specifically designed for parallel processing of multiple image enhancement tasks. The node learning module receives the weighted feature map from the self-attention module, performs feature modulation for at least two image enhancement tasks, and generates a composite feature representation containing both task-shared and task-specific features. By modeling task relationships, it dynamically assesses task synergy and determines whether to invoke a shared network or task adapter, achieving a balance between knowledge transfer and conflict avoidance, thus avoiding negative transfer problems in multi-task networks.
[0098] The aforementioned multi-scale feature extraction module refers to a network structure used to capture semantic information at different scales in an image based on the composite feature representation output by the node learning module. It can consist of parallel dilated convolutions, feature pyramids, or convolutions with different kernel sizes. This multi-scale feature extraction module can extract fine-grained noise points, edge contours, and global scene structure, fusing local detail restoration with global contextual understanding capabilities. This improves the ability to recover from complex degradations such as dense fog, low light, and noise superposition, and outputs high-quality intermediate feature representations.
[0099] The aforementioned target dimension can refer to the feature dimension that the self-attention module focuses on when performing attention weighting. This may include, but is not limited to, spatial dimension, which focuses on which pixel areas in the image have severe degradation, such as vignetting at the tunnel entrance or fog in front; and channel dimension, which focuses on which feature channels are more important to the current enhancement task, such as enhancing the color channel for dehazing tasks and enhancing the texture channel for denoising tasks.
[0100] The aforementioned weighted feature map refers to the reweighted feature representation output by the self-attention module after calculating attention weights in the spatial and channel dimensions of the original image. In the weighted feature map, the feature responses corresponding to degenerate regions are amplified, while irrelevant or redundant regions are suppressed, serving as high-quality input for subsequent task-oriented modules.
[0101] The aforementioned composite feature representation can refer to a feature tensor generated by the node learning module after multi-task collaborative modulation of the weighted feature map, containing shared features between tasks and / or task-specific features. Shared features reflect the basic information needed by multiple augmentation tasks, such as edges and structures, while task-specific features retain the recovery strategies unique to each task, such as scattering modeling for dehazing and prior noise distribution for denoising. Composite feature representation achieves parameter and feature reuse while preserving task independence.
[0102] As an optional implementation, when the original image is taken at night on a road without streetlights, the main degradation is the coexistence of low brightness and high-gain noise, indicating high synergy between illumination compensation and denoising tasks. The self-attention module first focuses on dark areas, such as roadside shadows, in the spatial dimension, and enhances texture-sensitive channels, such as grayscale channels, in the channel dimension, generating weighted feature maps so that the network prioritizes processing areas with weak information. The node learning module calls a shared network to jointly modulate the weighted feature maps, simultaneously increasing global brightness and suppressing high-frequency noise, forming a composite feature representation with a balance between brightness and noise. Subsequently, the multi-scale feature extraction module captures local noise, mid-scale edges, and global illumination trends through parallel dilated convolution, fusing them into an intermediate feature representation that preserves dark details and structural integrity.
[0103] As an alternative implementation, when the original image was taken at night in rain or fog, a feature conflict is identified between the dehazing and denoising tasks. The self-attention module prioritizes locating foggy areas, i.e., low-contrast regions, and texture edges, generating a weighted feature map. The node learning module activates two independent adapters for dehazing and denoising based on the task relationship matrix. The dehazing adapter enhances color channels and corrects color casts, while the denoising adapter suppresses high-frequency artifacts. Both adapters generate task-specific composite feature representations. The multi-scale module processes the two feature paths separately: the dehazing branch focuses on large-scale structure recovery, while the denoising branch focuses on preserving small-scale details. The intermediate feature representation is a weighted fusion of the two branches, achieving transparency in foggy areas and clear textures without mutual interference.
[0104] Through the above setup, the collaborative architecture achieves precise and differentiated feature processing by focusing attention, intelligently modulating tasks, and employing multi-scale fusion and multi-level progression. It improves efficiency in collaborative tasks and ensures accuracy in conflicting tasks, enabling efficient recovery of brightness and detail in complex nighttime scenes while avoiding semantic distortion caused by task interference. This enhances the adaptability, robustness, and practicality of image enhancement, providing high-quality and reliable visual input for autonomous driving perception systems.
[0105] In the above embodiments of this application, a self-attention module is used to perform target-dimensional attention processing on the original image to obtain a weighted feature map, including: performing spatial-dimensional attention processing on the degraded region in the original image to obtain a first feature map; performing channel-dimensional attention processing on the target image channel of the original image to obtain a second feature map, wherein the target image channel is determined based on multiple image enhancement tasks; and obtaining a weighted feature map based on the first feature map and / or the second feature map.
[0106] The aforementioned degraded regions refer to local spatial areas in the original image where visual information is severely lacking or distorted due to environmental or sensor limitations such as low light, fog, noise, high dynamic range, or motion blur. These degraded regions can manifest as: dark areas, such as tunnel entrances, under trees, or unlit road sections, where low illumination causes pixel values to approach black levels, making details almost invisible; foggy or scattering areas, where light scattering caused by water vapor, dust, or rain at night reduces image contrast, makes colors appear grayish-white, and blurs edges; noise-dominated areas, such as granular or speckled random interference appearing under high-gain shooting; and overexposed or glare edges, caused by local saturation or halo areas from strong light sources such as oncoming headlights or streetlights, resulting in the loss of effective information. The degraded regions are dynamically identified and located by the self-attention module based on features such as local gradients, brightness distribution, and texture entropy of the original image. The identification target allocates higher computational weights to subsequent processing, focusing on areas that need more restoration, thereby improving resource utilization efficiency and enhancement effects.
[0107] The aforementioned target image channels refer to the image feature channels that are dynamically determined during the image enhancement process, based on the multiple image enhancement tasks to be performed, such as image denoising, low-light compensation, and nighttime dehazing, and that have a greater impact on task performance. Different enhancement tasks have different dependencies on channels. For example, image dehazing relies more on color and saturation channels to restore the true colors that are obscured by scattering; image denoising focuses more on high-frequency texture channels, such as those with strong edge and detail responses, to distinguish between real structures and noise; and image light compensation relies more on luminance channels to uniformly increase overall brightness without distortion.
[0108] As an alternative implementation, when the original image is in a nighttime environment without streetlights, the degraded areas are mainly concentrated at road edges and in shadow areas, while the image channels show insufficient brightness channel information. The self-attention module first increases the weight of low-brightness areas in the spatial dimension, generating a first feature map, allowing the network to focus on the dark structure. Based on illumination compensation and denoising tasks, the target image channels are determined, and the channel attention module dynamically weights each target image channel to generate a second feature map. Finally, the first and second feature maps are multiplied element-wise to obtain a weighted feature map, achieving a synergistic improvement in dark area enhancement and noise channel suppression, thereby improving the brightness uniformity and signal-to-noise ratio of the target image.
[0109] As an alternative implementation, when the original image is affected by nighttime fog, the degraded areas are widely distributed, but at the channel level, the red and green channels are oversaturated due to fog scattering, while the blue channel information is severely attenuated. Based on the nighttime defogging task, the target image channel is identified as the blue channel, while the red and green channels need to be suppressed from interference. The spatial attention module, due to uniform degradation, applies only low weights; the channel attention module, on the other hand, increases the weight of the blue channel and decreases the weights of the red and green channels, generating a highly selective second feature map. The weighted feature map is dominated by the second feature map, with the spatial feature map providing auxiliary compensation. This process, driven by the channel dimension, accurately recovers the blue features obscured by fog, such as road signs and car lights, improving the recognizability of foggy areas and avoiding artifact enhancement caused by excessive spatial attention.
[0110] Through the above setup, the spatial and channel-based dual-dimensional attention mechanism overcomes the limitations of single-dimensional attention. The spatial dimension accurately locates where enhancement is needed, while the channel dimension intelligently identifies what information needs strengthening. The synergy of the spatial and channel-based dual-dimensional attention mechanism enables the enhancement process to possess both spatial localization and semantic guidance capabilities. When the degradation type is clear, the focus can be on the channel; when the degradation distribution is concentrated, the focus can be on the space, achieving adaptive weight allocation. This improves the responsiveness to complex degradation patterns, ensuring that the enhancement results retain structural integrity while enhancing semantic distinguishability, providing a high-fidelity, low-interference feature foundation for subsequent perception tasks.
[0111] In the above embodiments of this application, the composite feature representation includes task-shared features and / or task-specific features for multiple image enhancement tasks; using a node learning module and multiple image enhancement tasks, feature modulation processing is performed on the weighted feature map to obtain the composite feature representation, including: determining an inter-task collaboration index between any two image enhancement tasks in the multiple image enhancement tasks to obtain at least one inter-task collaboration index; if, among the at least one inter-task collaboration index, there is an inter-task collaboration index greater than a preset collaboration threshold, a shared network is invoked to perform feature modulation processing on the weighted feature map to obtain inter-task shared features; if, among the at least one inter-task collaboration index, there is an inter-task collaboration index less than or equal to the preset collaboration threshold, task adapters corresponding to the two image enhancement tasks are invoked to perform feature modulation processing on the weighted feature map to obtain task-specific features.
[0112] At least one of the aforementioned inter-task collaboration metrics can refer to a numerical metric that quantifies the similarity or correlation of feature representations or gradient update directions between multiple image enhancement tasks, obtained through learning or computation, and is used to measure whether two image enhancement tasks have complementary or conflicting knowledge in the feature space.
[0113] The aforementioned preset collaboration threshold can refer to an empirical or data-driven numerical boundary set during the training phase, used to determine whether the degree of collaboration between two image enhancement tasks is high enough, thereby deciding whether to enable the shared processing mechanism.
[0114] The aforementioned shared network refers to a set of lightweight parameterized feature processing units shared by multiple highly collaborative image enhancement tasks within the node learning module. It can consist of a set of convolutional layers, normalization layers, and nonlinear activation functions. For example, a shared network might learn a general pattern for enhancing the contrast of dark area edges, a pattern that simultaneously benefits low-light enhancement and denoising. The advantages of shared networks include parameter sharing, computational reuse, avoidance of redundancy, and reduced model computational overhead when tasks are highly collaborative.
[0115] The aforementioned inter-task shared features refer to high-level feature representations that are jointly utilized by multiple image enhancement tasks after the weighted feature maps are processed through a shared network. These inter-task shared features can contain general semantic information that is effective for multiple image enhancement tasks, such as structural edge features commonly relied upon by low-light and denoising tasks, and global brightness distribution features commonly relied upon by dehazing and low-light tasks.
[0116] The aforementioned task adapters can refer to lightweight feature modulation modules designed separately for a specific image enhancement task, existing in parallel with the shared network. Each task adapter can be a set of small convolutional layers, conditional normalization layers, gating units, or residual blocks, with parameters independent of other tasks, dedicated to learning feature transformation strategies unique to that task. For example, a dehazing task adapter learns how to model atmospheric scattering and recover the suppressed blue channel; a denoising task adapter learns how to preserve high-frequency textures while suppressing Gaussian noise patterns.
[0117] The aforementioned task-specific features refer to highly task-specific feature representations that serve a single image enhancement task after the weighted feature map is independently modulated by a task adapter. These task-specific features preserve the task-specific recovery mechanisms and semantic priors. For example, dehazing task-specific features include structures such as aerosol density estimation and color shift correction; low-light enhancement task-specific features include functions such as nonlinear brightness mapping and local contrast stretching.
[0118] As an optional implementation, when the original image is from a nighttime environment without streetlights, identifying image illumination compensation and image denoising are the current tasks, and calculating the inter-task collaboration index between image illumination compensation and image denoising. The node learning module determines that the tasks are highly collaborative, both requiring enhancement of dark area structure and suppression of high-frequency noise. Therefore, a shared network is invoked to uniformly modulate the weighted feature map. Through a set of shared convolutional layers and conditional normalization modules, brightness stretching and noise suppression are completed, outputting inter-task shared features. These inter-task shared features preserve the global structure and texture consistency of low-light areas, avoid redundant calculations in multiple branches, reduce computational overhead, and ensure that the enhancement results are highly consistent in brightness and noise control, providing a clean and unified intermediate representation for subsequent multi-scale fusion.
[0119] As an alternative implementation, when the original image is a rainy / foggy nighttime scene, the task-to-task coordination metrics between image dehazing and image denoising are identified. Image dehazing requires restoring the blue channel obscured by fog and global contrast, while image denoising requires filtering out high-frequency particles, creating a conflict between the goals of image dehazing and image denoising. The node learning module activates independent task adapters: the dehazing adapter is a lightweight channel attention network that enhances the blue channel and corrects color cast; the denoising adapter is a residual small convolutional block that locally suppresses texture noise. Different task adapters process the weighted feature maps separately to generate task-specific features. Dehazing features preserve distant contours, while denoising features preserve near details. Finally, these features are fused in a multi-scale module to avoid negative transfer problems caused by forced uniformity in shared networks, such as blurred edges in dehazing or loss of fog area structure in denoising.
[0120] The above settings enable the automatic selection of shared or independent processing paths based on the relationships between tasks, achieving an intelligent balance between resource efficiency and accuracy. Highly collaborative tasks share computation, improving efficiency and consistency; low-collaboration tasks are processed independently, avoiding feature conflicts and performance degradation. This enhances generalization ability and robustness under complex degradation combinations, reduces the computational burden on vehicles, and ensures the accuracy of critical perception tasks. It also helps to achieve multi-task adaptive enhancement, providing technical support for highly reliable visual perception in real and complex nighttime scenes.
[0121] Figure 2 This is a schematic diagram of an image processing system architecture according to an embodiment of the present invention, such as... Figure 2 As shown, it includes an image acquisition device, an image enhancement model, and a feedback adjustment module. The image acquisition device is used to acquire original images based on an image sensor and a vehicle supplementary lighting unit. The image enhancement model performs image enhancement processing based on multiple image enhancement tasks to obtain a target image. The feedback adjustment module obtains the confidence level of the task execution results based on image quality indicators and the execution of image application tasks, and obtains the device adjustment parameters corresponding to the original image to perform feedback adjustment on the image acquisition device.
[0122] Figure 3 This is a schematic diagram illustrating an image enhancement process for an original image according to an embodiment of the present invention, such as... Figure 3 As shown, the self-attention module in the target encoder is used to perform target dimension attention processing on the original image to obtain a weighted feature map; the node learning module and multiple image enhancement tasks in the target encoder are used to perform feature modulation processing on the weighted feature map to obtain a composite feature representation; the multi-scale feature extraction module in the target encoder is used to perform multi-scale feature extraction on the composite feature representation to obtain an intermediate feature representation; and the target decoder is used to perform feature reconstruction on the intermediate feature representation to obtain the target image.
[0123] Figure 4 This is a schematic diagram of a process for generating a composite feature representation according to an embodiment of the present invention, such as... Figure 4 As shown, an inter-task collaboration index is determined between any two image enhancement tasks in a plurality of image enhancement tasks, resulting in at least one inter-task collaboration index. If, among the at least one inter-task collaboration index, there exists an inter-task collaboration index greater than a preset collaboration threshold, a shared network is invoked to perform feature modulation processing on the weighted feature map to obtain inter-task shared features. If, among the at least one inter-task collaboration index, there exists an inter-task collaboration index less than or equal to the preset collaboration threshold, task adapters corresponding to the two image enhancement tasks are invoked to perform feature modulation processing on the weighted feature map to obtain task-specific features. A composite feature representation is determined based on the inter-task shared features and / or task-specific features.
[0124] Figure 5 This is a schematic diagram of an image processing procedure according to an embodiment of the present invention, such as... Figure 5 As shown, the system acquires vehicle driving status information and environmental status information of the vehicle's environment; based on the driving status information and / or environmental status information, initial control parameters are determined; in the absence of historical images, the image acquisition device is controlled based on the initial control parameters to obtain the original image; in the presence of historical images, the initial control parameters are adjusted based on the device adjustment parameters corresponding to the historical images to obtain target control parameters, and the image acquisition device is controlled based on the target control parameters to obtain the original image. Using a target encoder and multiple image enhancement tasks, image enhancement processing is performed on the original image to obtain intermediate feature representations. These multiple image enhancement tasks include at least two of the following: image denoising, image illumination compensation, and image dehazing; using a target decoder, feature reconstruction is performed on the intermediate feature representations to obtain the target image. The image quality index of the target image is determined; image application tasks are executed based on the target image to obtain task execution results; based on the image quality index and the confidence level of the task execution results, the device adjustment parameters corresponding to the original image are determined.
[0125] The technical solution proposed in this application is described below with reference to an optional embodiment. This application proposes an intelligent driving night vision enhancement device, system, and method. It relates to the field of intelligent vehicle autonomous driving technology, specifically to performance improvement technology for vehicle perception systems, and proposes a device, system, and method for enhancing the perception capabilities of autonomous driving systems in nighttime and low-light environments. This application aims to achieve all-weather, high dynamic range, and high-definition vehicle visual perception through closed-loop collaboration of intelligent hardware acquisition improvement, multi-knowledge intelligent enhancement, and intelligent lighting supplementation strategies, at least solving the following technical problems: improving image quality in nighttime and low-light environments; achieving a unified solution for processing different environmental conditions, solving the problem of single verification scenarios; utilizing deep parallel separable convolution and feature fusion mechanisms to achieve multi-receptive field enhancement; and maximizing the signal strength of moving objects.
[0126] The intelligent driving night vision enhancement system proposed in this application includes: a front-end hardware acquisition layer, a back-end algorithm improvement layer, and an application intelligent feedback layer. The front-end hardware acquisition layer is configured to dynamically adjust the exposure parameters and gain parameters of the image sensor and control the illumination of the automotive-grade supplementary lighting through an intelligent imaging controller to acquire improved raw image data from a physical level. The back-end algorithm improvement layer is configured to receive the raw image data and input it into a multi-knowledge-oriented image enhancement model for processing to output an improved high-definition image. The application intelligent feedback layer is configured to analyze the imaging quality and generate control parameters based on the improved high-definition image or the perception results obtained from upper-level visual applications based on it, and feed these parameters back to the intelligent imaging controller to dynamically adjust the front-end acquisition strategy and the supplementary lighting control signal, forming an improved closed loop of acquisition, enhancement, and feedback. The intelligent imaging controller is configured to dynamically decide and control the on / off state, brightness, color temperature, and beam mode of the vehicle's headlights and / or taillights based on at least one of ambient light sensor data, vehicle status signals, quality assessment results output by the back-end image enhancement model, and scene analysis results, to coordinate the target image acquisition quality and meet driving safety regulations. The backend algorithm improvement layer uses a deep learning-based network for a multi-knowledge-oriented image enhancement model, comprising: a backbone network based on an encoder and decoder; a task-oriented node learning module for handling at least three image restoration tasks: image denoising, low-light image enhancement, and nighttime image dehazing; a self-attention module embedded in the backbone network to enhance weighting of important image regions; and a multi-scale feature extraction and fusion module to aggregate feature information from different levels. The task-oriented node learning module employs a structure combining shared and separate networks to extract shared features across tasks and task-specific features. The multi-knowledge-oriented image enhancement model also includes: a self-attention module configured to enhance the network's focus on key image regions; and a multi-scale feature extraction module configured to fuse image features from different receptive fields to improve detail recovery. The image enhancement model is jointly improved through a hybrid loss function, which includes at least a loss term constraining image fidelity and a loss term constraining feature consistency.
[0127] The image enhancement method proposed in this application is applied to the aforementioned intelligent driving night vision enhancement system. The method includes: acquiring original images by adaptively controlling an image sensor and a supplementary light using an intelligent imaging controller; inputting the original images into a multi-knowledge-oriented image enhancement model for processing, wherein the image enhancement model uses a task-oriented node learning mechanism and integrates self-attention and multi-scale feature extraction to simultaneously handle multiple image degradation scenarios; outputting a target image processed by the enhancer; and generating adjustment instructions for the front-end acquisition parameters based on the target image or derived perception results, and feeding these instructions back to the intelligent imaging controller to intelligently control exposure or gain and the automotive-grade supplementary light to improve subsequent image acquisition. The training process of the multi-knowledge-oriented image enhancement model is supervised by a hybrid loss function, which jointly considers the reconstruction quality of the target image and the consistency of the feature layer.
[0128] The overall architecture of the intelligent driving night vision enhancement system includes: a front-end hardware acquisition layer, a back-end algorithm improvement layer, and an application intelligent feedback layer. The front-end hardware acquisition layer includes image sensors, such as automotive-grade image sensors, controllable lighting units, integrated headlight and taillight control, and an intelligent imaging controller. The intelligent imaging controller is configured to dynamically and collaboratively adjust the sensor's exposure time, analog and digital gain, and optionally activate multi-frame high dynamic range imaging modes based on signals from the ambient light sensor, vehicle controller area network bus, vehicle speed, steering angle, gear position, and instructions from the feedback layer. Simultaneously, it controls the brightness, color temperature, and beam pattern of the lighting units, such as the area illumination of matrix headlights. The goal of this process is to maximize the capture of effective light signals from the scene at the physical level and suppress noise, glare, and other interference, outputting improved raw image data. The back-end algorithm improvement layer is a knowledge-based image enhancement model. It receives raw image data and processes composite tasks such as image denoising, low-light enhancement, and nighttime defogging in parallel or selectively through a unified and dynamic neural network architecture, outputting high-definition images with rich details and true colors. The application employs an intelligent feedback layer, comprising upper-layer vision application modules such as object detection and lane recognition networks, and a perception quality analyzer. The perception quality analyzer generates a quantitative assessment based on objective quality metrics of the enhanced image, such as sharpness, contrast-to-noise ratio, or the confidence level of the upper-layer application results. This quantitative assessment is fed back to the intelligent imaging controller in real time to guide adjustments to acquisition parameters and lighting strategies for the next frame or subsequent scenes, thus forming a continuously improving intelligent closed loop of perception, acquisition, and enhancement.
[0129] This image enhancement model, designed for multi-knowledge processing, employs an encoder-decoder backbone network and integrates the following modules: a unified and dynamic base network. The encoder incorporates a dynamic computation mechanism for content and task awareness, specifically an alternating spatial-channel attention module. Spatial attention automatically identifies severely degraded regions in the image, such as dense fog or extremely dark shadows, and assigns higher feature computation weights to these regions. Channel attention dynamically adjusts the importance of each feature channel based on the current task focus, such as prioritizing color channels for dehazing and texture channels for denoising. This mechanism allows the network to adaptively concentrate limited computational resources on the spatiotemporal feature dimensions most in need of restoration, providing high-quality shared foundational features for subsequent processing. A task relationship modeling and management module introduces a learnable task relationship graph to address knowledge sharing and conflict issues among multiple tasks. During training, the network learns the correlation between loss gradients for different tasks (denoising, low-light enhancement, nighttime dehazing) or utilizes prior knowledge to construct a relationship matrix. This matrix quantifies the degree of collaboration (positive weights) or conflict (negative weights) between tasks. During inference, this relation matrix guides a dynamic feature routing controller. For highly collaborative task pairs, the feature flow is directed to a shared common path; for tasks with potential conflicts, the feature flow is partially isolated by generating task-specific gating signals or activating dedicated parameterized adapters, thus protecting task specificity and avoiding negative transfer while achieving knowledge transfer. The task-oriented node learning module, built on a shared dynamic base, provides lightweight, task-specific adapters for each task. These adapters can be viewed as a set of task-specific feature modulation layers, such as conditional normalization layers or small convolutional groups. Under the scheduling of the dynamic routing controller, the adapter corresponding to the current task is primarily activated, while other adapters are suppressed. This design ensures high task specificity of high-level semantic features and restoration functions, achieving broad sharing of low-level features and precise, specific high-level decisions. A multi-scale feature extraction and fusion module is embedded in the encoder and skip connections. This module simultaneously captures information at different scales in the image, such as small noise points, object edge contours, and global scene structure, by using dilated convolutions with different dilation rates in parallel or employing a feature pyramid structure. These multi-scale features are progressively fused in the decoder to ensure that the output image retains rich texture details while removing large areas of haze.
[0130] The workflow of this application's method is as follows: Scene perception and parameter pre-adjustment: When the vehicle enters a tunnel, the ambient light suddenly drops. The intelligent imaging controller learns the vehicle speed from the controller area network bus and, based on historical feedback, decides to adopt a short exposure and medium-high gain strategy to balance suppressing motion blur and ensuring sufficient light intake, adjusting the headlights to low beam mode. Image acquisition and task analysis: The sensor acquires a frame of raw image according to instructions. Based on positioning, time information, or rapid analysis of the raw image, the primary task is determined to be low-light enhancement, and the secondary associated task is image denoising. Dynamic network inference: The raw image is input into the image enhancement model module. The dynamic base network focuses on extracting features from dark areas; the task relationship management module guides the main feature flow to the shared path based on the high synergy between low-light enhancement and image denoising; the corresponding low-light enhancement adapter in the node learning module is strongly activated, performing specific corrections to color and contrast; the multi-scale module ensures that dark details are preserved. Result output and feedback: A bright, clear, and low-noise enhanced image is output for autonomous driving decision-making. Simultaneously, the perception quality analyzer calculates the signal-to-noise ratio and average gradient of the frame image. If the analysis reveals significant noise in the dark areas and a low signal-to-noise ratio, a feedback instruction is generated that allows for a slight increase in exposure time and a moderate reduction in gain within permissible limits. This instruction is then sent to the controller to improve the acquisition strategy for the next frame.
[0131] Task input can include preset rules, automatically determined based on other sensors such as ambient light sensors, time, and location. For example, if the sky is dark and humidity is high, it is determined to be a nighttime defogging task. Users or upper-layer applications can specify and manually select modes, or higher-layer perception modules, such as detectors, can infer modes based on requirements. For example, if the object detection module finds that the image is blurry and has low confidence, it requests defogging enhancement. Preliminary network analysis involves a rapid pre-analysis of the input image to determine the main degradation type. Shared feature input consists of common basic features extracted by a dynamic pedestal network, which have not undergone task-specific processing. The input image is processed through a shared encoder or feature extraction backbone network. The feature extraction backbone network is task-independent, aiming to extract basic information from the image that may be useful for subsequent tasks, such as geometric structures like edges and contours, basic appearance information like color and texture, and contextual information captured through multi-scale receptive fields. For training and improvement, the network used in the image enhancement model employs a hybrid loss function for end-to-end training. This function includes: L1 pixel loss to ensure overall fidelity; perceptual loss to ensure semantic authenticity; task-specific loss, such as prior loss for noise distribution for denoising tasks, to enhance the specialization of each node; and feature consistency loss to encourage shared features to maintain stable expression.
[0132] Through the aforementioned closed-loop design of deep hardware and software collaboration, this application effectively overcomes the limitations of a single technical path and enhances the night vision capability of vehicle-mounted vision systems in complex dynamic environments. The intelligent driving night vision enhancement device, system, and method proposed in this application achieve end-to-end image quality improvement from physical acquisition to algorithm processing. Specifically, this is reflected in the following aspects: End-to-end image quality improvement: Front-end hardware adaptively adjusts exposure, gain, and automotive-grade supplementary lighting to maximize the capture of effective signals and suppress noise, glare, and other interference during the acquisition stage. The back-end then utilizes a multi-knowledge-guided image enhancement model to perform deep restoration and fusion of complex degraded scenes, such as nighttime fog and low light, thereby achieving a clearer, more detailed, and lower-noise imaging effect overall. Strong adaptability to complex environments: Collaborative improvements have been made for various harsh imaging conditions such as nighttime, fog, and low light. Front-end hardware adaptive acquisition provides higher-quality raw data for the back-end algorithm, and the back-end enhancement network further performs targeted repair and fusion, enabling stable and high-quality image output even in various complex environments. A smart closed loop is formed, enabling adaptive improvement. A feedback loop is constructed from backend perception results to frontend acquisition parameters. The acquisition strategy can be dynamically adjusted based on the actual scenario and enhancement effect, achieving continuous improvement through acquisition, enhancement, feedback, and re-acquisition, further enhancing autonomous adaptability and overall stability in different scenarios. Hardware and software collaboration balances real-time performance and efficiency. Frontend hardware control ensures the quality foundation of the original signal, reducing the recovery difficulty of the backend algorithm; backend intelligent algorithm enhancement further taps the potential of image information. This collaborative approach improves the visual quality and usability of the final image while maintaining processing efficiency, making it suitable for intelligent driving scenarios that prioritize both real-time performance and quality. The system exhibits good scalability and versatility. The architecture is not dependent on specific hardware models or fixed scenarios. The adaptive acquisition strategy and multi-knowledge enhancement network can be configured and migrated according to different application needs, possessing strong scalability and industry applicability.
[0133] According to another aspect of the embodiments of this application, an image processing apparatus is also provided, which can execute the image processing method of the above embodiments. The specific implementation method and preferred application scenarios are the same as those of the above embodiments, and will not be repeated here.
[0134] Figure 6 This is a schematic diagram of an image processing apparatus according to an embodiment of this application, such as... Figure 6 As shown, the device includes the following: an acquisition module 602 and an image enhancement processing module 604.
[0135] The acquisition module 602 is used to control the vehicle's image acquisition device to acquire images of the vehicle's surrounding environment based on initial control parameters or target control parameters to obtain an original image. The target control parameters are obtained by adjusting the initial control parameters based on the device adjustment parameters corresponding to historical images. The image enhancement processing module 604 is used to perform image enhancement processing on the original image based on an image enhancement model to obtain a target image.
[0136] The acquisition module is also used to determine the image quality index of the target image; and to determine the device adjustment parameters corresponding to the original image based on the image quality index. The device adjustment parameters corresponding to the original image are used to determine the target control parameters to be used when acquiring the next original image.
[0137] The acquisition module is also used to perform image application tasks based on the target image and obtain task execution results. The image application task is used to represent visual detection and recognition tasks performed based on the target image. Based on the image quality index and the confidence level of the task execution results, the device adjustment parameters corresponding to the original image are determined.
[0138] The image acquisition device includes at least: an image sensor and a vehicle lighting unit; the acquisition module is also used to control the image acquisition device based on initial control parameters to obtain the original image when no historical image exists; and to adjust the initial control parameters based on the device adjustment parameters corresponding to the historical image to obtain target control parameters, and to control the image acquisition device based on the target control parameters to obtain the original image.
[0139] The acquisition module is also used to acquire vehicle driving status information and environmental status information of the vehicle's environment; and to determine initial control parameters based on the driving status information and / or environmental status information.
[0140] The image enhancement model includes a target encoder and a target decoder; the image enhancement processing module is also used to perform image enhancement processing on the original image using the target encoder and multiple image enhancement tasks to obtain intermediate feature representations. The multiple image enhancement tasks include at least two of the following: image denoising, image illumination compensation, and image dehazing; the target decoder is used to perform feature reconstruction on the intermediate feature representations to obtain the target image.
[0141] The target encoder includes a self-attention module, a node learning module, and a multi-scale feature extraction module. The image enhancement processing module is also used to apply attention processing to the target dimension of the original image using the self-attention module to obtain a weighted feature map. The node learning module and multiple image enhancement tasks are used to perform feature modulation processing on the weighted feature map to obtain a composite feature representation. The multi-scale feature extraction module is used to extract multi-scale features from the composite feature representation to obtain an intermediate feature representation.
[0142] The image enhancement processing module is further used to perform spatial dimension attention processing on the degraded regions in the original image to obtain a first feature map; to perform channel dimension attention processing on the target image channels of the original image to obtain a second feature map, wherein the target image channels are determined based on multiple image enhancement tasks; and to obtain a weighted feature map based on the first feature map and / or the second feature map.
[0143] The composite feature representation includes task-shared features and / or task-specific features for multiple image enhancement tasks. The image enhancement processing module is further configured to determine the task-to-task coordination index between any two image enhancement tasks to obtain at least one task-to-task coordination index. If, among the at least one task-to-task coordination index, there exists a task-to-task coordination index greater than a preset coordination threshold, a shared network is invoked to perform feature modulation processing on the weighted feature map to obtain task-shared features. If, among the at least one task-to-task coordination index, there exists a task-to-task coordination index less than or equal to the preset coordination threshold, task adapters corresponding to the two image enhancement tasks are invoked to perform feature modulation processing on the weighted feature map to obtain task-specific features.
[0144] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0145] According to another aspect of the embodiments of this application, a vehicle is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.
[0146] The aforementioned memory can refer to devices inside a computer used to store data and programs, including RAM, hard disks, etc. RAM can be used to temporarily store running programs and data, while hard disks can be used to store programs and data long-term. Memory enables the computer to read and write data and execute programs. The aforementioned processor is responsible for executing instructions in computer programs and performing data processing. It can also be responsible for controlling and executing various operations, including arithmetic operations, logical operations, and data transmission.
[0147] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.
[0148] The aforementioned computer storage media can refer to the media used in computer memory to store certain discontinuous physical quantities. Computer storage media mainly include semiconductors, magnetic cores, magnetic drums, magnetic tapes, laser discs, etc. Computer-readable storage media include stored programs, which can be a set of instructions that a computer can recognize and execute, running on an electronic computer to meet certain information needs.
[0149] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0150] The aforementioned computer program products can refer to software programs that have been written, tested, and released, and can run on computers or other devices. Computer program products can include application programs, operating systems, utility software, etc., used to achieve specific functions or solve specific problems.
[0151] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the methods in various embodiments of this application.
[0152] The aforementioned non-volatile computer-readable storage medium can refer to a medium for storing data. Non-volatile computer-readable storage media can retain data without loss when power is off and can be used to store long-term data, such as operating systems, applications, and user files. Non-volatile storage media can include hard disk drives, solid-state drives, optical disks, and flash memory storage devices, etc.
[0153] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.
[0154] The aforementioned computer program can refer to a set of instructions used to tell the computer to perform specific tasks or operations. Computer programs can be written by programmers using specific programming languages and can include algorithms, data structures, logic, and control flow. Computer programs can be used for a variety of purposes, including application software, operating systems, etc.
[0155] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0156] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0157] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0158] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0159] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0160] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. An image processing method, characterized in that, include: Based on initial control parameters or target control parameters, the vehicle's image acquisition device is controlled to acquire images of the vehicle's surrounding environment to obtain original images. The target control parameters are obtained by adjusting the initial control parameters based on the device adjustment parameters corresponding to historical images. The original image is enhanced using an image enhancement model to obtain the target image.
2. The image processing method according to claim 1, characterized in that, The method further includes: Determine the image quality index of the target image; Based on the image quality index, the device adjustment parameters corresponding to the original image are determined, wherein the device adjustment parameters corresponding to the original image are used to determine the target control parameters to be used when acquiring the next original image.
3. The image processing method according to claim 2, characterized in that, Determining the device adjustment parameters corresponding to the original image based on the image quality index includes: An image application task is performed based on the target image to obtain a task execution result, wherein the image application task is used to represent a visual detection and recognition task performed based on the target image; Based on the image quality index and the confidence level of the task execution result, the device adjustment parameters corresponding to the original image are determined.
4. The image processing method according to claim 1, characterized in that, The image acquisition device includes at least: an image sensor and a vehicle lighting unit; based on initial control parameters or target control parameters, the image acquisition device controls the vehicle to acquire images of the vehicle's surrounding environment to obtain raw images, including: In the absence of the historical image, the image acquisition device is controlled based on the initial control parameters to obtain the original image; If the historical image exists, the initial control parameters are adjusted based on the device adjustment parameters corresponding to the historical image to obtain the target control parameters, and the image acquisition device is controlled based on the target control parameters to obtain the original image.
5. The image processing method according to claim 4, characterized in that, The method further includes: Obtain the vehicle's driving status information and the environmental status information of the vehicle's surroundings; The initial control parameters are determined based on the driving condition information and / or the environmental condition information.
6. The image processing method according to any one of claims 1 to 5, characterized in that, The image enhancement model includes a target encoder and a target decoder; The original image is enhanced using an image enhancement model to obtain the target image, including: The original image is enhanced using a target encoder and multiple image enhancement tasks to obtain an intermediate feature representation. The multiple image enhancement tasks include at least two of the following: image denoising, image illumination compensation, and image dehazing. The target image is obtained by reconstructing the intermediate feature representation using the target decoder.
7. The image processing method according to claim 6, characterized in that, The target encoder includes a self-attention module, a node learning module, and a multi-scale feature extraction module; The original image is enhanced using a target encoder and multiple image enhancement tasks to obtain intermediate feature representations, including: Using the self-attention module, the original image is subjected to target dimension attention processing to obtain a weighted feature map; Using the node learning module and the multiple image enhancement tasks, the weighted feature map is subjected to feature modulation processing to obtain a composite feature representation; The multi-scale feature extraction module is used to extract multi-scale features from the composite feature representation to obtain the intermediate feature representation.
8. The image processing method according to claim 7, characterized in that, Using the self-attention module, the original image undergoes target-dimensional attention processing to obtain a weighted feature map, including: Spatial attention processing is performed on the degraded regions in the original image to obtain a first feature map; Attention processing along the channel dimension is performed on the target image channels of the original image to obtain a second feature map, wherein the target image channels are determined based on the multiple image enhancement tasks; The weighted feature map is obtained based on the first feature map and / or the second feature map.
9. The image processing method according to claim 7, characterized in that, The composite feature representation includes task-shared features and / or task-specific features for the multiple image enhancement tasks; Using the node learning module and the multiple image enhancement tasks, feature modulation processing is performed on the weighted feature map to obtain a composite feature representation, including: Determine the inter-task coordination index between any two image enhancement tasks among the plurality of image enhancement tasks to obtain at least one inter-task coordination index. If any of the at least one inter-task collaboration indicators is greater than a preset collaboration threshold, the shared network is invoked to perform feature modulation processing on the weighted feature map to obtain the inter-task shared features. If, among the at least one inter-task collaboration index, there exists an inter-task collaboration index that is less than or equal to the preset collaboration threshold, the task adapters corresponding to the two image enhancement tasks are invoked to perform feature modulation processing on the weighted feature map to obtain the task-specific features.
10. An image processing apparatus, characterized in that, include: The acquisition module is used to control the vehicle's image acquisition device to acquire images of the vehicle's surrounding environment based on initial control parameters or target control parameters, thereby obtaining original images. The target control parameters are obtained by adjusting the initial control parameters based on the device adjustment parameters corresponding to historical images. The image enhancement processing module is used to perform image enhancement processing on the original image based on the image enhancement model to obtain the target image.
11. A vehicle, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the image processing method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the storage medium is located to perform the image processing method according to any one of claims 1 to 9.