Image processing method, device and vehicle for vehicle
By deploying low-resolution infrared cameras on vehicles and using super-resolution neural network models for image processing, the problem of insufficient clarity and contrast of infrared cameras is solved, achieving high-quality image output in low-visibility environments, improving visual perception capabilities and driving safety, and making it suitable for complex scenarios of intelligent driving.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BYD CO LTD
- Filing Date
- 2026-03-16
- Publication Date
- 2026-06-05
AI Technical Summary
In existing technologies, low-resolution infrared cameras in high-end vehicles suffer from insufficient clarity and contrast in night vision perception, resulting in image quality that fails to meet requirements. Furthermore, existing contrast adjustment methods lead to the loss of highlight and shadow details, which are highly limited and applicable to a limited range of image types and scenarios.
By deploying a single low-resolution infrared camera on a vehicle, a super-resolution neural network model is used to process the acquired low-quality infrared images, including temporal alignment, spatial registration, and distortion correction, to train an infrared super-resolution model that outputs high-resolution, high-definition, and high-contrast infrared images.
Without significantly increasing hardware costs, it significantly improves visual perception capabilities and driving safety in low-visibility environments such as nighttime, fog, rain, and snow. It is applicable to various image types and scenarios, meeting the perception needs of intelligent driving in complex scenarios.
Smart Images

Figure CN122160603A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of vehicle technology, and in particular relates to an image processing method, apparatus and vehicle for vehicles. Background Technology
[0002] In recent years, with the rapid development of new energy vehicles, some high-end models have begun to be equipped with infrared cameras to achieve night vision perception. However, due to cost issues, most of them are equipped with low-resolution infrared cameras, and the clarity and contrast cannot meet the requirements. In order to improve the quality of infrared imaging, there are methods to adjust the contrast of infrared images using S-curves. However, this will lead to the loss of highlight and shadow details, and the overall brightness will be reduced. It lacks local adaptability, which will result in the limitation of contrast adjustment and the applicable image types and scenes are limited. Summary of the Invention
[0003] This invention aims to solve at least one of the technical problems existing in the prior art. To this end, this invention proposes an image processing method, apparatus, and vehicle for vehicles, which significantly improves visual perception capabilities and driving safety in low-visibility environments such as nighttime, fog, rain, and snow. It achieves the goal of compensating for hardware limitations with algorithm performance, is applicable to various image types and scenarios, can meet the perception needs of intelligent driving in complex scenarios, and has high application value.
[0004] In a first aspect, this application provides an image processing method for a vehicle, including a first infrared camera, the method comprising: Based on the first infrared camera, a first infrared image of the scene in which the vehicle is located is acquired; The first infrared image is input into the infrared super-resolution model to obtain the second infrared image output by the infrared super-resolution model. The image quality of the second infrared image is higher than that of the first infrared image. The second infrared image is displayed.
[0005] The image processing method for vehicles provided in this application can acquire environmental images by deploying a single low-resolution infrared camera on the vehicle and use a super-resolution neural network model to process the acquired low-quality infrared images in real time. This method can output infrared images with higher image quality, such as higher resolution, higher clarity, and higher contrast, without significantly increasing the cost of onboard hardware (only requiring the retention of a single low-cost infrared camera). This significantly improves visual perception capabilities and driving safety in low-visibility environments such as night, fog, rain, and snow, achieving the goal of compensating for hardware limitations with algorithm performance. It is applicable to various image types and scenarios, can meet the perception needs of intelligent driving in complex scenarios, and has high application value.
[0006] One embodiment of this application provides an image processing method for vehicles, wherein the image quality includes at least one of resolution, contrast, and sharpness; and / or, The first infrared camera is a low-resolution camera; and / or, The resolution of the first infrared camera is less than or equal to 650x512.
[0007] An embodiment of the image processing method for vehicles in this application, wherein the infrared super-resolution model is trained based on a low-resolution infrared image sequence corresponding to a sample scene acquired by a first sample camera and a high-resolution infrared image sequence corresponding to the sample scene acquired by a second sample camera. The high-resolution infrared image sequence has a higher resolution than the low-resolution infrared image sequence; and / or, The first sample camera is a low-resolution camera with a resolution of 650x512 or less, and the second sample camera is a high-resolution camera with a resolution greater than 650x512.
[0008] One embodiment of the image processing method for vehicles in this application, wherein the infrared super-resolution model is trained based on the following steps: The low-resolution infrared image sequence and the high-resolution infrared image sequence are registered to obtain training sample pairs; the training sample pairs include low-resolution infrared images and high-resolution infrared images. Using the low-resolution infrared images in the training sample pair as samples and the high-resolution infrared images corresponding to the low-resolution infrared images as sample labels, the convolutional neural network model is trained to obtain the infrared super-resolution model.
[0009] An embodiment of the image processing method for vehicles according to this application includes registering the low-resolution infrared image sequence and the high-resolution infrared image sequence to obtain training sample pairs, comprising: The low-resolution infrared image sequence and the high-resolution infrared image sequence are time-aligned. Spatial registration is performed on the time-aligned low-resolution infrared image sequence and the high-resolution infrared image sequence to obtain the training sample pair.
[0010] An embodiment of this application provides an image processing method for vehicles, wherein the time alignment processing of the low-resolution infrared image sequence and the high-resolution infrared image sequence includes: Acquire the motion state of the same moving target in the low-resolution infrared image sequence and the high-resolution infrared image sequence; Based on the temporal consistency of the motion state, the system time deviation between the first sample camera and the second sample camera is obtained; The low-resolution infrared image sequence and the high-resolution infrared image sequence are time-aligned based on the system time deviation; and / or, The spatial registration process performed on the time-aligned low-resolution infrared image sequence and the high-resolution infrared image sequence to obtain the training sample pair includes: Distortion correction processing is performed on the time-aligned low-resolution infrared image sequence and the high-resolution infrared image sequence; Feature points are obtained from distortion-corrected low-resolution infrared image sequences and high-resolution infrared image sequences, and multiple feature points are matched. Based on the matched feature points, the transformation relationship between the high-resolution infrared image sequence and the image space where the low-resolution infrared image sequence is located is determined. Based on the transformation relationship, the high-resolution infrared image sequence is spatially transformed to obtain a high-resolution infrared image sequence that is spatially aligned with the low-resolution infrared image sequence, thereby obtaining the training sample pair.
[0011] An embodiment of this application provides an image processing method for vehicles, wherein the distortion correction processing of the time-aligned low-resolution infrared image sequence and the high-resolution infrared image sequence includes: Obtain the camera intrinsic parameter matrix and distortion coefficients corresponding to the first sample camera and the second sample camera, respectively; Based on the camera intrinsic parameter matrix and distortion coefficients corresponding to the first sample camera, distortion correction processing is performed on the time-aligned low-resolution infrared image sequence; based on the camera intrinsic parameter matrix and distortion coefficients corresponding to the second sample camera, distortion correction processing is performed on the time-aligned high-resolution infrared image sequence.
[0012] An embodiment of this application provides an image processing method for vehicles, which, after spatial registration processing of a time-aligned low-resolution infrared image sequence and a high-resolution infrared image sequence, includes: The training sample pair is obtained by performing image enhancement processing on the spatially registered high-resolution infrared image sequence; or, the training sample pair is obtained by performing at least one of histogram equalization, contrast stretching, noise suppression, and image sharpening on the spatially registered high-resolution infrared image sequence.
[0013] An embodiment of the image processing method for vehicles according to this application includes obtaining a second infrared image output by the infrared super-resolution model, comprising: Receive the user's first input, which is used to determine the rate adjustment command; In response to the first input, the upsampling magnification parameter of the infrared super-resolution model is adjusted to obtain a second infrared image output by the infrared super-resolution model that corresponds to the magnification adjustment command.
[0014] One embodiment of this application provides an image processing method for vehicles, wherein the infrared super-resolution model is a lightweight model, comprising a shallow feature extraction module, a deep feature extraction module, and a feature reconstruction upsampling module connected in sequence; or... The infrared super-resolution model includes a shallow feature extraction module, a deep feature extraction module, and a feature recombination and upsampling module connected in sequence. The shallow feature extraction module processes the first infrared image to obtain shallow features, the deep feature extraction module processes the shallow features to obtain deep features, and the feature recombination and upsampling module performs feature recombination and fusion on the deep features to obtain the second infrared image. The deep feature extraction module includes a reparameterized residual module.
[0015] Secondly, this application provides an image processing apparatus for a vehicle, including a first infrared camera, the apparatus comprising: The first processing module is used to acquire a first infrared image of the scene where the vehicle is located based on the first infrared camera; The second processing module is used to input the first infrared image into the infrared super-resolution model and obtain the second infrared image output by the infrared super-resolution model. The image quality of the second infrared image is higher than that of the first infrared image. The third processing module is used to display the second infrared image.
[0016] The image processing device for vehicles provided in the embodiments of this application can acquire environmental images by deploying a single low-resolution infrared camera on the vehicle and use a super-resolution neural network model to process the acquired low-quality infrared images in real time. It can output infrared images with higher image quality, such as higher resolution, higher clarity and higher contrast, without significantly increasing the cost of on-board hardware (only requiring the retention of a single low-cost infrared camera). This significantly improves visual perception and driving safety in low-visibility environments such as night, fog, rain and snow, and achieves the goal of compensating for hardware limitations with algorithm performance. It is applicable to a variety of image types and scenarios, can meet the perception needs of intelligent driving in complex scenarios, and has high application value.
[0017] Thirdly, this application provides a vehicle, including: First infrared camera; Display device; A controller, which is connected to the first infrared camera and the display device respectively, is used to process a first infrared image acquired by the first infrared camera based on the image processing method for vehicles as described in the first aspect, to obtain a second infrared image, and to display the second infrared image on the display device; or, it includes the image processing device for vehicles as described in the second aspect.
[0018] According to the vehicle provided in the embodiments of this application, an environmental image can be acquired by a single low-resolution infrared camera deployed on the vehicle, and a super-resolution neural network model can be used to process the acquired low-quality infrared image in real time. This can output infrared images with higher image quality, such as higher resolution, higher clarity and higher contrast, without significantly increasing the cost of onboard hardware (only a single low-cost infrared camera needs to be retained). This significantly improves the visual perception capability and driving safety in low visibility environments such as night, fog, rain and snow, and achieves the goal of making up for hardware limitations with algorithm performance. It is applicable to a variety of image types and scenarios, can meet the perception needs of intelligent driving in complex scenarios, and has high application value.
[0019] One embodiment of the vehicle in this application includes: An input module, connected to the controller, is used to receive a first input from a user, the first input being used to determine a magnification adjustment command, the magnification adjustment command being used by the controller to acquire a second infrared image corresponding to the magnification adjustment command; and / or, A voice recognition module, connected to the controller, is used to receive user voice commands and process the voice commands into a magnification adjustment command. The magnification adjustment command is used by the controller to acquire a second infrared image corresponding to the magnification adjustment command; and / or, A physical interaction module, connected to the controller, is used to generate the magnification adjustment command based on the user's physical interaction operation. The magnification adjustment command is used by the controller to acquire a second infrared image corresponding to the magnification adjustment command; and / or, A roller mounted on the steering wheel of the vehicle is used to generate the magnification adjustment command based on the user's scrolling operation. The magnification adjustment command is used by the controller to acquire a second infrared image corresponding to the magnification adjustment command.
[0020] Fourthly, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the image processing method for a vehicle as described in the first aspect above.
[0021] Fifthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image processing method for a vehicle as described in the first aspect above.
[0022] The above-described one or more technical solutions in the embodiments of this application have at least one of the following technical effects: Environmental images can be acquired by a single low-resolution infrared camera deployed on a vehicle, and the acquired low-quality infrared images can be processed in real time using a super-resolution neural network model. This can output infrared images with higher image quality, such as higher resolution, higher clarity, and higher contrast, without significantly increasing the cost of onboard hardware (only requiring a single low-cost infrared camera). This significantly improves visual perception and driving safety in low-visibility environments such as night, fog, rain, and snow, achieving the goal of compensating for hardware limitations with algorithm performance. It is applicable to various image types and scenarios, can meet the perception needs of intelligent driving in complex scenarios, and has high application value.
[0023] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0024] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is one of the flowcharts illustrating the image processing method for vehicles provided in this application embodiment; Figure 2 This is a second schematic flowchart of the image processing method for vehicles provided in the embodiments of this application; Figure 3 This is the third schematic flowchart of the image processing method for vehicles provided in the embodiments of this application; Figure 4 This is a schematic diagram of the intrinsic parameter calibration data of the infrared camera in the image processing method for vehicles provided in this application embodiment; Figure 5 This is the fourth flowchart of the image processing method for vehicles provided in the embodiments of this application; Figure 6 This is one of the schematic diagrams illustrating the binarization of calibration data in the image processing method for vehicles provided in this application embodiment; Figure 7 This is the second schematic diagram of binarization processing of calibration data in the image processing method for vehicles provided in this application embodiment; Figure 8This is a schematic diagram of the circular target detection result in the image processing method for vehicles provided in this application embodiment; Figure 9 This is a schematic diagram comparing the effect of infrared image distortion correction in the image processing method for vehicles provided in this application embodiment; Figure 10 This is the fifth flowchart illustrating the image processing method for vehicles provided in this application embodiment; Figure 11 This is a schematic diagram of the feature points matched in the image processing method for vehicles provided in this application embodiment; Figure 12 This is the sixth schematic flowchart of the image processing method for vehicles provided in the embodiments of this application; Figure 13 This is a schematic diagram of the infrared super-resolution model structure of the image processing method for vehicles provided in this application embodiment; Figure 14 This is the seventh flowchart of the image processing method for vehicles provided in the embodiments of this application; Figure 15 This is the eighth schematic flowchart of the image processing method for vehicles provided in the embodiments of this application; Figure 16 This is a schematic diagram of the structure of the infrared super-resolution system provided in the embodiments of this application; Figure 17 This is one of the infrared super-resolution result comparison diagrams of the image processing method for vehicles provided in this application embodiment; Figure 18 This is the second schematic diagram comparing the infrared super-resolution results of the image processing method for vehicles provided in this application embodiment; Figure 19 This is a schematic diagram of the structure of an image processing device for a vehicle provided in an embodiment of this application; Figure 20 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0026] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0027] The following description, in conjunction with the accompanying drawings, details the image processing method for vehicles, image processing apparatus for vehicles, vehicles, electronic devices, and readable storage media provided in this application, through specific embodiments and application scenarios.
[0028] The image processing method for vehicles can be applied to a terminal, specifically executed by the hardware or software within the terminal.
[0029] The image processing method for vehicles provided in this application embodiment can be executed by an electronic device or a functional module or entity in an electronic device that can implement the image processing method for vehicles. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras, and wearable devices. The image processing method for vehicles provided in this application embodiment will be described below using an electronic device as the execution subject as an example.
[0030] like Figure 1 As shown, the image processing method for vehicles includes steps 110, 120 and 130.
[0031] It should be noted that the vehicle includes a first infrared camera, which is a device installed on the vehicle to sense and image the infrared radiation of the external environment. For example, it can be an uncooled microbolometer infrared camera, which is relatively inexpensive but has a lower resolution (e.g., an imaging resolution of 640×512 pixels). The first infrared camera can capture the heat emitted by objects themselves and can still generate images reflecting the temperature distribution of objects even in low-visibility conditions such as nighttime darkness, fog, rain, and snow.
[0032] The first infrared camera can be a low-resolution infrared camera; for example, the resolution of the first infrared camera is less than or equal to 650x512.
[0033] Step 110: Based on the first infrared camera, acquire the first infrared image of the scene where the vehicle is located; In this step, while the vehicle is in motion, the first infrared camera continuously images the areas of interest, such as the front, sides, or rear of the vehicle.
[0034] The raw infrared data collected by the first infrared camera is the first infrared image.
[0035] The first infrared image directly reflects the temperature distribution of the surfaces of various objects in the current scene, but due to the limitations of the camera's resolution, the image details are relatively blurry (i.e., the clarity is low) and the contrast may not be high.
[0036] Step 120: Input the first infrared image into the infrared super-resolution model to obtain the second infrared image output by the infrared super-resolution model; The image quality of the second infrared image is higher than that of the first infrared image.
[0037] In some embodiments, image quality includes at least one of resolution, contrast, and sharpness.
[0038] In this step, the infrared super-resolution model is a deep convolutional neural network model designed and trained specifically for the characteristics of infrared images. The infrared super-resolution model can reconstruct a corresponding image with higher resolution and more detail from a low-resolution input image.
[0039] Infrared super-resolution models can be deployed on automotive low-computing platforms, which can be high-performance system-on-chips (SoCs) suitable for advanced driver assistance systems (ADAS), autonomous driving (AD), and smart cockpit applications. These platforms integrate various processing units and interfaces, providing powerful computing capabilities, low power consumption, and high security.
[0040] The inventors' tests revealed that the vehicle-mounted low-computing-power platform can perform real-time super-resolution processing on 300,000-pixel infrared images with a computing power consumption of less than 2.7 TOPS.
[0041] The resolution of the second infrared image is greater than that of the first infrared image. For example, if the input is a first infrared image of 640×512 pixels, after model processing, the output second infrared image may reach 1280×1024 pixels or higher.
[0042] The contrast of the second infrared image is greater than that of the first infrared image, and the sharpness of the second infrared image is greater than that of the first infrared image.
[0043] Infrared super-resolution models can supplement the missing high-frequency details and textures in low-resolution images by learning the data features of a large number of real-world infrared images.
[0044] The infrared super-resolution model is trained using low-resolution infrared image sequences corresponding to sample scenes captured by a first sample camera and high-resolution infrared image sequences corresponding to sample scenes captured by a second sample camera. The resolution of the high-resolution infrared image sequences is higher than that of the low-resolution infrared image sequences.
[0045] For example, the low-resolution infrared image sequence is obtained by collecting images in various sample scenarios (such as urban roads, highways, rural night roads, and rainy / foggy weather) based on another infrared camera (i.e., the first sample camera, which can be regarded as a twin device) with the same model and performance as the first infrared camera on the vehicle.
[0046] In some embodiments, the first sample camera is a low-resolution camera with a resolution of 650x512 or less.
[0047] The high-resolution infrared image sequence was acquired synchronously with the low-resolution camera at the same time, location, and scene using another infrared camera with better performance and higher resolution (the second sample camera).
[0048] In some embodiments, the second sample camera is a high-resolution camera with a resolution greater than 650x512.
[0049] During the training of the infrared super-resolution model, the first and second sample cameras can be set to be close to each other. The relative positions of the cameras can be adjusted to ensure that the fields of view of the two sample cameras are basically consistent, so that there is no excessive offset or rotation of the fields of view of the two sample cameras. This avoids problems such as registration failure caused by large field of view deviation. Black tape can be used to fix the two infrared cameras to ensure that the cameras will not move relative to each other due to vibration, collision or other reasons during the data acquisition process.
[0050] Multiple training sample pairs can be constructed from synchronously acquired, content-aligned low-resolution and high-resolution images. The low-resolution images can be used as input to the neural network, while the corresponding high-resolution images serve as supervisory signals for the neural network to learn and approximate. The resulting trained model will then be capable of upscaling new, unseen low-resolution infrared images into high-resolution images.
[0051] Step 130: Display the second infrared image.
[0052] In this step, the processed second infrared image can be presented to the driver in real time through the vehicle's display screen (such as the central control screen, instrument panel, or head-up display).
[0053] The second infrared image is used for the driver's environmental observation and can assist the driver's visual perception. In situations where conventional visible light vision is limited (such as at night without streetlights or in dense fog), the enhanced, clear, and high-resolution second infrared image can help the driver detect potential hazards such as pedestrians, animals, vehicles, or obstacles earlier and more clearly, as well as identify road outlines and traffic signs, thereby making more timely and safer driving decisions.
[0054] In some embodiments, the second infrared image can also serve as a sensing input source for advanced driver assistance systems (ADAS), autonomous driving (AD) systems, or smart cockpits.
[0055] In this embodiment, for example, the second infrared image can be used as input to an onboard AI vision algorithm (such as YOLO (You Only Look Once) or Faster R-CNN, a target detector based on a convolutional neural network) for target detection and recognition, thereby improving the detection rate and recognition accuracy of the algorithm for key targets such as pedestrians, cyclists, animals, vehicles and obstacles.
[0056] The second infrared image can also enable algorithms to track multiple moving targets more stably and predict their future trajectories more accurately.
[0057] In extreme environments (such as when dense fog obscures lane lines), second infrared images can utilize the difference in thermal radiation between the road surface and the surrounding environment to help identify road boundaries and passable areas, providing supplementary information for vehicle positioning and route planning.
[0058] It can also be fused with the perception results from vehicle-mounted millimeter-wave radar, lidar, and visible light cameras. After fusion, the robustness, redundancy, and reliability of the entire perception system can be improved, enabling all-weather and all-climate environmental modeling.
[0059] The image processing method for vehicles provided in this application can acquire environmental images by deploying a single low-resolution infrared camera on the vehicle and use a super-resolution neural network model to process the acquired low-quality infrared images in real time. This method can output infrared images with higher image quality, such as higher resolution, higher clarity, and higher contrast, without significantly increasing the cost of onboard hardware (only requiring the retention of a single low-cost infrared camera). This significantly improves visual perception capabilities and driving safety in low-visibility environments such as night, fog, rain, and snow, achieving the goal of compensating for hardware limitations with algorithm performance. It is applicable to various image types and scenarios, can meet the perception needs of intelligent driving in complex scenarios, and has high application value.
[0060] In some embodiments, the infrared super-resolution model is trained based on the following steps: Registration processing is performed on low-resolution infrared image sequences and high-resolution infrared image sequences to obtain training sample pairs; the training sample pairs include low-resolution infrared images and high-resolution infrared images. Using low- and medium-resolution infrared images as training samples and high-resolution infrared images corresponding to the low-resolution infrared images as sample labels, a convolutional neural network model is trained to obtain an infrared super-resolution model.
[0061] In this embodiment, the low-resolution infrared camera and the high-resolution infrared camera used for acquisition are independent physical devices. Even when acquired synchronously, the two sets of image sequences they acquire are not strictly aligned in time (there may be slight deviations in the acquisition time) and space (content offsets, rotations, or perspective differences caused by different camera positions and lens angles). Through a series of image processing algorithms, these spatiotemporal deviations can be eliminated, ensuring that the two images in each training sample pair used for training are strictly corresponding to the same scene content at the same time.
[0062] The training sample pairs are image pairs obtained after registration processing. Each pair contains one image from a low-resolution infrared image sequence and one image from a high-resolution infrared image sequence. Their contents correspond, and only their resolution and some imaging details differ.
[0063] In each model training iteration, a low-resolution infrared image can be taken from the training sample set as the input data of the neural network, and the high-resolution infrared image that strictly corresponds to the input image can be taken as the true value that the model output should be close to in this iteration.
[0064] A convolutional neural network (CNN) model is a specific type of neural network architecture used to perform super-resolution tasks. It includes multiple convolutional layers and activation functions, among other things.
[0065] After receiving a sample (low-resolution image), the convolutional neural network model performs forward propagation to compute a predicted high-resolution image. Then, it calculates the difference (loss value) between this predicted image and the sample label (the true high-resolution image). This loss value is then backpropagated using the backpropagation algorithm, and the thousands of parameters (weights and biases) within the model are adjusted (optimized) accordingly. The model's parameters are continuously adjusted so that its predicted output increasingly approximates the true sample label.
[0066] Once the training process meets the preset termination conditions (such as reaching a certain number of iterations, or the loss value no longer decreasing significantly), the model has stored the optimal parameters learned from massive amounts of real data pairs. This neural network with solidified parameters is the infrared super-resolution model that can be deployed on the vehicle for real-time inference.
[0067] In some embodiments, registering a low-resolution infrared image sequence and a high-resolution infrared image sequence to obtain training sample pairs may include: Time alignment processing is performed on low-resolution infrared image sequences and high-resolution infrared image sequences; Spatial registration is performed on time-aligned low-resolution infrared image sequences and high-resolution infrared image sequences to obtain training sample pairs.
[0068] In this embodiment, the same motion state of the same moving target in two sequences can be found and matched. By comparing the pose or position of a significant moving object (such as a pedestrian or vehicle) in the two video streams, the temporal offset between the two sequences is calculated (for example, the 100th frame of the low-resolution sequence and the 105th frame of the high-resolution sequence depict the same instantaneous state). Then, the frame index of one of the sequences is slid or interpolated using this offset, so that the low-resolution infrared image sequence and the high-resolution infrared image sequence are aligned in the temporal flow of the content.
[0069] The first and second sample cameras cannot be physically installed in exactly the same spatial position. The images of the same scene captured by the two sample cameras will have differences in the content of the images, such as translation, rotation, scaling or perspective distortion. That is, the pixel coordinates of the same object may be different in the two images. A pixel-level or feature-level correspondence can be established between the two images. Based on the correspondence, one of the images can be resampled (such as by perspective transformation) to make its content completely aligned with the other image in space.
[0070] After the temporal alignment and spatial registration processes described above, for each set of temporally aligned frames, a pair of images with pixel-level content alignment is obtained: one is the original (or upsampled) low-resolution infrared image, and the other is a high-resolution infrared image that has undergone spatial transformation correction and whose content strictly corresponds to the former. This pair of images constitutes a training sample pair.
[0071] In some embodiments, time alignment processing of low-resolution infrared image sequences and high-resolution infrared image sequences may include: Acquire the motion state of the same moving target in low-resolution infrared image sequences and high-resolution infrared image sequences; Based on the temporal consistency of motion states, the system time deviation between the first sample camera and the second sample camera is obtained; Time alignment processing is performed on low-resolution infrared image sequences and high-resolution infrared image sequences based on system time deviation.
[0072] In this embodiment, the moving target is an object that undergoes changes in position or shape in the scene being captured, such as a walking pedestrian, a moving vehicle, or a swaying tree branch.
[0073] The motion state of a moving target includes information such as instantaneous posture, spatial position, direction of motion, and velocity.
[0074] A physical event (a specific state of a moving object) occurs only once in the real world. In the case of perfect synchronization, two images (one from a low-resolution camera and one from a high-resolution camera) depicting the same motion state of the event should have the same timestamp.
[0075] Because the clocks of the two cameras are not synchronized, two images describing the same physical instant will have different timestamps in the actual recording (e.g., T1 and T2). The difference between these two timestamps, ΔT = T1 - T2, is the system time deviation between the two cameras.
[0076] After obtaining the system time deviation ΔT, the timestamp of one of the image sequences can be globally compensated. For example, assuming ΔT = -0.1 seconds (that is, the timestamp of high-resolution camera images is on average 0.1 seconds later than that of low-resolution camera images), when searching for the corresponding high-resolution image for the low-resolution image, 0.1 seconds can be added to the timestamp of the low-resolution image, and then the image closest to the corrected time in the high-resolution sequence can be found as its corresponding frame.
[0077] Through the compensation operation described above, two image sequences that were originally misaligned on the timeline are aligned on the content timeline. For any given content moment, a pair of images depicting the same real instantaneous scene can be extracted from the two sequences.
[0078] In some embodiments, spatial registration processing is performed on time-aligned low-resolution infrared image sequences and high-resolution infrared image sequences to obtain training sample pairs, which may include: Distortion correction is performed on time-aligned low-resolution infrared image sequences and high-resolution infrared image sequences. Feature points are obtained from distortion-corrected low-resolution infrared image sequences and high-resolution infrared image sequences, and multiple feature points are matched. Based on the matched feature points, the transformation relationship between the high-resolution infrared image sequence and the low-resolution infrared image sequence in the image space is determined. Based on the transformation relationship, a spatial transformation is performed on the high-resolution infrared image sequence to obtain a high-resolution infrared image sequence that is spatially aligned with the low-resolution infrared image sequence, so as to obtain training sample pairs.
[0079] In this embodiment, distortion correction processing is performed on the time-aligned low-resolution infrared image sequence and the high-resolution infrared image sequence to reduce the impact of lens distortion on image geometry and restore the image to a state without geometric distortion that conforms to the ideal pinhole camera model. For example, the two cameras can be calibrated in advance to obtain their respective intrinsic parameter matrices and distortion coefficients, and then these parameters can be applied to perform inverse distortion transformation on the image.
[0080] Feature points are locations in an image that possess significant local characteristics, such as corners, edge intersections, or patch centers. For infrared images, algorithms sensitive to brightness variations and edges (such as SIFT (Scale-invariant feature transform) or SURF (Speeded-Up Robust Features)) can be used to extract stable and repeatable feature points. A set of feature points and vectors describing the image information surrounding the points can be extracted from both the distortion-corrected low-resolution and high-resolution images.
[0081] By calculating the distance between feature point vectors (such as Euclidean distance), for each feature point in a low-resolution image, the feature point with the most similar vector can be found in the high-resolution image as its matching point.
[0082] After obtaining a reliable set of matching feature point pairs (i.e., pixel coordinate pairs corresponding to the same physical location in two images), these coordinate pairs can be used to solve a spatial transformation model. For example, the homography matrix, a 3x3 matrix, can describe the perspective transformation relationship from one plane (high-resolution image plane) to another plane (low-resolution image plane), and can cover comprehensive deformations such as translation, rotation, scaling, and shearing.
[0083] Based on the coordinates of the matching point pairs, the optimal homography matrix, i.e. the transformation relation, is solved by algorithms such as direct linear transformation or RANSAC (Random Sample Consensus). The transformation relation defines how to map each point on the high-resolution image to the coordinate space of the low-resolution image.
[0084] Transformation relationships (such as homography matrices) can be applied to the entire high-resolution image. Geometric transformations of high-resolution images can be performed using image resampling techniques (such as forward or backward mapping combined with bilinear / bicubic interpolation).
[0085] A training sample pair can be formed by combining a raw (or upsampled) low-resolution infrared image with a high-resolution infrared image that has undergone distortion correction and spatial transformation and is time-aligned with the low-resolution image. The two images in the training sample pair depict the same viewpoint content of the same scene at the same time.
[0086] In some embodiments, distortion correction processing of time-aligned low-resolution infrared image sequences and high-resolution infrared image sequences may include: Obtain the camera intrinsic parameter matrix and distortion coefficients corresponding to the first and second sample cameras, respectively; Based on the camera intrinsic parameter matrix and distortion coefficients corresponding to the first sample camera, distortion correction is performed on the time-aligned low-resolution infrared image sequence; based on the camera intrinsic parameter matrix and distortion coefficients corresponding to the second sample camera, distortion correction is performed on the time-aligned high-resolution infrared image sequence.
[0087] In this embodiment, the camera intrinsic parameter matrix characterizes the internal geometric and optical properties of the camera, including focal length and principal point coordinates.
[0088] Distortion coefficients can be vectors used to characterize the degree and type of lens distortion. They can include radial distortion coefficients and tangential distortion coefficients, etc. Radial distortion coefficients are used to characterize the bending of light due to the shape of the lens, causing straight lines at the edge of the image to bend inward or outward. Tangential distortion coefficients are used to characterize the offset of light caused by the non-parallelism between the lens and the imaging sensor.
[0089] A calibration board with a known precise geometric pattern (such as a checkerboard or circular array) can be used. Multiple images of the calibration board are taken from the camera to be calibrated at different distances and angles. These images are then processed based on a calibration algorithm to obtain the camera's intrinsic parameter matrix and distortion coefficients.
[0090] By utilizing the camera's calibration parameters, a pixel coordinate mapping relationship can be established between the original distorted image and the ideal distortion-free image. For example, based on the intrinsic parameter matrix and distortion coefficients, the sampling position corresponding to each pixel in the ideal distortion-free image on the original distorted image can be calculated. Then, the pixel value at that position can be obtained through an interpolation algorithm, thereby generating a corrected image.
[0091] In some embodiments, after spatial registration of the time-aligned low-resolution infrared image sequence and the high-resolution infrared image sequence, the method may include: Image enhancement processing is performed on the spatially registered high-resolution infrared image sequence to obtain training sample pairs; or, Training sample pairs are obtained by performing at least one of the following processing steps on the spatially registered high-resolution infrared image sequence: histogram equalization, contrast stretching, noise suppression, and image sharpening.
[0092] In this embodiment, high-resolution infrared image sequences can be enhanced using digital image processing techniques to improve image quality. For example, histogram equalization, adaptive histogram equalization, gamma correction, and contrast stretching can be used to enhance image contrast, making the differences between bright and dark areas more obvious; alternatively, non-local mean filtering and wavelet denoising can be used to reduce random noise in the image; or, Laplacian sharpening, unsharpening masks, and high-lift filtering can be used to sharpen and enhance the image. The choice can be made based on user needs, and this application does not impose any limitations.
[0093] Image enhancement processing may include at least one of the following: histogram equalization, contrast stretching, noise suppression, and image sharpening.
[0094] Histogram equalization is used to adaptively adjust the grayscale distribution in different regions of an image, maximizing local contrast and solving the problems of noise amplification and local overexposure caused by global equalization.
[0095] Contrast stretching is used to further extend the grayscale range of an image linearly or non-linearly, making dark areas darker and bright areas brighter, thereby enhancing the overall contrast.
[0096] Noise suppression is used to effectively reduce image noise and improve the signal-to-noise ratio by utilizing denoising algorithms to preserve edges.
[0097] Image sharpening is used to enhance the clarity of high-resolution infrared image sequences by making the outlines and texture details of objects clearer by enhancing the high-frequency components of the image (such as edges).
[0098] The training sample pairs include raw or upsampled low-resolution infrared images, as well as high-resolution infrared images that have undergone time alignment, spatial registration, and contrast enhancement.
[0099] During the research and development process, the inventors discovered that some related technologies rely on fuzzy kernel degradation to synthesize training datasets, which leads to poor model generalization ability.
[0100] In this application, images are acquired simultaneously using a high-resolution infrared camera and a low-resolution infrared camera. By registering the acquired low-resolution infrared image sequence and the high-resolution image sequence, paired data of low-resolution infrared images and high-resolution infrared images of real scenes are obtained, resulting in training sample pairs. Infrared super-resolution models are then trained based on these training sample pairs, solving the problem that the fuzzy kernel degradation method cannot simulate real low-resolution infrared images, leading to poor model generalization ability.
[0101] In some embodiments, the infrared super-resolution model is a lightweight model, comprising a shallow feature extraction module, a deep feature extraction module, and a feature recombination upsampling module connected in sequence.
[0102] In this embodiment, the lightweight model has fewer parameters, lower computational complexity, faster inference speed, and smaller memory footprint, enabling the infrared super-resolution model to meet the constraints of the vehicle-mounted embedded platform, including limited chip computing power, power consumption budget, and real-time requirements.
[0103] The shallow feature extraction module is the first layer or first stage of the model processing, and can directly receive the original low-resolution infrared image as input.
[0104] The shallow feature extraction module can perform preliminary feature representation transformation. It can include one or a few convolutional layers to map the input single-channel or three-channel infrared image to a higher-dimensional feature space.
[0105] The deep feature extraction module can include multiple convolutional layers (which can be combined with activation functions, etc.) to perform deep nonlinear feature abstraction and fusion.
[0106] The feature recombination upsampling module can be used to upscale low-resolution high-level feature maps to the high-resolution size of the target using a specific algorithm. After upsampling, the high-level semantic features can be translated back to specific pixel values to generate the final high-resolution infrared image.
[0107] In this application, the constructed infrared super-resolution model includes 6 convolutional and activation layers, as well as 1 upsampling layer, which can achieve super-resolution from 640×512 resolution to 1280×1024 resolution and 30 frames of real-time inference with a board-side computing power consumption of 2.7 TOPS. This infrared super-resolution algorithm can be deployed on the vehicle-mounted board and has a wide range of applicable scenarios.
[0108] In actual implementation, such as Figure 2 As shown, during the training of the infrared super-resolution model, S10 can be executed first: high-resolution infrared data and low-resolution infrared data acquisition.
[0109] like Figure 3 The data acquisition process is illustrated in S101: The first sample camera (low-resolution infrared camera) and the second sample camera (high-resolution infrared camera) are combined and bound together so that the fields of view of the two infrared cameras are basically the same.
[0110] S102: Acquire intrinsic parameter calibration data from the first sample camera (low-resolution infrared camera) and the second sample camera (high-resolution infrared camera). For example, an open scene can be selected. Power on the thermal imaging calibration board and preheat it for two minutes to ensure that the thermal imaging calibration board's features are clearly imaged in the infrared camera. Then, acquire infrared camera data. By moving the camera, ensure that the thermal imaging calibration board is fully imaged in different areas of the infrared camera's view, and ensure that the thermal imaging calibration board is tilted to different degrees vertically, horizontally, and vertically at the same position, covering the entire camera's view. Figure 4 An example of infrared camera intrinsic parameter calibration data acquisition is provided.
[0111] S103: Install the low-resolution infrared camera and high-resolution infrared camera combination module in a suitable position on the vehicle. After fixing the two camera modules in S101, the two cameras can be fixed above the windshield of the car, with the viewing angle facing forward of the vehicle. Adjust the viewing angle so that the road surface occupies about a quarter of the area in the camera's image, without excessive rotation relative to the vehicle.
[0112] S104: Multi-Scene Data Acquisition. Diverse training scenario data can effectively improve the generalization performance of infrared super-resolution models, making tasks such as nighttime perception and severe weather recognition more reliable and accurate. Table 1 illustrates typical automotive perception scenarios suitable for infrared super-resolution, categorized by environment, traffic complexity, and task requirements. Vehicles can drive freely in different scenarios, and data acquisition personnel can use the data acquisition program to simultaneously collect data from both low-resolution and high-resolution infrared cameras.
[0113]
[0114] S20: Infrared image distortion correction to eliminate the impact of image distortion on image pairing. The specific process is as follows: Figure 5 As shown.
[0115] S201: Binarization of Infrared Calibration Data. The feature points of the thermal imaging calibration board in the acquired infrared camera calibration data are white, which may cause the circular target detection API in OpenCV to fail to accurately identify the infrared imaging target. Special processing can be performed. For example, the data can first be imported into ImageJ software, then Image---Type---8-bit can be selected to convert the image to a single-channel image. Then, Image---Adjust---Threshold can be selected for binarization. A pop-up window will appear... Figure 6 The window shown.
[0116] You can drag the maximum and minimum thresholds with the mouse to flip the image colors until the feature targets turn black and the background turns white, such as... Figure 7 As shown, all infrared camera data collected can be processed in the same way.
[0117] S202: Target Detection. The `cv2.findCirclesGrid()` function in OpenCV can be used to detect circular target points in an image. The `cv2.find4QuadCornerSubpix()` function can be used to refine the inner corner points. The accuracy of circular target detection is as follows: Figure 8 An example of the detection results for circular target points is provided.
[0118] S203: Intrinsic parameter calibration. Use the cv2.calibrateCamera() function to calibrate the intrinsic parameters of the infrared camera and output the calibrated intrinsic parameter matrix, distortion coefficients, and calibration error.
[0119] S204: Distortion Correction. After calibrating the intrinsic parameters of the infrared camera, you can use OpenCV's cv2.getOptimalNewCameraMatrix() function to input the calibrated camera intrinsic parameter matrix K and distortion coefficients D to obtain the distortion-corrected camera intrinsic parameter K_new. Then, use the cv2.undistort() function to input K, D, K_new, and the acquired original infrared image data to perform distortion correction on the infrared image. Figure 9 The distortion correction effect is illustrated in the example. Figure 9 The first row is a high-resolution infrared image, the second row is a low-resolution infrared image, the left column is the original image, and the right column is the image after distortion correction.
[0120] S30: Register low-resolution infrared images with high-resolution infrared images. For example... Figure 10 The registration process is illustrated.
[0121] S301: Time alignment of the low-resolution infrared camera (first sample camera) and the high-resolution infrared camera (second sample camera). The low-resolution and high-resolution infrared cameras use different platforms for data acquisition, and the infrared cameras do not support external triggering. Therefore, time alignment is required. Data from the same scene can be used to align the cameras based on identical moving objects such as people or vehicles captured by both cameras, ensuring that the movements and positions of moving objects captured by the low-resolution and high-resolution infrared cameras are consistent.
[0122] S302: Distortion Correction. Distortion correction can be performed on images acquired by low-resolution infrared cameras and high-resolution infrared cameras based on the camera intrinsic parameter matrix K and distortion coefficients D, respectively, to reduce the impact of image distortion on registration.
[0123] S303: Low-resolution image upsampling. This can be achieved using the bicubic interpolation function in OpenCV. The low-resolution image after distortion removal is upsampled so that the low-resolution infrared image and the high-resolution infrared image have the same size.
[0124] S304: Keypoint Detection in Infrared Images. Keypoints can be extracted from time-aligned low-resolution and high-resolution infrared images using OpenCV's SIFT feature extractor.
[0125] S305: Feature Matching. Based on the extracted key points, the features of the two infrared images are matched. The Euclidean distance between the key points in the two images is calculated, and the images are sorted according to their Euclidean distance. The top k images with the smallest distance (e.g., k=2) are selected as the matching results. Figure 11 As shown. Based on the matching results, the homography matrix between two planar point sets can be calculated using the RANSAC algorithm, i.e., the transformation relationship (perspective transformation matrix, M).
[0126] S306: Perspective Transformation. Based on the perspective transformation matrix M, the cv2.warpPerspective() function in OpenCV can be used to perform perspective transformation on a high-resolution infrared image, and then superimpose and fuse it with the corresponding low-resolution image to obtain an aligned fused image.
[0127] S307: Data Filtering. The fused images are filtered, removing those with misregistration. Based on the filtered fused data, a high-resolution infrared image and its corresponding low-resolution infrared image after perspective transformation are obtained.
[0128] S40: High-resolution infrared image enhancement, such as Figure 12 As shown, the acquired high-resolution infrared images can be processed with histogram equalization, contrast stretching, noise suppression, and image sharpening to enhance the clarity and contrast of the infrared super-resolution model's supervisory signal, creating image effects that better match human visual perception.
[0129] S401: Local histogram equalization enhances contrast by redistributing image pixel intensity. For example, local histogram equalization can be used to adaptively address the failure of global equalization in areas of large brightness difference. The image is divided into specified small blocks, and histogram equalization is performed independently on each block to adapt to the contrast characteristics of the local area. If the number of pixels at a certain gray level exceeds 2, the excess pixels are cropped, and the cropped portion is evenly distributed across all gray levels to avoid noise amplification and local over-enhancement. Finally, bilinear interpolation is performed on the transformation function of adjacent blocks to smooth the pixel value transition and eliminate block artifacts.
[0130] S402: Contrast Stretching. After local histogram equalization, infrared images may exhibit localized overexposure, excessive noise enhancement, and unnatural overall contrast. To further optimize image quality, a linear transformation can be used to adjust the pixel range, thereby enhancing the image's visual appeal. The main processing steps are: first, perform a linear transformation on the entire image; then, shift the overall pixel values to adjust the overall contrast, gently reducing it to prevent overexposure.
[0131] S403: Noise Suppression. After local histogram equalization and contrast stretching, the contrast and visual effect of high-resolution infrared images are greatly improved, but noise information may be amplified, affecting the visualization effect. To improve the clarity of infrared super-resolution model inference, a noise suppression module can be added to the high-resolution infrared image enhancement part. For example, BM3D can be used to suppress noise in high-resolution infrared images, and edge-preserving denoising can be achieved by utilizing non-local similarity and frequency domain sparsity.
[0132] S404: Image Sharpening. While BM3D processing can significantly suppress noise in high-resolution infrared images, object contours and textures may not be prominent, resulting in an overly smooth visual appearance. An image sharpening module can be added to enhance the contour information of the infrared image. Firstly, the infrared image... Gaussian filtering is performed to obtain the filtered infrared image. Then use minus Obtain infrared images Contour feature image Finally, the contour feature image and Adding them together yields an infrared image with sharpened edges. The processing procedure is shown in the following formula:
[0133]
[0134] Image sharpening enhances the high-frequency information of edges and details in an image, compensating for the contours and textures of objects, thereby improving the image's clarity and recognizability.
[0135] S50: Model structure design. The infrared super-resolution module can include a shallow feature extraction module, a deep feature extraction module, and a feature reconstruction module, such as... Figure 13 As shown.
[0136] The shallow feature extraction module processes the first infrared image to obtain shallow features, the deep feature extraction module processes the shallow features to obtain deep features, and the feature recombination and upsampling module performs feature recombination and fusion on the deep features to obtain the second infrared image. The deep feature extraction module also includes a reparameterized residual module.
[0137] Specifically, the shallow feature extraction module can use a simple layer. Convolution is used to upscale infrared data from a single channel to four channels, enabling simple shallow feature extraction.
[0138] The deep feature extraction module contains N layers The process uses convolutional layers and a HardTan activation function. The first convolutional module enhances the 4-channel shallow features into M (M>4)-channel deep features, and the remaining N-1 convolutional layers extract these deep features. Finally, a new layer is added. Convolution reduces M-channel deep features to 4-channel features and adds skip connections to shallow features, increasing the detail information of the shallow features. To enhance the feature extraction capability of the deep feature extraction module, related technologies often enhance a large number of residual connections in the model structure. However, applying residual connection structures to the board significantly increases the memory access latency of the model on the board, while also increasing performance overhead. Without residual connections, the model's accuracy and performance will significantly decrease. In this application, the first N layers... Convolutional layers incorporate initial residual connections to achieve implicit residual connections. A typical residual connection can be represented as... It can be further reparameterized into The mapping path of the input features is embedded as a convolution with a weight of 1 into a 3x3 convolution. Specifically, this can be achieved by adding a weight kernel with a center value of 1 and other values of 0 to the randomly initialized 3x3 convolution kernel. This application implicitly embeds the residual connection structure into the initialization method during training, directly eliminating the use of residual connections in the network design. This increases the feature extraction capability of the deep feature extraction module and avoids unnecessary performance overhead and latency.
[0139] The feature reorganization module first performs nearest-neighbor upsampling on the features output by the deep feature module to increase the feature size, and finally passes it through one layer. Convolutional processing combined with the HardTan activation function is used for feature recombination and fusion, while simultaneously reducing the 4-channel feature to a single channel to output the infrared super-resolution result. In related technologies, vehicle-mounted infrared super-resolution schemes are mostly limited to a fixed 2x super-resolution, which cannot meet users' zoom visualization needs. This application, by modifying the nearest neighbor upsampling parameters, can achieve 1 to 8x stepless infrared image super-resolution, breaking through the limitations of traditional fixed-magnification models and supporting clear imaging in both close-range and long-range scenes.
[0140] This application aims to implement a real-time infrared super-resolution algorithm on the vehicle-mounted board. The deep feature extraction module N is set to 3, and the feature channels M is set to 16. The inventors found through testing that the model can meet the real-time requirements under the low computing power platform TDA4VME board with a computing power consumption of less than 3 TOPS, and the super-resolution effect can meet the project requirements.
[0141] S60: Constructing the loss function. This application can use a fully supervised approach for model training, employing multiple loss combinations for model supervision, as shown in the following formula, including L1 loss, perceptual loss, frequency domain loss, and histogram loss to improve the model's contrast enhancement capability and super-resolution performance, thereby enhancing the dynamic range and feature information of infrared images.
[0142]
[0143] In image super-resolution, L1 loss measures the pixel-level absolute difference between the generated image and the high-resolution ground truth, as shown in the formula below. Its principle is to guide the model to reconstruct sharper edges and textures by minimizing the absolute error. Compared to L2 loss, it reduces blur and generates a sharper visual result.
[0144]
[0145] Perceptual loss utilizes pre-trained networks (such as VGG) to extract high-level features of images. The model can be optimized by calculating the difference between the generated image and the real image in the feature space (such as L1 distance). This application uses five layers of features of different sizes from VGG19 for simultaneous supervision, as shown in the formula below. A larger j indicates a deeper feature from VGG. Shallow features represent fine pixel-level details (such as texture and contour), while deep features ensure consistency between high-level semantic content and the overall structure, thus achieving a detailed and visually realistic reconstruction effect. The principle is to force the network to learn semantic information and texture details that are sensitive to the human eye, rather than simply pixel precision, thereby generating more natural and detailed results.
[0146]
[0147]
[0148] Frequency domain loss. While combining L1 loss and perceptual loss can achieve good super-resolution results, there is still a gap between the super-resolution image and the real image, especially in the frequency domain. Reducing the error in the frequency domain can further improve the quality of image super-resolution. Therefore, this application incorporates a frequency domain loss. First, a two-dimensional discrete Fourier transform is performed on the infrared image, and the image frequency information is decomposed into real and imaginary parts using Euler's formula. The real and imaginary parts correspond to the x-axis and y-axis in two-dimensional space, respectively. The Euclidean vectors of the super-resolution image and the real image are calculated using the real and imaginary parts. Finally, different frequencies can be weighted according to their difficulty, and the weights are dynamically determined based on the distribution of the image and the ground truth (GT).
[0149]
[0150] Histogram Loss. To further improve the contrast enhancement capability of the super-resolution model, this application incorporates histogram loss. First, the histogram distribution of the super-resolution image and the ground truth image is statistically analyzed. The number of pixel values from 0 to 255 is counted, and the absolute value of the number of pixel values in both the super-resolution and ground truth images is calculated as the histogram loss, as shown in the following formula. This represents the number of pixels with a value of i in the statistical image x. Histogram loss further constrains the histogram distribution of the inference results of the infrared super-resolution model. Combined with the contrast adjustment of the GT data in step S40, it further enhances the contrast enhancement capability of the super-resolution model.
[0151]
[0152] S70: Data Augmentation and Training. During the model training phase, this application can augment the data using left-right flipping, up-down flipping, random rotation, and random brightness transformation to increase the diversity of the training data. The left-right flipping, up-down flipping, and random rotation operations each have a 50% probability of being triggered. The random rotation angle is sampled from -180 degrees to 180 degrees, with the sampling probability following a uniform distribution. To expand the brightness distribution of the training data, this application incorporates random brightness transformation, where the brightness transformation factor can be randomly sampled from [1.0, 0.8, 0.7]. The data augmentation process is as follows: Figure 14 As shown.
[0153] This application can use the standard Adam optimizer for model optimization, and the initial learning rate can be set to 1e-4. =0.9, =0.999, trained for 600 epochs with a batch size of 32. The core scheduling strategy is a step-by-step decay of the learning rate, which balances convergence speed and final performance by periodically reducing the learning rate: every 70 epochs, the learning rate is decayed by a factor of 0.5. This configuration means that the learning rate will undergo 8 step-by-step decreases during training (70*8=560), eventually dropping to about 1 / 256 of the initial value, which helps to achieve fine convergence of the model in the later stages of training.
[0154] To achieve 1 to 8 times super-resolution of infinite infrared images, this application can randomly change the hyperparameters of the nearest neighbor upsampling of the model feature reorganization module during model training by using random hyperparameters. This involves training the infinite infrared image super-resolution model by upsampling, while also adjusting the resolution of the corresponding supervision image to match the model output for loss calculation. Hyperparameters The value range can be [1, 8], for example, randomly selecting hyperparameters. If the value is 4, the feature remodeling module will upsample the features output by the deep feature module to four times their original value, and the resolution of the supervision image will also be adjusted to the corresponding size, and so on. The training process is as follows: Figure 15 As shown.
[0155] During the research and development process, the inventors discovered that infrared images have low resolution, low contrast, and low signal-to-noise ratio. In related technologies, the above problems are solved by enhancing the brightness contrast in the middle of the image through traditional methods, but this leads to the loss of highlight and shadow details.
[0156] In this application, by performing contrast enhancement and denoising processing on high-resolution infrared image data, the contrast, edge sharpening, and visualization effect of supervised data (i.e., training sample pairs) can be improved. Furthermore, by using L1 loss, perceptual loss, frequency domain loss, and histogram loss for model training, the global and local contrast enhancement mapping capabilities of the model are enhanced, as well as the edge sharpening processing capabilities are improved. This results in improved dynamic range and image detail, and enhanced sensory effects of infrared data on the human eye.
[0157] In some embodiments, acquiring the second infrared image output by the infrared super-resolution model may include: Receive the user's first input; In response to the first input, the upsampling magnification parameter of the infrared super-resolution model is adjusted to obtain the second infrared image output by the infrared super-resolution model corresponding to the magnification adjustment command.
[0158] In this embodiment, the generation of the final high-resolution image (i.e., the second infrared image) can be flexibly and dynamically adjusted according to the user's real-time instructions.
[0159] The first input is used to determine the magnification adjustment command, which includes a specified target super-resolution factor, i.e., the magnification ratio of the second infrared image relative to the first infrared image in terms of size, such as 2x, 3.5x, or 8x.
[0160] The first input can be in at least one of the following ways: Firstly, the first input can be a touch operation, including but not limited to click, swipe, and press operations.
[0161] In this embodiment, receiving the user's first input can be receiving the user's touch operation on the display area of the vehicle's central control display screen.
[0162] To reduce user error rates, the effective area of the first input can be limited to a specific area, such as the edge adjustment control area of the infrared image display pane; or, while displaying a real-time infrared image, a target control such as a magnification adjustment slider or plus / minus buttons can be displayed on the current interface, and touching the target control will enable the first input; or the first input can be set to a series of taps on the display area within a target time interval.
[0163] Secondly, the first input can be a physical button input.
[0164] In this embodiment, the vehicle's steering wheel or central control area is provided with physical buttons corresponding to the magnification adjustment, which can receive the user's first input, such as receiving the user's operation of pressing the corresponding physical button; the first input can also be a combination operation of pressing multiple physical buttons simultaneously.
[0165] Thirdly, the first input can be the steering wheel scroll wheel input.
[0166] In this embodiment, receiving the user's first input can be receiving the user's operation of rotating a multi-function scroll wheel mounted on the steering wheel. The scrolling direction (e.g., forward / backward) or the number of scroll steps of this scroll wheel can be mapped to an instruction to increase or decrease the super-resolution.
[0167] Fourth, the first input can be voice input.
[0168] In this embodiment, the vehicle's voice recognition module can trigger the generation of a corresponding magnification adjustment command when it receives a voice command containing magnification keywords (such as "magnify the infrared image by two times", "reduce it a little", or "maximum magnification").
[0169] Of course, in other embodiments, the first input may also be in other forms, including but not limited to gesture recognition input, gaze point control input, etc., which can be determined according to actual needs, and this application embodiment does not limit it.
[0170] The infrared super-resolution model is a model with variable upsampling capability. Within the network structure of the infrared super-resolution module, there is one or a set of configurable parameters that control the size of the final output image, namely the upsampling factor parameter. This parameter determines whether the model enlarges the input image to 2x, 4x, or any other preset factor.
[0171] The system can dynamically modify the values of parameters in the model according to the specific values indicated by the received magnification adjustment command, so as to obtain output images at different magnifications.
[0172] In actual implementation, continue to refer to Figure 2 S80: Vehicle-mounted board deployment. A trained infrared super-resolution model can be deployed on the vehicle-mounted board to build a 1 to 8x infinite infrared image super-resolution system. This system and equipment can run infrared super-resolution in real-time on a low-computing-power vehicle platform, such as... Figure 16 An example of an infrared super-resolution system is given, which includes: a low-resolution infrared sensor (first infrared camera), an onboard low-computing-power platform, a central control screen (display device), a voice recognition module, and a steering wheel scroll wheel.
[0173] The automotive low-computing platform is a high-performance system-on-a-chip (SoC) specifically designed for advanced driver assistance systems (ADAS), autonomous driving (AD), and smart cockpit applications. It integrates various processing units and interfaces, aiming to provide powerful computing capabilities, low power consumption, and high security.
[0174] The voice recognition module can recognize the user's voice requests, such as enabling 1.5x zoom, 3x magnification, and 2x reduction, and can extract hyperparameters for stepless infrared image super-resolution from the user's voice. The data is then fed into an in-vehicle low-computing-power platform to achieve infrared super-resolution at the corresponding magnification. Users can also achieve stepless super-resolution from 1x to 8x by adjusting the steering wheel scroll wheel.
[0175] After the infrared super-resolution model is trained and converged, a model parameter file is obtained. After setting the required parameter configuration file for model conversion (including basic parameters, preprocessing parameters, quantization parameters, and inference parameters), the TIDL-RT model conversion tool can be used to perform model conversion, exporting a TIDL model file that can run on the TDA4WME. The quantization stage uses a post-training quantization strategy to quantize to 8 bits. After model conversion, the TIDL model file can be applied for model inference on the TDA4VME board. Testing showed that super-resolution of infrared images from 640×512 resolution to 1280×1024 resolution on the board achieved 30 frames of real-time inference with a computing power consumption of 2.7 TOPS, providing a possibility for the application of infrared super-resolution algorithms on vehicle-mounted systems. Figure 17 and Figure 18The comparison between board-side infrared super-resolution and bicubic interpolation (left column shows bicubic interpolation results, right column shows infrared super-resolution results) shows that the method provided in this application can guarantee image quality while ensuring performance. The infrared image after super-resolution processing has significantly improved in terms of clarity and contrast.
[0176] In this application, an infrared super-resolution model is deployed on the vehicle-mounted board. Real-time infrared images are acquired through a low-resolution infrared camera, sent to a low-computing-power vehicle platform for infrared image super-resolution processing, and then displayed on the central control screen. Stepless super-resolution control from 1 to 8 times can be achieved through the steering wheel scroll wheel and voice recognition module. Only one hyperparameter of the infrared super-resolution model needs to be changed to flexibly achieve super-resolution enhancement at any magnification within the range of 1 to 8 times, breaking through the limitations of traditional fixed magnification models. At the same time, it can achieve clear imaging of both near and far-field scenes, allowing the driver to dynamically adjust the visual focal length according to actual road conditions and observation needs. This provides users with a visual zoom capability that can be freely adjusted according to actual needs, improves the human-computer interaction experience, and enhances the adaptability and practicality of infrared image analysis.
[0177] The image processing method for vehicles provided in this application can be executed by an image processing device for vehicles. This application uses an image processing device for vehicles executing the image processing method for vehicles as an example to illustrate the image processing device for vehicles provided in this application.
[0178] This application also provides an image processing apparatus for vehicles.
[0179] like Figure 19 As shown, the vehicle includes a first infrared camera, and the image processing device for the vehicle includes a first processing module 1910, a second processing module 1920, and a third processing module 1930.
[0180] The first processing module 1910 is used to acquire a first infrared image of the scene where the vehicle is located based on the first infrared camera. The second processing module 1920 is used to input the first infrared image into the infrared super-resolution model and obtain the second infrared image output by the infrared super-resolution model; the image quality of the second infrared image is higher than that of the first infrared image; the infrared super-resolution model is trained based on the low-resolution infrared image sequence corresponding to the sample scene acquired by the first sample camera and the high-resolution infrared image sequence corresponding to the sample scene acquired by the second sample camera, and the first sample camera is of the same type as the first infrared camera. The third processing module 1930 is used to display the second infrared image.
[0181] The image processing device for vehicles provided in the embodiments of this application can acquire environmental images by deploying a single low-resolution infrared camera on the vehicle and use a super-resolution neural network model to process the acquired low-quality infrared images in real time. It can output infrared images with higher image quality, such as higher resolution, higher clarity and higher contrast, without significantly increasing the cost of on-board hardware (only requiring the retention of a single low-cost infrared camera). This significantly improves visual perception and driving safety in low-visibility environments such as night, fog, rain and snow, and achieves the goal of compensating for hardware limitations with algorithm performance. It is applicable to a variety of image types and scenarios, can meet the perception needs of intelligent driving in complex scenarios, and has high application value.
[0182] In some embodiments, the image processing apparatus for the vehicle may further include a fourth processing module for: Image quality includes at least one of resolution, contrast, and sharpness; and / or, The first infrared camera is a low-resolution camera; and / or, The resolution of the first infrared camera is less than or equal to 650x512.
[0183] In some embodiments, the image processing device for the vehicle may further include a fifth processing module for: The infrared super-resolution model is trained based on the low-resolution infrared image sequence corresponding to the sample scene acquired by the first sample camera and the high-resolution infrared image sequence corresponding to the sample scene acquired by the second sample camera. High-resolution infrared image sequences have a higher resolution than low-resolution infrared image sequences; and / or, The first sample camera is a low-resolution camera with a resolution of 650x512 or less. The second sample camera is a high-resolution camera with a resolution greater than 650x512.
[0184] In some embodiments, the image processing apparatus for the vehicle may further include a sixth processing module for training an infrared super-resolution model based on the following steps: Registration processing is performed on low-resolution infrared image sequences and high-resolution infrared image sequences to obtain training sample pairs; the training sample pairs include low-resolution infrared images and high-resolution infrared images. Using low- and medium-resolution infrared images as training samples and high-resolution infrared images corresponding to the low-resolution infrared images as sample labels, a convolutional neural network model is trained to obtain an infrared super-resolution model.
[0185] In some embodiments, the sixth processing module can also be used for: Time alignment processing is performed on low-resolution infrared image sequences and high-resolution infrared image sequences; Spatial registration is performed on time-aligned low-resolution infrared image sequences and high-resolution infrared image sequences to obtain training sample pairs.
[0186] In some embodiments, the sixth processing module can also be used for: Acquire the motion state of the same moving target in low-resolution infrared image sequences and high-resolution infrared image sequences; Based on the temporal consistency of motion states, the system time deviation between the first sample camera and the second sample camera is obtained; Time alignment processing is performed on low-resolution infrared image sequences and high-resolution infrared image sequences based on system time deviation.
[0187] In some embodiments, the sixth processing module can also be used for: Distortion correction is performed on time-aligned low-resolution infrared image sequences and high-resolution infrared image sequences. Feature points are obtained from distortion-corrected low-resolution infrared image sequences and high-resolution infrared image sequences, and multiple feature points are matched. Based on the matched feature points, the transformation relationship between the high-resolution infrared image sequence and the low-resolution infrared image sequence in the image space is determined. Based on the transformation relationship, a spatial transformation is performed on the high-resolution infrared image sequence to obtain a high-resolution infrared image sequence that is spatially aligned with the low-resolution infrared image sequence, so as to obtain training sample pairs.
[0188] In some embodiments, the sixth processing module can also be used for: Obtain the camera intrinsic parameter matrix and distortion coefficients corresponding to the first and second sample cameras, respectively; Based on the camera intrinsic parameter matrix and distortion coefficients corresponding to the first sample camera, distortion correction is performed on the time-aligned low-resolution infrared image sequence; based on the camera intrinsic parameter matrix and distortion coefficients corresponding to the second sample camera, distortion correction is performed on the time-aligned high-resolution infrared image sequence.
[0189] In some embodiments, the image processing apparatus for the vehicle may further include a seventh processing module, configured to perform image enhancement processing on the spatially registered high-resolution infrared image sequence after spatial registration processing of the time-aligned low-resolution infrared image sequence and the high-resolution infrared image sequence, to obtain training sample pairs; or... Training sample pairs are obtained by performing at least one of the following processes on the spatially registered high-resolution infrared image sequence: histogram equalization, contrast stretching, noise suppression, and image sharpening.
[0190] In some embodiments, the second processing module 1920 may also be used for: Receive the user's first input, which is used to determine the rate adjustment command; In response to the first input, the upsampling magnification parameter of the infrared super-resolution model is adjusted to obtain the second infrared image output by the infrared super-resolution model corresponding to the magnification adjustment command.
[0191] In some embodiments, the image processing apparatus for the vehicle may further include a sixth processing module for making the infrared super-resolution model a lightweight model, comprising a shallow feature extraction module, a deep feature extraction module, and a feature reconstruction upsampling module connected in sequence; or... The infrared super-resolution model includes a shallow feature extraction module, a deep feature extraction module, and a feature recombination and upsampling module connected in sequence. The shallow feature extraction module processes the first infrared image to obtain shallow features, the deep feature extraction module processes the shallow features to obtain deep features, and the feature recombination and upsampling module performs feature recombination and fusion on the deep features to obtain the second infrared image. The deep feature extraction module includes a reparameterized residual module.
[0192] The image processing device for vehicles in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the specific device.
[0193] The image processing device for vehicles in this application embodiment can be a device with an operating system. This operating system can be a Microsoft (Windows) operating system, an Android operating system, an iOS operating system, or other possible operating systems; this application embodiment does not specifically limit the specific operating system.
[0194] The image processing device for vehicles provided in this application embodiment can achieve... Figures 1 to 18 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0195] In some embodiments, this application also provides a vehicle, including: a first infrared camera, a display device, and a controller.
[0196] In this embodiment, the first infrared camera is a device installed on the vehicle to sense and image the infrared radiation of the external environment. For example, it can be an uncooled microbolometer infrared camera, which is relatively low in cost and has a lower resolution (e.g., an imaging resolution of 640×512 pixels). The first infrared camera can capture the heat emitted by the object itself and can still generate an image reflecting the temperature distribution of the object even under conditions of poor visibility such as nighttime darkness, fog, rain, and snow.
[0197] The first infrared camera can be a low-resolution infrared camera.
[0198] Display devices may include central control screens, etc.
[0199] The controller is connected to the first infrared camera and the display device respectively, and is used to process the first infrared image acquired by the first infrared camera based on the image processing method for vehicles described in any of the above embodiments to obtain a second infrared image, and display the second infrared image on the display device.
[0200] In some embodiments, the vehicle may include an image processing apparatus for the vehicle as described in any of the above embodiments.
[0201] According to the vehicle provided in the embodiments of this application, an environmental image can be acquired by a single low-resolution infrared camera deployed on the vehicle, and a super-resolution neural network model can be used to process the acquired low-quality infrared image in real time. This can output infrared images with higher image quality, such as higher resolution, higher clarity and higher contrast, without significantly increasing the cost of onboard hardware (only a single low-cost infrared camera needs to be retained). This significantly improves the visual perception capability and driving safety in low visibility environments such as night, fog, rain and snow, and achieves the goal of making up for hardware limitations with algorithm performance. It is applicable to a variety of image types and scenarios, can meet the perception needs of intelligent driving in complex scenarios, and has high application value.
[0202] In some embodiments, the vehicle may further include: The input module, connected to the controller, receives a first input from the user. This first input determines a magnification adjustment command, which is then used by the controller to acquire a second infrared image corresponding to the magnification adjustment command; and / or, A voice recognition module, connected to the controller, receives user voice commands and processes them into magnification adjustment commands. These magnification adjustment commands are used by the controller to acquire a second infrared image corresponding to the magnification adjustment command; and / or, The physical interaction module, connected to the controller, is used to generate magnification adjustment commands based on user physical interaction operations. These magnification adjustment commands are used by the controller to acquire a second infrared image corresponding to the command; and / or, A scroll wheel installed on the vehicle's steering wheel is used to generate a magnification adjustment command based on the user's scrolling operation. The magnification adjustment command is used by the controller to obtain a second infrared image corresponding to the magnification adjustment command.
[0203] In this embodiment, the voice recognition module can recognize the user's voice requests, such as achieving 1.5x zoom, 3x magnification, and 2x reduction, and can extract hyperparameters for stepless infrared image super-resolution from the user's voice. The data is then fed into an in-vehicle low-computing-power platform to achieve infrared super-resolution at the corresponding magnification. Users can also achieve stepless super-resolution from 1x to 8x by adjusting the steering wheel scroll wheel.
[0204] like Figure 16 As shown, the vehicle may include a low-resolution infrared sensor (first infrared camera), an onboard low-computing-power platform (controller), a central control screen (display device), a voice recognition module, and a steering wheel scroll wheel.
[0205] In some embodiments, such as Figure 20 As shown, this application embodiment also provides an electronic device 2000, including a processor 2001, a memory 2002, and a computer program stored in the memory 2002 and executable on the processor 2001. When the program is executed by the processor 2001, it implements the various processes of the above-described embodiment of the image processing method for vehicles and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0206] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0207] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described image processing method for vehicles and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0208] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0209] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described image processing method for vehicles.
[0210] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0211] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described image processing method embodiment for vehicles, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0212] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0213] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0214] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0215] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
[0216] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0217] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.
Claims
1. An image processing method for vehicles, characterized in that, Including a first infrared camera, the method includes: Based on the first infrared camera, a first infrared image of the scene in which the vehicle is located is acquired; The first infrared image is input into the infrared super-resolution model to obtain the second infrared image output by the infrared super-resolution model. The image quality of the second infrared image is higher than that of the first infrared image. The second infrared image is displayed.
2. The image processing method for vehicles according to claim 1, characterized in that, The image quality includes at least one of resolution, contrast, and sharpness; and / or, The first infrared camera is a low-resolution camera; and / or, The resolution of the first infrared camera is less than or equal to 650x512.
3. The image processing method for vehicles according to claim 1 or 2, characterized in that, The infrared super-resolution model is trained based on the low-resolution infrared image sequence corresponding to the sample scene acquired by the first sample camera and the high-resolution infrared image sequence corresponding to the sample scene acquired by the second sample camera. The resolution of the high-resolution infrared image sequence is higher than that of the low-resolution infrared image sequence; And / or, The first sample camera is a low-resolution camera with a resolution of 650x512 or less, and the second sample camera is a high-resolution camera with a resolution greater than 650x512.
4. The image processing method for vehicles according to claim 3, characterized in that, The infrared super-resolution model is trained based on the following steps: The low-resolution infrared image sequence and the high-resolution infrared image sequence are registered to obtain training sample pairs; the training sample pairs include low-resolution infrared images and high-resolution infrared images. Using the low-resolution infrared images in the training sample pair as samples and the high-resolution infrared images corresponding to the low-resolution infrared images as sample labels, the convolutional neural network model is trained to obtain the infrared super-resolution model.
5. The image processing method for vehicles according to claim 4, characterized in that, The registration process of the low-resolution infrared image sequence and the high-resolution infrared image sequence to obtain training sample pairs includes: The low-resolution infrared image sequence and the high-resolution infrared image sequence are time-aligned. Spatial registration is performed on the time-aligned low-resolution infrared image sequence and the high-resolution infrared image sequence to obtain the training sample pair.
6. The image processing method for vehicles according to claim 5, characterized in that: The step of performing time alignment processing on the low-resolution infrared image sequence and the high-resolution infrared image sequence includes: acquiring the motion state of the same moving target in the low-resolution infrared image sequence and the high-resolution infrared image sequence; obtaining the system time deviation between the first sample camera and the second sample camera based on the temporal consistency of the motion state; performing time alignment processing on the low-resolution infrared image sequence and the high-resolution infrared image sequence based on the system time deviation; and / or, The step of spatially registering the time-aligned low-resolution infrared image sequence and the high-resolution infrared image sequence to obtain the training sample pair includes: performing distortion correction processing on the time-aligned low-resolution infrared image sequence and the high-resolution infrared image sequence; obtaining feature points in the distortion-corrected low-resolution infrared image sequence and the high-resolution infrared image sequence, and matching multiple of the feature points; determining the transformation relationship between the high-resolution infrared image sequence and the image space where the low-resolution infrared image sequence is located based on the matched feature points; and performing spatial transformation on the high-resolution infrared image sequence based on the transformation relationship to obtain a high-resolution infrared image sequence that is spatially aligned with the low-resolution infrared image sequence, thereby obtaining the training sample pair.
7. The image processing method for vehicles according to claim 6, characterized in that, The distortion correction processing of the time-aligned low-resolution infrared image sequence and the high-resolution infrared image sequence includes: Obtain the camera intrinsic parameter matrix and distortion coefficients corresponding to the first sample camera and the second sample camera, respectively; Based on the camera intrinsic parameter matrix and distortion coefficients corresponding to the first sample camera, distortion correction processing is performed on the time-aligned low-resolution infrared image sequence; based on the camera intrinsic parameter matrix and distortion coefficients corresponding to the second sample camera, distortion correction processing is performed on the time-aligned high-resolution infrared image sequence.
8. The image processing method for vehicles according to any one of claims 5-7, characterized in that, After performing spatial registration processing on the time-aligned low-resolution infrared image sequence and the high-resolution infrared image sequence, the method includes: The training sample pairs are obtained by performing image enhancement processing on the spatially registered high-resolution infrared image sequence; or... The training sample pairs are obtained by performing at least one of the following processes on the spatially registered high-resolution infrared image sequence: histogram equalization, contrast stretching, noise suppression, and image sharpening.
9. The image processing method for vehicles according to any one of claims 1-8, characterized in that, The step of acquiring the second infrared image output by the infrared super-resolution model includes: Receive the user's first input, which is used to determine the rate adjustment command; In response to the first input, the upsampling magnification parameter of the infrared super-resolution model is adjusted to obtain a second infrared image output by the infrared super-resolution model that corresponds to the magnification adjustment command.
10. The image processing method for a vehicle according to any one of claims 1-9, characterized in that, The infrared super-resolution model is a lightweight model, comprising a shallow feature extraction module, a deep feature extraction module, and a feature recombination and upsampling module connected in sequence; or, The infrared super-resolution model includes a shallow feature extraction module, a deep feature extraction module, and a feature recombination and upsampling module connected in sequence. The shallow feature extraction module processes the first infrared image to obtain shallow features, the deep feature extraction module processes the shallow features to obtain deep features, and the feature recombination and upsampling module performs feature recombination and fusion on the deep features to obtain the second infrared image. The deep feature extraction module includes a reparameterized residual module.
11. An image processing apparatus for a vehicle, characterized in that, The device includes a first infrared camera and comprises: The first processing module is used to acquire a first infrared image of the scene where the vehicle is located based on the first infrared camera; The second processing module is used to input the first infrared image into the infrared super-resolution model and obtain the second infrared image output by the infrared super-resolution model. The image quality of the second infrared image is higher than that of the first infrared image. The third processing module is used to display the second infrared image.
12. A vehicle, characterized in that, include: First infrared camera; Display device; A controller, connected to both the first infrared camera and the display device, is configured to process a first infrared image captured by the first infrared camera according to the image processing method for vehicles as described in any one of claims 1-10, to obtain a second infrared image, and then display the second infrared image on the display device; or... Includes the image processing apparatus for a vehicle as described in claim 11.
13. The vehicle according to claim 12, characterized in that, include: An input module, connected to the controller, is used to receive a first input from a user, the first input being used to determine a magnification adjustment command, the magnification adjustment command being used by the controller to acquire a second infrared image corresponding to the magnification adjustment command; and / or, A voice recognition module, connected to the controller, is used to receive user voice commands and process the voice commands into a magnification adjustment command. The magnification adjustment command is used by the controller to acquire a second infrared image corresponding to the magnification adjustment command; and / or, A physical interaction module, connected to the controller, is used to generate the magnification adjustment command based on the user's physical interaction operation. The magnification adjustment command is used by the controller to acquire a second infrared image corresponding to the magnification adjustment command. And / or, A roller mounted on the steering wheel of the vehicle is used to generate the magnification adjustment command based on the user's scrolling operation. The magnification adjustment command is used by the controller to acquire a second infrared image corresponding to the magnification adjustment command.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the image processing method for vehicles as described in any one of claims 1-10.
15. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the image processing method for a vehicle as described in any one of claims 1-10.