Target perception method and device based on image enhancement, equipment and storage medium
By enhancing the brightness and optimizing the transmittance of image data under backlight and fog conditions, the image data quality is restored and fused with point cloud data, thus solving the problem of image data quality degradation under backlight and fog conditions and improving the performance of advanced driver assistance systems.
Patent Information
- Application Number
- CN202511328135.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-12-09
AI Technical Summary
In backlit and foggy conditions, the quality of image data captured by the camera deteriorates, affecting the fusion quality of point cloud data generated by the advanced driver assistance system and radar, resulting in a decline in the overall system performance.
By enhancing image brightness and optimizing image transmittance, backlight restoration and dehazing restoration image data are recovered and fused with point cloud data to improve image data quality.
In backlit and foggy conditions, the quality of image data captured by the camera is improved, enhancing the overall performance of the advanced driver assistance system.
Smart Images

Figure CN121095129A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of driving assistance, and in particular to a target perception method and device based on image enhancement, equipment and storage medium. BACKGROUND
[0002] Advanced Driver Assistance Systems (ADAS) use a series of sensors such as radar and cameras to monitor the vehicle's surroundings in real time, and warn the driver when potential dangers are detected, or even automatically perform operations in certain emergency situations, thereby greatly improving the safety and comfort of driving.
[0003] The implementation of current advanced driving assistance system functions is highly dependent on the fusion of point cloud data generated by radar and image data captured by cameras. By fusing these data, the advanced driving assistance system can accurately detect the traffic conditions in front and identify the type and motion state of the perceived target. Therefore, the fusion quality between point cloud data and image data directly determines the work efficiency of the advanced driving assistance system.
[0004] However, in the case of backlight and fog, the quality of image data captured by the camera will decrease significantly. This not only affects the accuracy of the image data itself, but also interferes with the fusion between the image data and the point cloud data, ultimately leading to a decline in the overall performance of the advanced driving assistance system. SUMMARY
[0005] The present application provides a target perception method and device based on image enhancement, equipment and storage medium, which realizes improving the quality of image data captured by the camera in the case of backlight and fog, and improves the overall performance of the advanced driving assistance system through the fusion between the backlight restored image, the dehazing restored image and the point cloud data.
[0006] The first aspect of the present application provides a target perception method based on image enhancement, which comprises:
[0007] obtaining point cloud data generated by radar and original image data captured by a camera;
[0008] determining whether the original image data is captured by the camera in a backlight and fog environment;
[0009] when the original image data is captured by the camera in a backlight and fog environment, performing image brightness value enhancement processing on the original image data to obtain a backlight restored image, and performing image transmittance optimization processing on the original image data to obtain a dehazing restored image;
[0010] The back light recovery image, the defogging recovery image and the point cloud data are fused to obtain fused data, and a perception target is identified according to the fused data to obtain a type and a motion state of the perception target.
[0011] In a possible design, the format of the original image data is an RGB format.
[0012] The original image data is subjected to image brightness value enhancement processing to obtain a back light recovery image, including:
[0013] The original image data is converted from the RGB format to an HSV format, and an original brightness component is extracted from the original image data in the HSV format.
[0014] The original brightness component is subjected to non-overlapping block processing according to a preset block scale to obtain a plurality of image blocks.
[0015] A part of the plurality of image blocks is subjected to image block brightness value enhancement processing, and the plurality of image blocks before enhancement and the plurality of image blocks after enhancement are subjected to image block merging processing to obtain an updated brightness component.
[0016] The updated brightness component is subjected to color space inverse conversion processing to obtain updated image data in the RGB format.
[0017] The updated image data is subjected to contrast expansion processing by a preset contrast expansion function to obtain the back light recovery image.
[0018] In a possible design, the part of the plurality of image blocks is subjected to image block brightness value enhancement processing, including:
[0019] An image main body region in which a perception target is located is determined from the original brightness component, and a plurality of candidate image blocks that are completely located within the image main body region are determined from the plurality of image blocks.
[0020] A brightness median value is obtained according to a brightness value of each candidate image block, and a brightness threshold interval is obtained according to the brightness median value and a preset adjustment parameter; the brightness median value is located within the brightness threshold interval.
[0021] A plurality of target image blocks in which a brightness value is located within the brightness threshold interval are determined from the plurality of candidate image blocks, and each target image block is subjected to image block brightness value enhancement processing.
[0022] In a possible design, a plurality of block scales are preset.
[0023] The first block scale is any one of the plurality of block scales, and for the first block scale, the updated image data is subjected to contrast expansion processing by the preset contrast expansion function to obtain the back light recovery image, including:
[0024] The first block scale corresponding to the updated image data is subjected to contrast expansion processing through a preset contrast expansion function, to obtain a candidate inverse light restoration image corresponding to the first block scale;
[0025] When the first block scale is any one of the multiple block scales except the last one, the candidate inverse light restoration image corresponding to the first block scale is determined as the original image data corresponding to the second block scale; wherein the second block scale is located next to the first block scale;
[0026] When the first block scale is the last one of the multiple block scales, the candidate inverse light restoration image with the minimum distortion degree is determined as the inverse light restoration image from the candidate inverse light restoration images corresponding to each block scale.
[0027] In a possible design, the format of the original image data is an RGB format;
[0028] The original image data is subjected to image transmittance optimization processing to obtain a defogging restoration image, comprising:
[0029] The original image data in the RGB format is converted into a grayscale image;
[0030] A gradient weight map is obtained according to the gradient amplitude of the grayscale image, and a brightness weight map is obtained according to the grayscale value of the grayscale image;
[0031] The gradient weight map and the brightness weight map are subjected to dynamic weight fusion processing to obtain a dynamic weight map;
[0032] An atmospheric light value is determined from the original image data; wherein the atmospheric light value refers to an RGB channel vector representing atmospheric light in the original image data;
[0033] An original transmittance map of the original image data is obtained according to the atmospheric light value through a preset dark channel prior algorithm, and the original transmittance map is subjected to optimization processing through the dynamic weight map to obtain an updated transmittance map;
[0034] The original image data is subjected to defogging processing according to the updated transmittance map to obtain a defogging restoration image.
[0035] In a possible design, the atmospheric light value is determined from the original image data, comprising:
[0036] A plurality of candidate pixel points satisfying a preset condition in terms of brightness and saturation are determined from the original image data;
[0037] A target pixel point with the maximum RGB channel vector is determined from the plurality of candidate pixel points;
[0038] The atmospheric light value is obtained according to the RGB channel vector of the target pixel point.
[0039] In a possible design, the back light recovery image, the de-fog recovery image and the point cloud data are fused to obtain fused data, including:
[0040] The back light enhancement features are extracted from the back light recovery image, the de-fog enhancement features are extracted from the de-fog recovery image, and the geometric features are extracted from the point cloud data; wherein the back light enhancement features include color distribution features, texture features and brightness features, and the de-fog enhancement features include contrast features, detail features and color saturation features.
[0041] The back light enhancement features and the de-fog enhancement features are fused to obtain image features, and the image features and the geometric features are fused to obtain the fused data.
[0042] A second aspect of the present application provides a target perception device based on image enhancement, which comprises:
[0043] A data acquisition module is configured to acquire point cloud data generated by a radar and original image data captured by a camera.
[0044] An environment judgment module is configured to judge whether the original image data is captured by the camera in a back light and fog environment.
[0045] An image processing module is configured to perform image brightness value enhancement processing on the original image data to obtain a back light recovery image, and perform image transmittance optimization processing on the original image data to obtain a de-fog recovery image, when the original image data is captured by the camera in the back light and fog environment.
[0046] A feature fusion module is configured to fuse the back light recovery image, the de-fog recovery image and the point cloud data to obtain fused data, and identify a perception target according to the fused data to obtain a type and a motion state of the perception target.
[0047] A third aspect of the present application provides an electronic device, comprising a memory and a processor in communication connection with the memory.
[0048] The memory stores computer execution instructions.
[0049] The processor, when executing the computer execution instructions stored in the memory, is configured to implement the target perception method based on image enhancement of any one of the first aspect.
[0050] A fourth aspect of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores computer execution instructions, and the computer execution instructions, when executed by a processor, are configured to implement the target perception method based on image enhancement of any one of the first aspect.
[0051] The fifth aspect of the present application provides a computer program product comprising a computer program, which, when executed by a processor, is configured to implement the image enhancement-based target perception method of any one of the first aspect.
[0052] The present application provides an image enhancement-based target perception method, device, equipment and storage medium. The method comprises: acquiring point cloud data and original image data; determining whether the original image data is captured by a camera in a backlight and fog environment; if so, performing image brightness value enhancement processing on the original image data to obtain a backlight recovery image, and performing image transmittance optimization processing on the original image data to obtain a de-fog recovery image; fusing the backlight recovery image, the de-fog recovery image and the point cloud data to obtain fused data, and identifying a perception target according to the fused data to obtain a type and a motion state of the perception target. The following technical effects are achieved: through image brightness value enhancement processing and image transmittance optimization processing, the quality of image data captured by the camera in a backlight and fog environment is improved; through fusion among the backlight recovery image, the de-fog recovery image and the point cloud data, the overall performance of an advanced driving assistance system is improved; determining whether the original image data is captured by the camera in a backlight and fog environment avoids image enhancement on the original image data in a backlight and / or fog-free environment. BRIEF DESCRIPTION OF DRAWINGS
[0053] In order to more clearly illustrate the technical solutions of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0054] Figure 1 The flowchart of the image enhancement-based target perception method provided by the embodiments of the present application is shown in the figure.
[0055] Figure 2 The structure diagram of the image enhancement-based target perception device provided by the embodiments of the present application is shown in the figure.
[0056] Figure 3 The structure diagram of the electronic device provided by the embodiments of the present application is shown in the figure.
[0057] REFERENCE SIGNS:
[0058] 210-data acquisition module; 220-environment determination module; 230-image processing module; 240-feature fusion module;
[0059] 310-processor; 320-memory; 330-communication component; 340-bus. DETAILED DESCRIPTION
[0060] The exemplary embodiments will be described in detail herein with reference to the attached drawings. In the following description, like reference numerals refer to like elements, unless the context clearly dictates otherwise. The following description is not meant to limit the application to all of the embodiments set forth herein. Rather, the following description is meant to provide examples of apparatus and methods consistent with the application as detailed in the appended claims.
[0061] In the present application, the terms "first", "second", etc. are used to distinguish between the same or similar items or elements having substantially the same function and role. Those skilled in the art can understand that the terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. do not necessarily mean different. It should be noted that the words "exemplary" or "for example" in the present application are used to represent an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the use of the words "exemplary" or "for example" is intended to present the relevant concept in a specific manner. In the present application, "at least one" means one or more, and "multiple" means two or more.
[0062] It should be noted that "at the time of" in the present application can be at the moment when a certain condition occurs, or within a certain period of time after the occurrence of a certain condition, which is not specifically limited in the present application. In addition, the image enhancement-based target perception method provided in the present application is only an example, and the image enhancement-based target perception method can include more or less content. The user information (including but not limited to user device information and user personal information, etc.) and data (including but not limited to data for analysis, stored data and displayed data, etc.) involved in one or more embodiments of the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards, and provide corresponding operation portal for user to choose authorization or refusal.
[0063] In order to clearly describe the technical solutions of the present application, the following briefly introduces some terms and technologies involved in the present application:
[0064] Radar: refers to a technology that measures the distance, speed and angle of a target by transmitting and receiving radio waves. In advanced driver assistance systems, radar is used to monitor the environment around the vehicle in real time, helping to identify the type and motion state of the perceived target.
[0065] Point Cloud Data: Refers to the dataset generated by radar or other scanning devices. Point cloud data is composed of a series of points in space, which represent the position information of the object surface. In advanced driving assistance systems, point cloud data is used to accurately depict the three-dimensional structure of the vehicle's surroundings.
[0066] Camera: Refers to a device used to capture images or continuous video streams. In advanced driving assistance systems, cameras are used for visual perception, allowing the system to understand the surrounding environment and make appropriate decisions by analyzing the image data captured by the camera.
[0067] Image Data: Refers to the images captured and digitized by the camera. In advanced driving assistance systems, image data is processed to identify various objects and conditions.
[0068] Backlight: Refers to the situation where the light source is located behind the subject, opposite the direction of the camera. Images taken in backlight conditions will make the subject appear dark, lose details, and reduce contrast.
[0069] Fog: Refers to the situation where a large number of small water droplets or ice crystals suspended in the atmosphere cause visual obstruction. Images taken in foggy conditions will reduce the quality of the images captured by the camera, making distant objects appear blurry.
[0070] The technical solutions of the present application will be described in detail below with specific examples. The following specific examples can be combined with each other, and the same or similar concepts or processes may not be described again in some examples. The present application will be described below with reference to the accompanying drawings.
[0071] In order to clearly understand the technical solutions of the present application, the prior art solutions will be described in detail first.
[0072] Advanced driving assistance systems use a series of sensors such as radar and cameras to monitor the vehicle's surroundings in real time and warn the driver when potential dangers are detected, or even automatically perform operations in certain emergency situations, thereby greatly improving driving safety and comfort.
[0073] The current implementation of advanced driver assistance system functions highly depends on the fusion of point cloud data generated by radar and image data captured by camera. By fusing these data, the advanced driver assistance system can accurately detect the front traffic condition and identify the type and motion state of the perceived target. For example, when the ego vehicle is getting closer to the perceived target, the advanced driver assistance system, such as the automatic emergency braking system (AEBS), will give a warning to the driver according to the distance, and if the driver does not respond, the system will reduce the collision or avoid the collision through emergency braking. Therefore, the fusion quality between the point cloud data and the image data directly determines the work efficiency of the advanced driver assistance system.
[0074] In the fusion process between the point cloud data and the image data, the radar and the camera work independently to collect the point cloud data and the image data of the surrounding environment, respectively. The camera can identify the boundary, color and texture of the perceived target, and the radar can identify the speed and distance of the perceived target. Then, the point cloud data and the image data are matched to determine whether they point to the same target. If so, a data layer or image layer fusion strategy is adopted to fuse the point cloud data and the image data, and based on the fused data, a decision basis is provided for the advanced driver assistance system.
[0075] However, for the advanced driver assistance system equipped with a single radar and a single camera, the quality of the image data captured by the camera will decrease significantly in poor optical environment, such as backlight and fog. This is because the backlight and fog area is often the area where the perceived target is located, which usually has the characteristics of low visual quality, incomplete detail expression and serious color loss, while the background area outside this area usually has the characteristics of overexposure, detail loss and poor contrast. The radar, such as laser radar or millimeter wave radar, has lower requirements for optical environment, but it provides point cloud data, which has low discrimination for different perceived targets. This not only affects the accuracy of the image data itself, but also interferes with the fusion between the image data and the point cloud data, ultimately leading to the decline of the overall performance of the advanced driver assistance system.
[0076] Therefore, in order to solve the above technical problems, it is found in the research that in order to solve the problem, the embodiment of the present application provides a method for image brightness value enhancement processing and image transmittance optimization processing of image data, so that the processed image realizes backlight restoration and dehazing restoration, and then the point cloud data and the processed image data are fused. In the case of backlight and fog, the quality of the image data captured by the camera is improved, and through the fusion between the backlight restoration image, the dehazing restoration image and the point cloud data, the overall performance of the advanced driver assistance system is improved.
[0077] Based on the above creative findings, the technical solutions of the present application are proposed. The embodiments of the present application will be introduced below in conjunction with the drawings of the specification.
[0078] Figure 1 The flowchart of the target perception method based on image enhancement provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, in the embodiments of the present application, the execution subject can be a target perception device based on image enhancement, which can be located in an electronic device, which can be a data processing server. Then, the target perception method based on image enhancement provided by the embodiments of the present application includes the following steps: Figure 1
[0079] S101, acquiring point cloud data generated by a radar and raw image data captured by a camera.
[0080] Specifically, the target perception device based on image enhancement scans the environment through a laser radar or a millimeter wave radar to generate three-dimensional point cloud data to record the spatial coordinates and motion information of the perceived target. In addition, the target perception device based on image enhancement takes pictures of the environment through a camera to capture raw image data to record the image information of the perceived target.
[0081] S102, judging whether the raw image data is captured by the camera in a backlit and foggy environment.
[0082] Specifically, considering that in a backlit environment, the subject often presents a clear outline, its brightness is usually much darker than the background, and the color saturation of some transparent or semi-transparent objects is greater. Therefore, the target perception device based on image enhancement can judge whether the raw image data is captured by the camera in a backlit environment according to the outline of the subject, the brightness difference between the subject and the background, and the color saturation change of the raw image data.
[0083] Considering that in a foggy environment, the near view is usually much clearer than the far view, the contrast of the far view is smaller, and the overall color tone of the image is biased towards cold tone. Therefore, the target perception device based on image enhancement can judge whether the raw image data is captured by the camera in a foggy environment according to the spatial sense and contrast between the near view and the far view, and the color tone shift of the raw image data.
[0084] When the raw image data is captured by the camera in a backlit and foggy environment, the target perception device based on image enhancement continues to execute the target perception method based on image enhancement of the embodiments of the present application, that is, continues to execute S103. When the raw image data is captured by the camera in a front light and / or non-foggy environment, the target perception method based on image enhancement of the embodiments of the present application is not continued, that is, the raw image data and the point cloud data are directly fused, and the perceived target is recognized to obtain the type and motion state of the perceived target.
[0085] S103, when the original image data is captured by the camera in the backlight and fog environment, the original image data is processed by image brightness value enhancement, and the backlight recovery image is obtained, and the original image data is processed by image transmittance optimization, and the defogging recovery image is obtained.
[0086] Specifically, the image brightness value enhancement processing refers to adjusting the pixel brightness distribution of the original image data, improving the light and dark contrast of the whole or local area, and restoring the details lost in the backlight environment. The target perception device based on image enhancement enhances the dark details of the whole or local area of the original image data through histogram equalization, Retina-Cortex (Retinex) or deep learning model, and obtains the backlight recovery image.
[0087] The image transmittance optimization processing refers to adjusting the transmittance of the original image data, reducing the influence of fog on the original image data, and restoring the real color and contrast in the fog environment. The transmittance refers to the proportion of light penetrating the fog to reach the camera. The target perception device based on image enhancement calculates the transmittance and atmospheric light value through dark channel priori or atmospheric scattering model, and performs defogging processing on the original image data to obtain the defogging recovery image.
[0088] S104, the backlight recovery image, the defogging recovery image and the point cloud data are fused to obtain the fusion data, and the type and motion state of the perception target are obtained according to the fusion data.
[0089] Specifically, data fusion refers to the alignment and integration of data from multiple sensors in space and time to generate more comprehensive and reliable environmental perception results. The backlight recovery image, the defogging recovery image and the point cloud data are fused, which can be image-level fusion, target-level fusion or signal-level fusion. Image-level fusion refers to taking the backlight recovery image and the defogging recovery image as the main body, converting the point cloud data into image features, and then fusing with the backlight recovery image and the defogging recovery image; target-level fusion refers to weighting the comprehensive reliability of the backlight recovery image, the defogging recovery image and the point cloud data, and then adaptively searching and matching after fusion output with precision calibration information; signal-level fusion refers to fusing the data sources output by the electronic control unit (Electronic Control Unit, ECU) of the camera and radar respectively. The signal-level fusion data loss is the smallest and the reliability is the highest, but a large amount of calculation is required.
[0090] After fusion, based on the fused data, pre-stored target detection algorithms identify the perceived targets and assign them specific category labels, such as cars or pedestrians. Then, algorithms such as optical flow or point cloud tracking are used to continuously track the perceived targets to accurately calculate their motion states, including speed, acceleration, and direction of motion.
[0091] This application provides a target perception method based on image enhancement. The method includes: acquiring point cloud data and original image data; determining whether the original image data was captured by a camera in a backlit and foggy environment; if so, performing image brightness enhancement processing on the original image data to obtain a backlit restored image, and performing image transmittance optimization processing on the original image data to obtain a dehazed restored image; fusing the backlit restored image, the dehazed restored image, and the point cloud data to obtain fused data, and identifying the perceived target based on the fused data to obtain the target's type and motion state. This achieves the following technical effects: by enhancing image brightness and optimizing image transmittance, the quality of image data captured by the camera is improved in backlit and foggy conditions; by fusing the backlit restored image, the dehazed restored image, and the point cloud data, the overall performance of the advanced driver assistance system is improved; and by determining whether the original image data was captured by the camera in a backlit and foggy environment, image enhancement of the original image data is avoided in front-lit and / or fog-free environments.
[0092] In one possible design, the image enhancement-based target perception method provided in this application embodiment is... Figure 1 This embodiment further refines the target perception method based on image enhancement. The original image data is in RGB format, i.e., the image color mode, composed of three colors: red, green, and blue, used for digital image processing such as screen display. The target perception method based on image enhancement provided in this embodiment includes the following steps.
[0093] S201. Acquire point cloud data generated by radar and raw image data captured by camera.
[0094] S202. Determine whether the original image data was captured by the camera in a backlit and foggy environment.
[0095] When the original image data was captured by the camera in a backlit and foggy environment, steps S203 and S204 continue to be executed. The execution order of S203, S204, and S205 can be as follows: S203 is executed first, then S204, and finally S205; S204 is executed first, then S203, and finally S205; or S203 and S204 are executed simultaneously, and S205 is executed last.
[0096] S203, performing image brightness value enhancement processing on the original image data to obtain a back-light recovery image.
[0097] In a possible design, S203 includes:
[0098] S2031, converting the original image data from an RGB format to an HSV format, and extracting an original brightness component from the original image data in the HSV format.
[0099] Specifically, the original image data captured by the camera is generally output in an RGB format after internal processing. A set of red, green, and blue colors is a minimum display unit, and any color on the image can be recorded and expressed by a set of RGB values. However, the RGB format has poor uniformity, and it is difficult to accurately infer the RGB values for a certain color. Therefore, the RGB format is suitable for display systems but is not very suitable for image processing.
[0100] The HSV format is more suitable for image processing and is composed of a hue (Hue), a saturation (Saturation), and a value (Value) color space. The V component directly reflects the image brightness information.
[0101] For the original image data in the RGB format, the middle axis of the RGB three-dimensional coordinates is erected and flattened to form a conical model of the HSV format, and the original image data in the HSV format is obtained. Then, the V component is extracted from the original image data in the HSV format to obtain the original brightness component, so as to separate the color and brightness information and facilitate independent processing.
[0102] S2032, performing non-overlapping block processing on the original brightness component according to a preset block scale to obtain a plurality of image blocks.
[0103] Specifically, the block scale can be one of a plurality of preset scales, for example, one of 32x32, 16x16, 8x8, 6x6, 4x4, and 2x2. The target perception device based on image enhancement performs non-overlapping block processing on the original brightness component according to the preset block scale, divides the original brightness component into rectangular regions that do not overlap with each other, each rectangular region corresponds to an image block, and each pixel belongs to only one image block, so as to reduce the computational complexity and retain local features.
[0104] S2033, performing image block brightness value enhancement processing on a part of the plurality of image blocks.
[0105] In a possible design, S2033 includes:
[0106] S20331, determine an image subject region where the perception target is located from the original luminance component, and determine a plurality of candidate image blocks which are completely located within the image subject region from the plurality of image blocks.
[0107] Specifically, the image subject region refers to a core region containing the perception target such as a pedestrian or a vehicle, which can be determined by a target detection algorithm. Then, the perception target in the original image data is identified by the target detection algorithm, and its position in the original image data is determined, and then the position is mapped to the original luminance component to determine the image subject region where the perception target is located. Then, a plurality of candidate image blocks are determined from the blocking result, i.e. the plurality of image blocks, to exclude the background region and reduce the calculation amount.
[0108] S20332, obtain a luminance median value according to the luminance values of each candidate image block, and obtain a luminance threshold interval according to the luminance median value and a preset adjustment parameter; wherein the luminance median value is located within the luminance threshold interval.
[0109] Specifically, the luminance median value refers to the median of the luminance values of each candidate image block, which is used to reflect the typical luminance level. The luminance threshold interval refers to a dynamic range including the luminance median value, which is determined by the adjustment parameter to filter the target image blocks that need to be enhanced.
[0110] The calculation of the luminance threshold interval can be adding or subtracting an adjustment parameter based on the luminance median value. For example, when the luminance median value is 100, the adjustment parameter is determined to be ±20, and then the luminance threshold interval is 100-20 to 100+20, i.e. 80 to 120.
[0111] The calculation of the luminance threshold interval can also be scaling based on the luminance median value. For example, when the luminance median value is 100, the adjustment parameter is determined to be 0.2, and then the luminance threshold interval is 100×(1-0.2) to 100×(1+0.2), i.e. 80 to 120.
[0112] It should be noted that when setting the adjustment coefficients of the upper limit and the lower limit of the interval, the two can be the same or different, and the embodiments of the present application do not limit this.
[0113] S20333, determine a plurality of target image blocks whose luminance values are located within the luminance threshold interval from the plurality of candidate image blocks, and perform image block luminance value enhancement processing on each target image block.
[0114] Specifically, the target image block refers to the candidate image block whose luminance value is located within the threshold interval, which is considered to contain the details of the perception target that need to be enhanced. Then, the image block luminance value enhancement processing is performed on each target image block by algorithms such as gamma (Gamma) correction, histogram equalization or linear transformation, so as to highlight the details of the perception target.
[0115] The technical effect of the embodiment of the application is that a plurality of target image blocks are screened from the original luminance component, and image block luminance value enhancement processing is performed on each target image block, thereby improving the utilization rate of computing resources, ensuring that the perception target is processed preferentially, and meanwhile, overexposure or noise amplification caused by global luminance enhancement is avoided.
[0116] After S20333 is executed, S2034 is executed.
[0117] S2034, performing image block merging processing on the plurality of image blocks before enhancement and the plurality of image blocks after enhancement, to obtain an updated luminance component.
[0118] Specifically, based on the order of the above block processing, the plurality of image blocks before enhancement and the plurality of image blocks after enhancement are spliced according to the original positions, to obtain a complete V component, that is, the updated luminance component.
[0119] S2035, performing color space inverse conversion processing on the updated luminance component, to obtain updated image data in RGB format.
[0120] S2036, performing contrast expansion processing on the updated image data by using a preset contrast expansion function, to obtain a back light restoration image.
[0121] Specifically, the contrast expansion function refers to stretching the pixel value distribution range by using linear or nonlinear stretching, to enhance the global contrast of the updated image data, so as to improve the visibility of details in the back light area.
[0122] The technical effect of the embodiment of the application is that by separating the luminance component and performing local enhancement, the dark detail loss caused by back light is improved, and meanwhile, overexposure in the highlight area is avoided; by block processing, the amount of calculation is reduced, and by contrast expansion, the global level is enhanced, so that the outline of the perception target is clearer.
[0123] In a possible design, a plurality of block scales are preset, for example, six block scales of 32x32, 16x16, 8x8, 6x6, 4x4 and 2x2 are preset, covering processing requirements from coarse granularity to fine granularity.
[0124] The first block scale is any one of the plurality of block scales, and for the first block scale, S2036 includes:
[0125] The updated image data corresponding to the first block scale is subjected to contrast expansion processing by using the preset contrast expansion function, to obtain a candidate back light restoration image corresponding to the first block scale;
[0126] When the first block size is any one of the plurality of block sizes except the last one, the candidate inverse light restoration image corresponding to the first block size is determined as the original image data corresponding to the second block size, where the second block size is next to the first block size.
[0127] When the first block size is the last one of the plurality of block sizes, the candidate inverse light restoration image with the minimum distortion is determined as the inverse light restoration image from the candidate inverse light restoration images corresponding to each block size.
[0128] Specifically, starting from the block size of 32x32, the image brightness value enhancement processing is iteratively performed on the original brightness component, and S2031 to S2036 are repeatedly executed to generate the candidate inverse light restoration image corresponding to each block size.
[0129] When the first block size is any one of 32x32, 16x16, 8x8, 6x6 and 4x4, the candidate inverse light restoration image corresponding to 32x32 is determined as the original image data corresponding to 16x16, and so on until the candidate inverse light restoration image corresponding to 4x4 is determined as the original image data corresponding to 2x2.
[0130] When the first block size is 2x2, the iteration processing is ended to obtain the candidate inverse light restoration image corresponding to each block size, and then the distortion of the candidate inverse light restoration image corresponding to each block size is calculated by using an algorithm such as Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM) or Feature Similarity Index (FSIM). The candidate inverse light restoration image corresponding to 32x32 has a serious distortion and a visual effect worse than the original image. The candidate inverse light restoration images corresponding to 16x16 and 8x8 have a larger improvement in brightness compared with the original image, and the main body region of the image can be clearly seen, but the non-main body region of the image is still seriously distorted. The candidate inverse light restoration images corresponding to 6x6 and 4x4 have clear lines in the main body region of the image, and the brightness is good compared with the original image, but the non-main body region of the image has a large distortion. The candidate inverse light restoration image corresponding to 2x2 has a good visual effect in the main body region and the non-main body region of the image except for a small degree of distortion in the edge region compared with the candidate inverse light restoration images corresponding to other block sizes.
[0131] The candidate inverse light restoration image with the minimum distortion is determined as the inverse light restoration image from the candidate inverse light restoration images corresponding to each block size. For example, the candidate inverse light restoration image corresponding to 2x2 is determined as the inverse light restoration image.
[0132] The technical effect of the embodiments of the present application is that: by multiple block scales, different back light intensities are adapted, and by distortion degree, the back light restoration image is determined from the candidate back light restoration images corresponding to each block scale, reducing the missed detection and false detection of the perceived target.
[0133] S204, performing image transmittance optimization processing on the original image data to obtain a defogging restoration image.
[0134] In a possible design, S204 includes:
[0135] S2041, converting the original image data in RGB format into a gray scale image.
[0136] Specifically, the gray scale image refers to a single channel image containing only brightness information, which is used to simplify the calculation of gradient and brightness weight.
[0137] S2042, obtaining a gradient weight map according to the gradient amplitude of the gray scale image, and obtaining a brightness weight map according to the gray scale value of the gray scale image.
[0138] Specifically, the gradient amplitude refers to the intensity of the local brightness change of the gray scale image, which is used to reflect the texture complexity. The target perception device based on image enhancement calculates the gradient amplitude of the gray scale image through Sobel operator, Prewitt operator or Canny edge detection algorithm, gives high weight to the high gradient area such as vehicle edge, and obtains the gradient weight map to preferentially retain details. At the same time, the target perception device based on image enhancement directly uses the gray scale value of the gray scale image to obtain the brightness weight map, and gives higher weight to the low brightness area such as foggy dim area, so as to focus on defogging.
[0139] S2043, performing dynamic weight fusion processing on the gradient weight map and the brightness weight map to obtain a dynamic weight map.
[0140] Specifically, the gradient weight map and the brightness weight map are fused according to a preset proportion, for example, the gradient weight proportion is 0.6 and the brightness weight proportion is 0.4, to obtain a dynamic weight map considering texture and brightness, which is used to balance detail retention and defogging intensity.
[0141] S2044, determining an atmospheric light value from the original image data; wherein the atmospheric light value refers to an RGB channel vector representing atmospheric light in the original image data.
[0142] In a possible design, S2044 includes:
[0143] S20441, determining a plurality of candidate pixel points from the original image data, the brightness and saturation of which satisfy a preset condition.
[0144] Specifically, the candidate pixel points refer to pixel points with high brightness and low saturation in the original image data, which can represent atmospheric light regions such as the sky or a heavy fog area. The preset condition can be a preset number before brightness sorting, such as the top 0.1%, and the saturation is lower than a first preset threshold, such as 0.3. The preset condition can also be that the brightness is higher than a second preset threshold, such as higher than 220, and the saturation is lower than the first preset threshold. Through the preset condition, a plurality of candidate pixel points can be screened out, wherein the first preset threshold is used to exclude colored objects such as vehicles or traffic signs, and the second preset threshold is used to exclude dark regions.
[0145] S20442, determining a target pixel point with the maximum RGB channel vector from the plurality of candidate pixel points.
[0146] S20443, obtaining an atmospheric light value according to the RGB channel vector of the target pixel point.
[0147] Specifically, the RGB channel vector, that is, the RGB value, is the pixel point with the maximum value, that is, the pixel point that is the brightest and closest to white. From the plurality of candidate pixel points, the pixel point is taken as the target pixel point, and the RGB channel vector of the target pixel point is taken as the atmospheric light value, which is used for subsequent transmission rate map estimation.
[0148] The technical effect of the embodiment of the application is that the candidate pixel points are screened by combining the brightness and saturation conditions to avoid misjudging high-brightness objects as atmospheric light regions.
[0149] After S20443 is executed, S2045 is continued.
[0150] S2045, obtaining an original transmission rate map of the original image data according to the atmospheric light value through a preset dark channel prior algorithm, and optimizing the original transmission rate map through a dynamic weight map to obtain an updated transmission rate map.
[0151] Specifically, the dark channel prior algorithm refers to a dehazing algorithm for estimating a transmission rate map based on the statistical law that the intensity of at least one color channel of a local region of a haze-free image tends to 0. The original transmission rate map is estimated through the dark channel prior algorithm, and the transmission rate value of the original transmission rate map is locally adjusted, such as increasing the transmission rate of the edge region to enhance the details, to obtain an updated transmission rate map.
[0152] S2046, performing dehazing processing on the original image data according to the updated transmission rate map to obtain a dehazed restored image.
[0153] Specifically, the original image data is dehazed through an atmospheric scattering model according to the updated transmission rate map to obtain a dehazed restored image.
[0154] The technical effect of the embodiments of the present application is that the fog removal strength is adapted to the image content through the dynamic weight map; the color deviation caused by the foggy weather is corrected through the estimation of the atmospheric light value, and the real scene color is restored.
[0155] After S2046 is executed, S205 is executed.
[0156] S205, fusing the back light restoration image, the de-fog restoration image and the point cloud data to obtain the fusion data.
[0157] In a possible design, S205 includes:
[0158] S2051, extracting the back light enhancement feature from the back light restoration image, extracting the de-fog enhancement feature from the de-fog restoration image, and extracting the geometric feature from the point cloud data.
[0159] Specifically, the back light enhancement feature includes a color distribution feature, a texture feature and a brightness feature. The color distribution feature refers to the statistical distribution of the RGB channel in the back light restoration image, such as a histogram, and is used to reflect the color balance after back light restoration. The texture feature refers to the local texture information extracted by the local binary pattern (LBP) or the Gabor filter, and is used to reflect the surface details of the perceived target such as the license plate. The brightness feature refers to the average brightness and brightness variance of the global or local region of the image, and is used to measure the exposure uniformity after back light restoration.
[0160] The de-fog enhancement feature includes a contrast feature, a detail feature and a color saturation feature. The contrast feature refers to the intensity difference of the pixels in the local window, such as the local variance, and is used to reflect the distinction between the perceived target and the background after de-fogging. The detail feature refers to the edge information extracted by the Canny edge detection or the gradient amplitude, and is used to reflect the contour clarity of the perceived target after de-fogging. The color saturation feature refers to the mean value of the S channel in the HSV space, and is used to reflect the vividness of the color after de-fogging, such as the recovery of the blue color of the sky.
[0161] The geometric feature refers to the three-dimensional information extracted from the point cloud data, including the spatial coordinates, the reflection intensity and the normal vector. The spatial coordinates refer to the position of the perceived target in the three-dimensional space. The reflection intensity refers to the radar wave reflection intensity, and is used to reflect the material of the perceived target, such as metal material or plastic material. The normal vector refers to the surface orientation of the perceived target, and is used for shape analysis, such as ground plane fitting.
[0162] S2052, fusing the back light enhancement feature and the de-fog enhancement feature to obtain the image feature, and fusing the image feature and the geometric feature to obtain the fusion data.
[0163] Specifically, the back light enhancement feature and the defogging enhancement feature are fused by a weighted fusion or principal component analysis to obtain image features. Then, the image features and the geometric features are fused by data layer fusion, feature layer fusion and decision layer fusion to obtain fused data.
[0164] The technical effect of the embodiment of the application is that the back light restoration and the defogging restoration of the original image data and the fusion between the image features and the set features are realized through the fusion between the back light restoration image, the defogging restoration image and the point cloud data.
[0165] After S2052 is executed, S206 is continuously executed.
[0166] S206, identifying a perception target according to the fused data to obtain a type and a motion state of the perception target.
[0167] Figure 2 A structure diagram of a target perception device based on image enhancement provided by the embodiment of the application is shown in FIG. 2. Figure 2 As shown in FIG. 2, in the embodiment of the application, the target perception device based on image enhancement can be located in an electronic device. The target perception device based on image enhancement includes:
[0168] A data acquisition module 210 is configured to acquire point cloud data generated by a radar and original image data captured by a camera.
[0169] An environment judgment module 220 is configured to judge whether the original image data is captured by the camera in a back light and fog environment.
[0170] An image processing module 230 is configured to perform image brightness value enhancement processing on the original image data to obtain a back light restoration image and perform image transmittance optimization processing on the original image data to obtain a defogging restoration image when the original image data is captured by the camera in the back light and fog environment.
[0171] A feature fusion module 240 is configured to fuse the back light restoration image, the defogging restoration image and the point cloud data to obtain fused data, and identify a perception target according to the fused data to obtain a type and a motion state of the perception target.
[0172] The target perception device based on image enhancement provided by the embodiment of the application can execute the technical scheme of the method embodiment shown in FIG. 3. Figure 1 The implementation principle and the technical effect of the method embodiment shown in FIG. 3 are similar to those of the method embodiment shown in FIG. 1, and the embodiment of the application will not be described herein. Figure 1
[0173] Meanwhile, the target perception device based on image enhancement provided by the embodiment of the present application is further refined on the basis of the target perception device based on image enhancement provided by the previous embodiment.
[0174] In a possible design, the format of the original image data is an RGB format.
[0175] The image processing module 230 comprises:
[0176] The first conversion module is configured to convert the original image data from the RGB format to an HSV format, and extract an original luminance component from the original image data in the HSV format.
[0177] The block processing module is configured to perform non-overlapping block processing on the original luminance component according to a preset block size, to obtain a plurality of image blocks.
[0178] The first luminance enhancement module is configured to perform image block luminance value enhancement processing on a part of the plurality of image blocks, and perform image block merging processing on the plurality of image blocks before enhancement and the plurality of image blocks after enhancement, to obtain an updated luminance component.
[0179] The second conversion module is configured to perform color space inverse conversion processing on the updated luminance component, to obtain updated image data in the RGB format.
[0180] The extension processing module is configured to perform contrast extension processing on the updated image data by using a preset contrast extension function, to obtain a back light recovery image.
[0181] In a possible design, the first luminance enhancement module comprises:
[0182] The image block screening module is configured to determine an image main body region in which the target is located from the original luminance component, and determine a plurality of candidate image blocks that are completely located within the image main body region from the plurality of image blocks.
[0183] The interval determination module is configured to obtain a luminance median value according to the luminance value of each candidate image block, and obtain a luminance threshold interval according to the luminance median value and a preset adjustment parameter; and the luminance median value is located within the luminance threshold interval.
[0184] The second luminance enhancement module is configured to determine a plurality of target image blocks in which the luminance value is located within the luminance threshold interval from the plurality of candidate image blocks, and perform image block luminance value enhancement processing on each target image block.
[0185] In a possible design, a plurality of block sizes are preset.
[0186] The first sub-block scale is any one of a plurality of sub-block scales, and for the first sub-block scale, the extension processing module is configured to perform contrast extension processing on the updated image data corresponding to the first sub-block scale by using a preset contrast extension function, to obtain a candidate inverse light restoration image corresponding to the first sub-block scale.
[0187] When the first sub-block scale is any one of the plurality of sub-block scales except the last one, the candidate inverse light restoration image corresponding to the first sub-block scale is determined as the original image data corresponding to a second sub-block scale; and the second sub-block scale is located next to the first sub-block scale.
[0188] When the first sub-block scale is the last one of the plurality of sub-block scales, the candidate inverse light restoration image with the minimum distortion degree is determined as the inverse light restoration image from the candidate inverse light restoration images corresponding to each sub-block scale.
[0189] In a possible design, the format of the original image data is an RGB format.
[0190] The image processing module 230 includes:
[0191] The third conversion module is configured to convert the original image data in the RGB format into a grayscale image.
[0192] The first weight map determination module is configured to obtain a gradient weight map according to the gradient amplitude of the grayscale image, and obtain a brightness weight map according to the grayscale value of the grayscale image.
[0193] The second weight map determination module is configured to perform dynamic weight fusion processing on the gradient weight map and the brightness weight map, to obtain a dynamic weight map.
[0194] The atmospheric light value determination module is configured to determine an atmospheric light value from the original image data; the atmospheric light value refers to an RGB channel vector representing atmospheric light in the original image data.
[0195] The transmittance optimization module is configured to obtain an original transmittance map of the original image data according to the atmospheric light value by using a preset dark channel prior algorithm, and perform optimization processing on the original transmittance map by using the dynamic weight map, to obtain an updated transmittance map.
[0196] The defogging repair module is configured to perform defogging processing on the original image data according to the updated transmittance map, to obtain a defogging restoration image.
[0197] In a possible design, the atmospheric light value determination module includes:
[0198] The first pixel point screening module is configured to determine a plurality of candidate pixel points from the original image data, the brightness and saturation of which satisfy a preset condition.
[0199] The second pixel point screening module is configured to determine a target pixel point with the maximum RGB channel vector from the plurality of candidate pixel points.
[0200] The channel vector mapping module is configured to obtain the atmosphere light value according to the RGB channel vector of the target pixel point.
[0201] In a possible design, the feature fusion module 240 includes:
[0202] The first fusion module is configured to extract a back-light enhancement feature from the back-light restored image, extract a de-fog enhancement feature from the de-fog restored image, and extract a geometry feature from the point cloud data; wherein the back-light enhancement feature includes a color distribution feature, a texture feature and a brightness feature, and the de-fog enhancement feature includes a contrast feature, a detail feature and a color saturation feature.
[0203] The second fusion module is configured to perform feature fusion on the back-light enhancement feature and the de-fog enhancement feature to obtain an image feature, and perform feature fusion on the image feature and the geometry feature to obtain the fusion data.
[0204] The target perception apparatus based on image enhancement provided by the embodiments of the present application can perform Figure 1 The technical solutions of the method embodiments shown in the drawings have similar implementation principles and technical effects to Figure 1 The method embodiments shown in the drawings are similar, and the embodiments of the present application will not be described herein.
[0205] The embodiments of the present application further provide an electronic device, Figure 3 The structure of the electronic device provided by the embodiments of the present application is shown in the drawings. As Figure 3 The electronic device includes at least one processor 310 and a memory 320. The electronic device further includes a communication component 330. The processor 310, the memory 320 and the communication component 330 are connected through a bus 340.
[0206] In the specific implementation process, the at least one processor 310 executes the computer execution instructions stored in the memory 320, so that the at least one processor 310 is configured to implement the target perception method based on image enhancement of the above-mentioned embodiments.
[0207] The specific implementation process of the processor 310 can be referred to the above-mentioned method embodiments, and the implementation principles and technical effects are similar, and the embodiments of the present application will not be described herein.
[0208] In the above embodiments, it should be understood that the processor 310 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the application can be directly embodied as hardware processor execution, or executed by a combination of hardware and software modules in the processor.
[0209] The memory 320 can include a high-speed RAM memory, and can also include a non-volatile storage NVM, such as at least one disk memory.
[0210] The bus 340 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 340 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus 340 in the drawings of the present application does not limit to only one bus or one type of bus.
[0211] The functions implemented by the electronic device and the master device described above are introduced for the scheme provided by the embodiments of the present application. It can be understood that the electronic device or the master device includes the hardware structure and / or software modules corresponding to the execution of each function in order to implement the above functions. The units and algorithm steps of each example described in combination with the embodiments disclosed in the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is implemented in hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the technical solution of the embodiments of the present application.
[0212] The embodiments of the present application also provide a computer readable storage medium, and the computer readable storage medium stores computer execution instructions. When the computer execution instructions are executed by a processor, the computer execution instructions are used to implement the target perception method based on image enhancement of the above embodiments. In the specific implementation of the above target perception method based on image enhancement, each module can be implemented as a processor.
[0213] The aforementioned readable storage medium can be implemented by any type of volatile or nonvolatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special purpose computer.
[0214] An exemplary readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in an electronic device or a host device.
[0215] The embodiments of the present application also provide a computer program product, comprising a computer program, which is executed by a processor to implement the image enhancement-based target perception method of the above-mentioned embodiments.
[0216] The computer program is stored in a readable storage medium, and at least one processor can read the computer program from the readable storage medium, and the at least one processor executes the computer program to perform the scheme provided by any of the above-mentioned embodiments.
[0217] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments of the application can be completed by program instructions related to hardware. The aforementioned program can be stored in a computer readable storage medium. The program is executed to perform the steps of the above-mentioned method embodiments; and the aforementioned storage medium includes ROM, RAM, magnetic disk or optical disk and various media that can store program codes.
[0218] So far, the technical scheme of the present application has been described in combination with the preferred embodiments shown in the drawings, but those skilled in the art can easily understand that the protection scope of the present application is obviously not limited to these specific embodiments, and the above embodiments are only used to illustrate the technical scheme of the present application, but not to limit it; although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that they can still modify the technical scheme recorded in the above-mentioned embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical scheme deviate from the scope of the technical scheme of the embodiments of the present application.
Claims
1. A target perception method based on image enhancement, characterized in that, The method includes: Acquire point cloud data generated by radar, as well as raw image data captured by a camera; Determine whether the original image data was captured by the camera in a backlit and foggy environment; When the original image data is captured by the camera in a backlit and foggy environment, the original image data is subjected to image brightness enhancement processing to obtain a backlit restored image, and the original image data is subjected to image transmittance optimization processing to obtain a defogging restored image. The backlit restored image, the dehazed restored image, and the point cloud data are fused to obtain fused data. Based on the fused data, the perceived target is identified, and the type and motion state of the perceived target are obtained.
2. The target perception method based on image enhancement according to claim 1, characterized in that, The original image data is in RGB format; The step of enhancing the image brightness value of the original image data to obtain the backlight restored image includes: The original image data is converted from RGB format to HSV format, and the original luminance component is extracted from the original image data in HSV format. According to the preset block size, the original brightness component is divided into multiple image blocks by non-overlapping block processing. A portion of the multiple image blocks is subjected to image block brightness value enhancement processing, and the multiple image blocks before enhancement and the multiple image blocks after enhancement are subjected to image block merging processing to obtain updated brightness components; The updated luminance component is subjected to inverse color space conversion to obtain updated image data in RGB format; The updated image data is subjected to contrast expansion processing using a preset contrast expansion function to obtain the backlight restoration image.
3. The target perception method based on image enhancement according to claim 2, characterized in that, The step of enhancing the brightness value of a portion of the plurality of image blocks includes: From the original brightness components, determine the image subject region where the perceived target is located, and from the plurality of image blocks, determine a plurality of candidate image blocks that are completely located within the image subject region; Based on the brightness value of each candidate image block, a median brightness value is obtained, and based on the median brightness value and a preset adjustment parameter, a brightness threshold range is obtained; wherein, the median brightness value is within the brightness threshold range; From the candidate image blocks, a number of target image blocks whose brightness values are within the brightness threshold range are determined, and each target image block is subjected to image block brightness enhancement processing.
4. The target perception method based on image enhancement according to claim 2 or 3, characterized in that, Multiple block sizes are preset; The first block scale is any one of the multiple block scales. For the first block scale, the step of performing contrast expansion processing on the updated image data using a preset contrast expansion function to obtain the backlight restoration image includes: By using a preset contrast expansion function, the updated image data corresponding to the first block scale is subjected to contrast expansion processing to obtain the candidate backlight restoration image corresponding to the first block scale. When the first block scale is any one of the multiple block scales except the last one, the candidate backlight restoration image corresponding to the first block scale is determined as the original image data corresponding to the second block scale; wherein, the second block scale is located next to the first block scale. When the first block scale is the last of the multiple block scales, the candidate backlight restoration image with the smallest distortion is determined as the backlight restoration image from the candidate backlight restoration images corresponding to each block scale.
5. The target perception method based on image enhancement according to claim 1, characterized in that, The original image data is in RGB format; The step of performing image transmittance optimization processing on the original image data to obtain a dehazed and restored image includes: Convert raw image data in RGB format to grayscale image; A gradient weight map is obtained based on the gradient magnitude of the grayscale image, and a brightness weight map is obtained based on the grayscale values of the grayscale image. The gradient weight map and the brightness weight map are subjected to dynamic weight fusion processing to obtain a dynamic weight map; From the original image data, atmospheric light values are determined; wherein, the atmospheric light values refer to the RGB channel vectors representing atmospheric light in the original image data; Using a pre-set dark channel prior algorithm, the original transmittance map of the original image data is obtained based on the atmospheric light value. The original transmittance map is then optimized using the dynamic weight map to obtain an updated transmittance map. Based on the updated transmittance map, the original image data is dehazed to obtain the dehazed restored image.
6. The target perception method based on image enhancement according to claim 5, characterized in that, Determining atmospheric light values from the original image data includes: From the original image data, determine multiple candidate pixels whose brightness and saturation both meet preset conditions; From the plurality of candidate pixels, determine the target pixel with the largest RGB channel vector; The atmospheric light value is obtained based on the RGB channel vector of the target pixel.
7. The target perception method based on image enhancement according to claim 1, characterized in that, The process of fusing the backlight-restored image, the dehazed image, and the point cloud data to obtain fused data includes: Backlight enhancement features are extracted from the backlight restored image, dehazing enhancement features are extracted from the dehazing restored image, and geometric features are extracted from the point cloud data; wherein, the backlight enhancement features include color distribution features, texture features, and brightness features, and the dehazing enhancement features include contrast features, detail features, and color saturation features; The backlight enhancement feature and the dehazing enhancement feature are fused to obtain image features, and the image features and the geometric features are fused to obtain the fused data.
8. A target perception device based on image enhancement, characterized in that, The device includes: The data acquisition module is used to acquire point cloud data generated by radar and raw image data captured by camera; An environment judgment module is used to determine whether the original image data was captured by the camera in a backlit and foggy environment; The image processing module is used to perform image brightness enhancement processing on the original image data when the original image data is captured by the camera in a backlit and foggy environment to obtain a backlit restored image, and to perform image transmittance optimization processing on the original image data to obtain a defogging restored image. The feature fusion module is used to fuse the backlight restoration image, the dehazing restoration image, and the point cloud data to obtain fused data, and to identify the sensing target based on the fused data to obtain the type and motion state of the sensing target.
9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; When the processor executes the computer execution instructions stored in the memory, it is used to implement the image enhancement-based target perception method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the image-enhanced target perception method as described in any one of claims 1 to 7.