Target Image Recognition and Target Detection Method Based on Video Enhancement Algorithm

By generating effectively enhanced images using a visual sensor and combining an improved YOLO model with dual recognition verification, the problems of noise residue and false positives in target detection under adverse weather conditions are solved, achieving clear recognition of target contours and textures and reducing the false positive rate.

CN121392260BActive Publication Date: 2026-04-03BEIJING LISIDA NEW TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing target detection technologies suffer from lagging filtering parameters in adverse weather conditions such as rain, snow, and fog, resulting in residual noise, obscuring of target details, and a high false positive rate. Furthermore, the lack of feature quality assessment in the detection model leads to coordinate calculation errors.

Method used

The system receives video stream signals through a visual sensor, generates effectively enhanced images, and inputs them into an improved YOLO model for feature extraction and category recognition. Combined with bidirectional calibration and dual recognition verification, it ensures image quality and the accuracy of detection results.

Benefits of technology

It effectively reduces residual noise interference in dense rain scenes, ensures clear target outlines and textures, reduces artifacts and noise effects, lowers the detection misjudgment rate, and provides accurate target information support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121392260B_ABST
    Figure CN121392260B_ABST
Patent Text Reader

Abstract

This invention discloses a target image recognition and target detection method based on video enhancement algorithms, belonging to the field of image processing technology. The method first receives a video stream of a rain, snow, or fog scene using a visual sensor, extracts target images from the video frames, and performs video enhancement processing to eliminate rain and snow occlusion, fog blurring, and noise, generating an effective enhanced image with complete target contours, clear details, and suitability for subsequent detection. Next, this image is input into an improved YOLO model for rain, snow, and fog scenes. An optimized feature extraction network completes feature extraction and category recognition, outputting preliminary target bounding boxes and category labels to determine the presence of a target and its specific category. Finally, if a target is found, the detection data is statistically analyzed and validity is verified. This method achieves high accuracy, low false positives, and strong real-time performance for target detection under adverse weather conditions such as rain, snow, and fog, effectively solving the problem of high false positive rates in existing technologies due to parameter adjustment lag in adaptive filtering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method for target image recognition and target detection based on video enhancement algorithms. Background Technology

[0002] With the surge in demand for real-time target detection in fields such as intelligent driving, security monitoring, and drone inspection, robustness in complex scenarios has become a core challenge. In adverse weather conditions such as rain, snow, and fog, video images are prone to problems such as a sharp drop in contrast, blurred details (fog), rain obscuring (rain), and snow noise (snow), leading to a significant increase in false positives and false negatives for traditional target detection algorithms. Therefore, video enhancement algorithms for adverse weather conditions have become crucial. Among them, the improved YOLO (You Only Look Once) model for rain, snow, and fog scenarios is the most representative. By integrating the enhancement module with the detection network, it improves image quality from the source, laying the foundation for accurate target identification.

[0003] Existing technologies first employ video enhancement techniques such as the Retinex (Retina-Cortex Algorithm) to restore image contrast, eliminate rain and snow noise and fog blur, and restore target detail information. Then, the enhanced image is input into improved YOLO models, primarily YOLOv5 / YOLOv7 / YOLOv8. For example, a differentiable image processing module is embedded in the YOLOv5 backbone to dynamically predict image enhancement parameters in rain, snow, and fog scenes; or multiple frames of optical flow features are fused in the YOLOv7 neck layer to aggregate spatiotemporal information and enhance dynamic target features; or attention-guided feature extraction branches are designed for YOLOv8 to focus on key textures and edges of the target. Finally, the enhancement module and the improved YOLO detection network are jointly trained end-to-end, allowing them to work together to effectively reduce interference from severe weather.

[0004] For example, Chinese invention patent application CN119338857A discloses a target tracking control method, device, and camera based on image recognition, which includes: performing color recognition on a target image using an adaptive color recognition algorithm, identifying color regions and calculating the coordinates of the color center; recognizing the target image using a pre-trained YOLOv2 target detection model, calculating the coordinates of the object center based on the identified object region; calculating the tracking center coordinates based on the color center coordinates and the object center coordinates, converting the tracking center coordinates into motion control signals to drive the rotation of a vision sensor, calculating the actual rotation angle of the vision sensor using a PID control algorithm, and correcting the rotated vision sensor based on the calculation results to achieve target tracking.

[0005] The above-mentioned technology has at least the following technical problems:

[0006] Existing image filtering processes often rely on preset rules or fixed mapping tables. When rain line density increases sharply, a complete parameter optimization process needs to be re-executed. If the optimization algorithm is not optimized for convergence speed in real-time scenarios, parameter adjustments will lag behind the rhythm of rain line changes. This leads to the risk of residual noise in existing adaptive filtering methods in dense rain line scenes, which can easily interfere with feature extraction and obscure key details of the target (such as vehicle outlines and pedestrian limbs). Furthermore, existing detection models do not incorporate feature quality assessment and lack filtering for features affected by artifacts and noise. These features are then directly used for tracking, failing to avoid coordinate calculation biases caused by low-quality features, resulting in a high misclassification rate. Summary of the Invention

[0007] To address the technical problems in the prior art, embodiments of the present invention provide a method for target image recognition and target detection based on video enhancement algorithms. The technical solution is as follows:

[0008] Step 1: After the vision sensor receives the video stream signal from the specified scene, it converts the video stream signal into a digital image sequence to generate continuous video frames. Simultaneously, it extracts the target image from the video frames and performs target region detection to obtain an effective enhanced image. Step 2: The obtained effective enhanced image is input into the improved YOLO model for image feature extraction and category recognition. It outputs preliminary target bounding boxes and category labels to determine whether the target to be detected and its specific category exist in the effective enhanced image. Step 3: If the target to be detected and its specific category exist, the effectiveness is verified to ensure the accuracy and reliability of the detection results. Otherwise, reverse correction and verification are performed to reduce the distortion of target features in rain, snow, and fog scenes.

[0009] The beneficial effects of the technical solutions provided by the embodiments of the present invention include at least the following:

[0010] 1. This invention receives video streams of rain, snow, and fog scenes using a visual sensor, extracts target images from video frames, and performs target region detection. After video enhancement processing, an effectively enhanced image is generated, ensuring complete target contours, clear textures, and brightness and contrast suitable for subsequent detection. This effectively reduces the lag in feature extraction caused by residual noise in dense rain scenes, and the recognition of key details such as vehicle outlines and pedestrian limbs that are obscured. The enhanced image is input into an improved YOLO model for rain, snow, and fog scenes. An optimized feature extraction network completes feature extraction and category recognition, outputting preliminary target bounding boxes and category labels to determine the target's existence and specific category. The improved model enhances feature extraction capabilities for complex scenes, and combined with high-quality image data, reduces the impact of artifacts and noise on features, avoiding recognition bias caused by the lack of feature quality assessment in existing models. If a target to be detected exists, the detection data is statistically analyzed and validity is verified. This step implicitly filters low-quality features affected by artifacts and noise, avoiding coordinate calculation biases caused by them, ensuring accurate and reliable detection results, and effectively reducing the high misjudgment rate in existing technologies. Ultimately, this provides accurate target information support for subsequent decision-making such as obstacle avoidance in intelligent driving and early warning in security monitoring.

[0011] 2. By using the preset acquisition resolution and frame rate of the visual sensor as a benchmark, video frame data containing pixel grayscale matching values ​​and distortion adaptation matching values ​​of the current scene are collected as bidirectional calibration input samples to ensure that the calibration data is compatible with the sensor hardware parameters and avoid judgment bias caused by inconsistent data acquisition standards. Next, the system's built-in standard parameter table of video frames containing pixel grayscale value distribution ranges and image distortion degree reference values ​​is extracted to judge the calibration samples and select video frames that meet the grayscale adaptation and distortion standards. This process reduces the risk of residual noise interfering with feature extraction in dense rain scenes from the source, preventing key details such as vehicle outlines and pedestrian limbs from being obscured. Simultaneously, it replaces the feature quality pre-verification missing in the existing model with dual evaluation, avoiding coordinate calculation bias caused by low-quality features. During judgment, if both values ​​meet the standards, the calibration is qualified and enters target detection; if neither meets the standards, the Gamma correction coefficient and the parameters of the special distortion removal algorithm are dynamically adjusted step by step according to a preset multiple, without relying on a fixed mapping table, avoiding parameter optimization lag when rain line density increases sharply; if only a single parameter fails to meet the standards, it is adjusted specifically and re-judged. At the same time, a maximum number of corrections is set, and if the target is not met, an early warning is issued and detection is suspended to further ensure the image quality of the input target area and reduce the false judgment rate of subsequent detections from the source.

[0012] 3. By acquiring bidirectionally calibrated target images, if edge contour gaps exist, feature enhancement and filling are performed first. Then, an adaptive threshold segmentation algorithm is used to divide the target and background regions, initially locating the target range and recording the coordinate area. If there are no gaps, the target is located directly, avoiding the loss of target features due to residual noise from dense rain lines and solving the detail masking problem caused by the lag in the adjustment of existing filter parameters. The gray-level co-occurrence matrix contrast is calculated. If it is not less than the reference value, it is recorded as a valid region, and information is extracted to generate a valid enhanced image. Otherwise, it is recorded as an invalid region and returned for relocalization verification until a valid region is selected or the maximum number of retries is reached. If the standard is not met, an alert is issued and the process is paused. Among them, gray-level contrast verification can replace the feature quality assessment missing in the existing model, filter low-quality features affected by artifacts and noise, reduce the subsequent detection misjudgment rate from the data source, and ensure the image quality input to the improved YOLO model.

[0013] 4. By statistically analyzing the total number of dimensions of candidate target features output by the improved YOLO model, the overlap ratio between these dimensions and the training reference target features is calculated to obtain the feature matching degree. This metric accurately quantifies the consistency between candidate features and standard features, effectively identifying feature deviations caused by residual noise and artifacts, and preventing interfered features from entering subsequent processes. The bounding box confidence score output by the model is obtained, reflecting the probability of the bounding box locating the true target. The feature matching degree and the bounding box confidence score are then multiplied, and a two-dimensional interval is constructed using Cartesian combination. The product result is then double-integrated to obtain a dual recognition verification value that quantifies the comprehensive reliability of the candidate target. This process, through dual verification of feature matching degree and bounding box confidence score, constructs a comprehensive feature quality evaluation mechanism, making up for the lack of feature quality evaluation in existing detection models. It can filter out low-matching features caused by residual noise from dense rain lines and avoid coordinate calculation deviations caused by inaccurate bounding box positioning, thus providing dual protection for the reliability of candidate targets from both feature and positioning perspectives.

[0014] 5. Initial bounding boxes and category labels are output based on the morphology of the target region: For rectangular regions, only those regions that meet the dual recognition verification criteria and whose bounding box coordinates are within the image range are output with edge-fitting coordinates and corresponding labels to ensure accurate positioning. For irregular regions, the coordinates of the minimum bounding rectangle, the original contour set, and labels are output, balancing standardization and morphological details. If the verification fails or the coordinates are out of range, the target is deemed invalid and not output, filtering low-quality data from the results. Next, the presence and specific category of the target in the image are determined: the bounding box and label information from the structured output are obtained; if the target exists, its category is extracted; otherwise, it is not. Simultaneously, the number of targets by category is counted to form a conclusion. Throughout this process, dual recognition verification avoids coordinate bias caused by the lack of quality assessment in existing models; classification output and statistics further ensure the accuracy of the results, providing high-quality target information for subsequent tracking and decision-making, solving the core problems of lagging filtering parameters and the lack of quality assessment in existing detection models.

[0015] 6. By acquiring the bounding box coordinates, category labels, and dual-recognition verification values ​​of valid targets, it is determined whether they are within the reliability range, providing a basis for feature correction. Next, three types of correction factors—texture, edge, and grayscale—are mapped and output to enhance texture levels, fill in contour gaps, and adjust grayscale range, respectively, specifically addressing the issue of target detail obscuring caused by rain, snow, and fog. Finally, the algorithm parameters are adjusted using the correction factors to output the corrected features, and the feature deviation value and dual-recognition verification value are recalculated. If the verification value meets the standard, the result is retained; if it does not meet the standard but the deviation decreases by more than a preset amount, the factor is optimized and the test is repeated; otherwise, it is marked. The entire process effectively avoids coordinate calculation errors caused by low-quality features, significantly reducing the detection misjudgment rate. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart of a target image recognition and target detection method based on video enhancement algorithm provided in an embodiment of the present invention;

[0018] Figure 2 A flowchart of target area detection provided in an embodiment of the present invention;

[0019] Figure 3 A flowchart illustrating the output of the preliminary target bounding box and category labels provided in this embodiment of the invention;

[0020] Figure 4 This is a schematic diagram of the output results of the preliminary target bounding box and category labels provided in an embodiment of the present invention;

[0021] Figure 5 A flowchart illustrating the reverse correction and verification process provided in this embodiment of the invention. Detailed Implementation

[0022] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0023] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0024] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0025] In response to the problems of existing target detection technologies in adverse weather conditions such as rain, snow, and fog, which are prone to noise residue, target detail obscuration, and high false positive rates due to lag in filtering parameters and lack of feature quality assessment, this invention provides a target image recognition and detection method based on video enhancement algorithms, which can be efficiently adapted to core application scenarios such as intelligent driving and security monitoring.

[0026] like Figure 1 The flowchart shown is for a target image recognition and target detection method based on video enhancement algorithms. The processing flow of this method may include:

[0027] Step 1: When the vision sensor receives the video stream signal under the specified scene (such as rain, snow, fog and severe weather), it converts the video stream signal into a digital image sequence to generate continuous video frames. Simultaneously, it extracts the target image in the video frame and performs target region detection to obtain an effective enhanced image. The effective enhanced image represents an optimized image that has been eliminated by rain and snow occlusion, fog blurring and noise interference after video enhancement processing, and whose contrast and brightness are adapted to subsequent target detection.

[0028] Step 2: Input the acquired enhanced image into the improved YOLO model for rain, snow and fog scenes. The feature extraction network in the improved YOLO model performs image feature extraction and category recognition, and outputs preliminary target bounding boxes and category labels to determine whether there is a target to be detected in the enhanced image and the specific category of the target. The improved YOLO model is used to optimize feature extraction capabilities for complex scenes and improve the positioning accuracy and category discrimination of targets in low-quality environments.

[0029] Step 3: If there is a target to be detected and its specific category, then the validity is verified to ensure the accuracy and reliability of the detection results, providing accurate target information support for subsequent decisions (such as obstacle avoidance in intelligent driving and early warning in security monitoring). Otherwise, reverse correction and verification are performed to reduce the distortion of target features in rain, snow and fog scenarios.

[0030] Specifically, visual sensors can cover vehicle cameras (intelligent driving scenarios), security monitoring cameras (park / road monitoring scenarios), industrial cameras (special environment detection scenarios), etc. They need to have lens coatings that resist rain and snow interference and imaging parameters that are adapted to harsh weather (such as wide dynamic range and high sensitivity) to ensure stable acquisition of video streams in low light and high humidity environments.

[0031] The image signal processor (ISP) built into the vision sensor first performs noise reduction and white balance correction on the analog video stream signal, and then uses an analog-to-digital converter (ADC) to quantize the analog signal into digital pixel data. Subsequently, the digital signal is segmented into frames according to a preset frame rate (such as 25fps), with each frame corresponding to a digital image, and finally forming a continuous digital image sequence, providing a standardized data format for subsequent processing.

[0032] Based on the inter-frame difference method, the pixel changes of consecutive video frames are compared to locate potential target areas that are moving or static. Then, by combining grayscale thresholding (excluding overly bright / dark backgrounds) and edge detection (such as the Canny algorithm), image blocks containing the core features of the target (such as vehicle outlines and pedestrian shapes) are cropped from the region, thus completing the target image extraction.

[0033] Based on the YOLOv8 architecture, attention mechanisms (such as CBAM, Convolutional Block Attention Module) are added to the feature extraction network (Backbone) to strengthen the weight of target features in rain, snow and fog scenarios; a multi-scale feature fusion module is added to the neck network to improve the recognition ability of small targets (such as pedestrians in the distance); at the same time, the model parameters are fine-tuned and the loss function is optimized (such as introducing CIoU loss for bounding box regression) using a rain, snow and fog scenario labeled dataset (containing rain, snow and fog samples of different densities), ultimately achieving improved localization accuracy and class discrimination in complex scenarios.

[0034] In a specific embodiment, namely intelligent transportation scenario (road monitoring in rain and snow), after receiving road video stream signals under rain and snow conditions, the visual sensor (high-definition road camera) first converts them into digital image sequences to generate video frames, simultaneously extracting target images containing vehicles and pedestrians. After video enhancement processing to eliminate rain line obstruction and snow reflection noise, an effectively enhanced image with appropriate brightness is output. This image is then input into an improved YOLO model, which accurately locates low-speed vehicles and pedestrians crossing the road on icy surfaces through multi-scale feature fusion, outputting bounding boxes and category labels. If targets such as vehicles illegally changing lanes or pedestrians running red lights are detected, the validity verification process confirms the reliability of the results and pushes them to the traffic control platform in real time, providing data support for dynamic traffic light control and on-site traffic management by traffic police. If no valid targets are detected, the system reverse-corrects for feature distortion caused by fog, ensuring uninterrupted road monitoring under severe weather conditions.

[0035] In another specific embodiment, namely the intelligent security scenario (park perimeter protection in rain, snow, and fog), infrared visual sensors at the park perimeter receive monitoring video streams under rain, snow, and fog conditions. After converting them into digital image sequences, target images containing fences, people, and vehicles are extracted. Video enhancement processing is used to eliminate fog blur and noise caused by raindrops, generating effectively enhanced images with appropriate contrast. An improved YOLO model, combined with an attention mechanism, focuses on enhancing the limb features of people climbing over fences and the outline features of vehicles illegally entering, accurately outputting target bounding boxes and category labels such as personnel intrusion and vehicle violations. If a valid intrusion target is detected, the park security alarm is immediately triggered after the validity verification is passed, and the target location is simultaneously pushed to the security terminal. If no valid target is detected, the feature weakening caused by fog is corrected in reverse to avoid security omissions due to weather interference, ensuring reliable 24-hour perimeter protection for the park.

[0036] Preferably, before target area detection, a bidirectional video frame calibration judgment is performed. The specific judgment process is as follows: Based on the preset acquisition resolution and frame rate of the visual sensor, video frame data extracted in the current scene is acquired and used as input samples for bidirectional calibration data. The video frame data includes pixel grayscale matching values ​​and distortion adaptation matching values. The bidirectional calibration data input samples are judged based on the video frame standard parameter table to select video frame data that meets the expected requirements for subsequent detection. The video frame standard parameter table includes the pixel grayscale value distribution range and image distortion degree reference values. The pixel grayscale matching value represents the result of calculating the actual grayscale value distribution of the current video frame and the pixel grayscale value distribution under the corresponding weather type in the video frame standard parameter table. The overlap ratio of the grayscale range reflects the degree of matching between the current video frame's grayscale distribution and the standard grayscale range. The higher the value, the better the grayscale distribution meets the requirements of subsequent target detection for image brightness and contrast. The distortion fit matching value represents the deviation calculation result based on the proportion of rain and snow occlusion area in the current video frame and the image distortion reference value in the standard parameter table of the video frame. That is, the ratio of the absolute value of the difference between the proportion of rain and snow occlusion area and the image distortion reference value to the image distortion reference value. It is used to reflect the degree of distortion of the current video frame caused by severe weather such as rain, snow, and fog and its fit with the reference distortion allowable range. The higher the value, the lower the degree of image distortion, which is more conducive to the accurate identification of target contours and texture details in subsequent target area detection.

[0037] Among them, the discrimination of the bidirectional calibration data input sample is based on the video frame standard parameter table. The specific steps are as follows: If the pixel gray level matching value is within the pixel gray level value distribution range and the distortion adaptation matching value is not greater than the image distortion degree reference value, the video frame bidirectional calibration determination result is recorded as qualified for video frame bidirectional calibration, and the target area detection can be directly carried out; If the pixel gray level matching value is not within the pixel gray level value distribution range and the distortion adaptation matching value is greater than the image distortion degree reference value, a step-by-step correction strategy is adopted to synchronously optimize the video frame gray level parameter (i.e., Gamma correction coefficient) and the video frame distortion parameter (i.e., structural element size); If only the pixel gray level matching value is not within the pixel gray level value distribution range, only the video frame gray level parameter is adjusted, and after adjustment, a new bidirectional calibration data input sample is generated and determined; If only the distortion adaptation matching value is greater than the image distortion degree reference value, only the video frame distortion parameter is adjusted, and after adjustment, a new bidirectional calibration data input sample is generated and determined; If it is not satisfied within the preset maximum number of corrections (such as 3 times), a bidirectional calibration warning is issued to prompt manual intervention to check the status of the vision sensor (such as lens contamination, abnormal exposure parameters), and at the same time, the subsequent target area detection process is paused to avoid misjudgment caused by low-quality images entering the detection link.

[0038] The step-by-step correction strategy is adopted as follows: The video frame gray level parameter is adjusted based on the pixel gray level matching value to improve the pixel gray level matching value. Specifically, the Gamma correction coefficient is increased by a preset multiple of the pixel gray level matching value (usually 1.2 times) to enhance the overall brightness of the image; The video frame distortion parameter is adjusted based on the distortion adaptation matching value to improve the distortion adaptation matching value. Specifically, the structural element size of the adaptive morphological filtering is increased by a preset multiple of the distortion adaptation matching value (usually 1.1 times) to strengthen the separation of rain lines / snowflakes; After the video frame gray level parameter and the video frame distortion parameter are adjusted, the video frame data is re-acquired to generate a new bidirectional calibration data input sample, and the bidirectional calibration determination is carried out again. If the video frame bidirectional calibration determination result is still unqualified, a video frame parameter abnormality warning is issued to prompt the preset personnel for manual intervention, such as hardware inspection of the vision sensor.

[0039] Specifically, the standard parameter table for video frames is constructed by scene classification, data collection, and statistical optimization: First, scenes such as rain, snow, and fog are divided, and 5000+ high-quality video frames are collected in each scene. Data such as grayscale distribution and occlusion ratio are statistically analyzed. Combined with the detection accuracy for reverse verification, an initial table is formed and dynamically iteratively optimized. The pixel grayscale value distribution range is determined by statistically analyzing the grayscale histograms of high-quality frames in similar scenes, selecting the interval with a frequency ≥ 85%. This is to filter out interference from extremely bright / dark pixels. Taking a moderate snow scene as an example, this frequency range covers 98% of vehicle / pedestrian target pixels, eliminating <15% of extreme bright / dark noise points, ensuring grayscale adaptation for detection requirements. The image distortion reference value is calculated by adding a 20% error tolerance to the average proportion of rain and snow occlusion in high-quality frames in similar scenes. This is to balance distortion control with scene adaptability. For example, in a light rain scene, the average is 3%, and after adding, it becomes 3.6%, which can accommodate sudden increases in raindrops (such as short-term showers) and avoid misjudging slight occlusions. The allowable range of reference distortion is ±10% of the historical distortion reference value. This is to adapt to dynamic scene changes. If the reference value is 3.6%, a fluctuation of 3.24%-3.96% is allowed to adapt to instantaneous raindrop residue in the lens and reduce unnecessary parameter corrections.

[0040] The maximum number of corrections is preset to 3, as 1-2 corrections can resolve 90% of the deviation frames. If the standard is not met after 3 corrections, it is mostly a hardware problem, and further corrections will affect real-time performance. The Gamma correction coefficient is a parameter built into the visual sensor ISP chip (default 1.0, range 0.4-2.0). The initial value is set by the manufacturer and is dynamically adjusted according to the grayscale matching value. The preset multiple for pixel grayscale matching value is 1.2 times and the multiple for distortion adaptation matching value is 1.1 times. Both have been determined through comparative experiments: at this multiple, 80% of the frames can meet the standard in one attempt, and the adjustment range is too large, which will cause new deviations.

[0041] The aforementioned grayscale frequency of 85%, error tolerance of 20%, allowable fluctuation of ±10%, number of corrections of 3, and parameter adjustment multiples (1.2x and 1.1x) are all optimal experimental results verified through laboratory scene simulation and field testing. In practical applications, these parameters can be slightly adjusted according to differences in visual sensor models, target scene weather intensity (such as blizzard / dense fog), and detection accuracy requirements to adapt to specific application environments.

[0042] In this embodiment, the process improves target detection quality from the source, adapting to harsh scenarios such as rain, snow, and fog. Firstly, bidirectional calibration pre-screens qualified video frames, judging based on both grayscale and distortion dimensions, preventing low-quality images from entering the detection stage and significantly reducing the false detection rate caused by brightness imbalances and rain / snow occlusion. Secondly, the strategy combining step-by-step correction and single-parameter adjustment can specifically address grayscale deviation or distortion issues: dynamic adjustment of the Gamma correction coefficient quickly optimizes image brightness, while adaptive morphological filter parameter adjustment efficiently separates rain lines and snowflakes, balancing correction efficiency and effectiveness. Thirdly, a pre-set early warning mechanism and manual intervention prompts can promptly identify sensor hardware faults (such as lens contamination) and simultaneously pause the detection process to prevent invalid calculations, providing a high-quality data foundation for subsequent target area detection and YOLO model recognition, especially suitable for scenarios with high real-time and accuracy requirements such as intelligent transportation and security monitoring.

[0043] like Figure 2 The target region detection flowchart shown below has the following design logic: First, by judging the relationship between the acquired Euclidean distance and the corresponding set value, it is determined whether to perform target feature enhancement and region segmentation to initially locate the target range. Next, the validity of the target region is verified by comparing the gray-level co-occurrence matrix contrast with the reference value: if the conditions are met, it is marked as valid and the enhanced image is output, completing the detection; if not, it is recorded as an invalid region. If the maximum number of retries has not been reached, the target region is re-located and verified again; if it is still invalid, an alert is triggered. This logic ensures the accuracy and reliability of target detection through a multi-level verification and retry mechanism, and is suitable for image analysis in complex scenes.

[0044] It is necessary to further understand that the specific process of target region detection is as follows:

[0045] The first step is to acquire the target image after bidirectional calibration of the video frames. If the Euclidean distance between adjacent edge pixels in the target image is greater than the set Euclidean distance (usually set to 3 pixels), it indicates that there is an edge contour gap in the target image. Target feature enhancement processing is then performed to fill in the edge contour features of the target in the target image. Based on an adaptive threshold segmentation algorithm, the target image after target feature enhancement processing is divided into a target region and a background region to initially locate the target region range. The corresponding pixel coordinates and area information are recorded. The target region represents the set of pixels containing the target to be detected and whose features are valid. If the Euclidean distance between adjacent edge pixels in the target image is not greater than the set Euclidean distance, it indicates that there is no edge contour gap in the target image, and the target region range is directly and initially located.

[0046] The second step involves feature verification of the initially located target area to identify effective target areas. Specifically, this involves obtaining the gray-level co-occurrence matrix contrast of the target area to quantify the degree of difference in pixel gray levels within the target area. The higher the contrast value, the more obvious the pixel gray-level changes and the more prominent the texture details, which better meet the requirements of an effective enhanced image for target texture. If the gray-level co-occurrence matrix contrast is not less than the corresponding reference contrast, the corresponding target area is recorded as an effective target area, and the pixel coordinates, area, and gray-level distribution information of the area are extracted to obtain an effective enhanced image, i.e., a high-quality enhanced image. Otherwise, the corresponding target area is recorded as an invalid target area, and the process returns to the first step to re-perform the initial target area location and feature verification until all effective target areas are selected or the preset maximum number of retries (e.g., 3 times) is reached, ensuring that the final effective target areas can support the generation of an effective enhanced image. If the corresponding target area is still recorded as an invalid target area within the preset maximum number of retries, a target area identification warning is issued to prompt manual inspection of the visual sensor acquisition parameters (e.g., focal length, exposure time), and the effective enhanced image generation process is paused to prevent invalid data from entering subsequent detection stages.

[0047] Specifically, the Euclidean distance is set to 3 pixels because discontinuities smaller than this threshold are usually caused by slight noise or gradient changes and can be considered as continuous edges; discontinuities larger than this value indicate significant gaps in the contour, requiring feature enhancement to complete and ensure the continuity of the target boundary. The maximum number of retries is preset to 3 because 1-2 retries can resolve most positioning errors. If the target is still not met after 3 retries, it is mostly due to sensor parameters or scene factors. Continuing to retries will increase latency and reduce real-time performance. The reference contrast is obtained by calculating the gray-level co-occurrence matrix contrast of the target area of ​​high-quality samples in the same scene, taking the mean or the upper limit of the 95% confidence interval, to ensure matching with the texture sensitivity of the detection model. Pixel coordinates, area, and gray-level distribution information are extracted for precise cropping of the target area, calculation of region features (such as contrast), and enhancement processing such as gray-level stretching, thereby generating a high-quality enhanced image that focuses on the target and removes background noise.

[0048] In this embodiment, edge continuity is quantized using Euclidean distance of 3 pixels, and the gaps are filled by feature enhancement, which improves the integrity of the target outline in rain, snow and fog scenes and provides a reliable boundary for region segmentation. Existing technologies directly use preliminary region segmentation, which is prone to introducing interference. This method uses gray-level co-occurrence matrix contrast verification + 3 retry mechanism to filter effective regions, effectively avoiding invalid regions from entering subsequent stages and improving data quality.

[0049] Preferably, before outputting the initial target bounding box and category label, a dual verification of feature matching degree and bounding box confidence is performed. Specifically, this involves: counting the total number of dimensions corresponding to the candidate target features output by the feature extraction network of the improved YOLO model, i.e., the total number of dimensions contained in the feature vector generated after the feature extraction network encodes the candidate target (e.g., 512 dimensions, 1024 dimensions, preset by the network structure); calculating the overlap ratio between the candidate target features and the reference target features trained by the improved YOLO model, i.e., by comparing the vectors corresponding to the candidate target features and the vectors corresponding to the reference target features dimension by dimension, counting the number of dimensions where the numerical deviation is less than a preset value (e.g., 0.1), and recording this as the number of consistent dimensions; and calculating the result based on the ratio of the number of consistent dimensions to the total number of dimensions. The process involves obtaining the feature matching degree; acquiring the bounding box confidence score, which reflects the probability that the candidate target is the real target, from the output of the improved YOLO model. The bounding box confidence score reflects the improved YOLO model's confidence in the accuracy of the candidate target's bounding box localization; a higher value indicates a greater likelihood that the bounding box contains the real target. The acquired feature matching degree and bounding box confidence score are multiplied, and the actual value corresponding to 0 and the feature matching degree is set as the first integration interval, while the actual value corresponding to 0 and the bounding box confidence score is set as the second integration interval. The first and second integration intervals are combined using a Cartesian combination to obtain a two-dimensional interval, and the result of the multiplication operation is double-integrated within the two-dimensional interval to obtain a double recognition verification value, which serves as a quantitative indicator of the overall reliability of the candidate target.

[0050] Specifically, the process of obtaining the dual recognition verification value is as follows: the feature matching degree and the bounding box confidence degree are used as two variables of a binary function to construct a continuous function on a two-dimensional plane. Its expression is: In the formula, x represents the feature matching degree, y represents the bounding box confidence degree, and the continuous function is calculated in the interval... and Double integrals on That is, first integrate variable a over the interval [0, x] (treating b as a constant) to obtain an integral expression for b, and then integrate the variable b in this expression over the interval [0, y]. da and db are the differentials of the integration variables, representing the small changes in a and b respectively when integrating. By accumulating the small steps of da and db over the regions of a from 0 to x and b from 0 to y, the result of the double integral is obtained. This value serves as a dual recognition verification value, used to measure the comprehensive reliability of candidate target feature matching and bounding box localization. The larger the value, the better the synergistic performance of candidate target feature consistency and localization reliability, and the more it meets the accuracy requirements of target detection in rain, snow and fog scenarios.

[0051] In this embodiment, compared to existing technologies that rely solely on confidence level verification and are susceptible to interference from rain, snow, and fog, this mechanism combines feature matching degree and bounding box confidence. By using double integral quantification to assess the collaborative reliability of both, it can filter out falsely detected targets with high confidence but low feature matching (such as trees in fog being mistaken for pedestrians), effectively improving target recognition accuracy in rain, snow, and fog scenarios. In scenarios where rain, snow, and fog cause target features to become blurred and bounding boxes to easily shift, double integral can highlight targets with both superior features and localization, suppressing candidate targets that meet the criteria in one dimension but are unreliable overall.

[0052] like Figure 3 The flowchart for outputting the initial target bounding box and category label is shown below. Its design logic is as follows: First, determine the target type and apply condition a and condition b accordingly. Then, comprehensively assess the satisfaction of conditions a and b: if both are satisfied, the target is marked as valid and its bounding box and category label are output; if either condition is not satisfied, the target is marked as invalid and no related information is output. Finally, the valid target information is structured for output, ensuring the accuracy and standardization of the detection results, suitable for diverse target detection scenarios. Condition a indicates that the obtained double-recognition verification value is not less than the set double-recognition verification value, and the bounding box coordinates of the candidate target do not exceed the pixel range of the effective enhanced image. Condition b indicates that the obtained double-recognition verification value is not less than the set double-recognition verification value, and the bounding box coordinates of the candidate target's minimum bounding rectangle do not exceed the pixel range of the effective enhanced image.

[0053] like Figure 4 The diagram shown illustrates the initial target bounding box and category labels output. Figure 4 The image presents a foggy road scene, showcasing the initial target bounding boxes and category labels. For vehicles on the road, the model outputs yellow rectangular bounding boxes that perfectly match the actual edges of the vehicles. The top-left and bottom-right corner coordinates accurately match the vehicle's pixel boundaries, with no extraneous background, and the category label "car". Although the fog blurs some target details, the bounding boxes accurately surround the vehicles, demonstrating that the model can effectively output bounding boxes and category labels for rectangular targets even in adverse weather conditions.

[0054] What needs further understanding is that the process of outputting the initial target bounding box and category label is as follows:

[0055] First, for a rectangular target region, if the obtained dual recognition verification value is not less than the set dual recognition verification value, and the bounding box coordinates of the candidate target do not exceed the pixel range of the effective enhanced image (i.e., the horizontal x1 and vertical y1 coordinates of the upper left corner of the bounding box are not less than 0, the horizontal x2 coordinate of the lower right corner is not greater than the image width pixel value, and the vertical y2 coordinate is not greater than the image height pixel value), then the corresponding candidate target is determined to be a valid target to be detected, and the corresponding preliminary target bounding box (recorded with the upper left corner coordinates (x1, y1) and the lower right corner coordinates (x2, y2) that fit the actual edge of the rectangular target, the coordinates accurately match the pixel boundary of the target rectangle, and there is no extra background area) and category label (such as vehicles of regular shape, square traffic signs, rectangular equipment, etc., the label is consistent with the target category system during the training of the improved YOLO model), which not only ensures that the bounding box accurately wraps the rectangular target, but also simplifies the subsequent localization calculation and category recognition process through standardized coordinates.

[0056] For irregular target regions, if the obtained dual recognition verification value is not less than the set dual recognition verification value, and the bounding box coordinates of the minimum bounding rectangle of the candidate target do not exceed the pixel range of the effective enhanced image (i.e., the horizontal x3 and vertical y3 coordinates of the upper left corner of the bounding rectangle are not less than 0, the horizontal x4 coordinate of the lower right corner is not greater than the image width pixel value, and the vertical y4 coordinate is not greater than the image height pixel value), then the corresponding candidate target is determined to be a valid target to be detected, and the corresponding preliminary target bounding box (recorded with the upper left corner coordinates (x3, y3) - the lower right corner coordinates (x4, y4) of the minimum bounding rectangle, and the pixel coordinate set of the original irregular contour of the target is attached) and category label (such as a pedestrian in a curved shape, an irregularly stacked obstacle, etc., the label is consistent with the target category system during the training of the improved YOLO model), which satisfies the standardized localization requirements through the bounding rectangle and preserves the target shape details through the original contour coordinates.

[0057] Then, if the obtained dual recognition verification value is less than the set dual recognition verification value, or the bounding box coordinates of the candidate target exceed the pixel range of the effective enhanced image, the corresponding candidate target is determined to be an invalid target to be detected, and the corresponding preliminary target bounding box and category label are not output. After all candidate targets have completed dual recognition verification and bounding box range verification, the preliminary target bounding boxes and category labels of all valid targets to be detected are summarized, classified and sorted according to the target category (e.g., vehicles first, then pedestrians), and the corresponding dual recognition verification value is labeled (usually retaining three decimal places), and finally a structured output result is formed.

[0058] Finally, from the structured output, the preliminary target bounding boxes and category labels of valid targets are obtained, including the precise bounding box coordinates of rectangular targets, the minimum bounding rectangle coordinates and original contour coordinate set of irregular targets, and the corresponding category label content. If the preliminary target bounding box data and category label information contain the preliminary target bounding boxes and category labels of valid targets, it is determined that there are targets to be detected, and the content of each label (such as vehicles, pedestrians, traffic signs, etc.) is extracted to determine the specific category of the target. Otherwise, it is determined that there are no targets to be detected. For the targets to be detected, the number of each category is counted according to the category label (such as 2 vehicles and 1 pedestrian), forming a conclusion on the existence and quantity of target categories, thus completing the judgment on the existence and specific category of targets to be detected.

[0059] In this embodiment, compared to existing technologies that often use a uniform rectangular bounding box to process all targets, easily losing details of irregular targets or causing redundancy in the positioning of rectangular targets, this mechanism achieves differentiated and accurate output by dividing the target regions into rectangular and irregular types, significantly improving the completeness of target representation and positioning accuracy. For rectangular targets (such as vehicles and traffic signs), the bounding box is directly output with coordinates that fit the edge, eliminating redundant background and reducing subsequent computational redundancy; for irregular targets (such as pedestrians and obstacles), while outputting the minimum bounding rectangle to meet standardized positioning, the original contour coordinates are added to preserve shape details.

[0060] Meanwhile, this mechanism combines dual recognition verification values ​​and bounding box range checks to quantitatively analyze the effectiveness of the two types of targets. It ensures the reliability of target features and positioning through verification values, while eliminating invalid targets outside the image by checking coordinate ranges, avoiding erroneous outputs caused by existing technologies relying solely on confidence scores. The final structured classification and ranking results also provide clear data support for subsequent decisions (such as obstacle avoidance priority determination in intelligent driving), thereby effectively improving the accuracy and usability of target output in rain, snow, and fog scenarios.

[0061] Preferably, the specific process for validity verification is as follows: obtain the detection data of the valid target in the structured output results, including bounding box coordinates, category labels, and dual recognition verification values; if the dual recognition verification value of each target is within the preset reliability range (e.g., ≥0.7), the detection result of the target is directly determined to be valid; otherwise, determine whether the current dual recognition verification value is less than the minimum value of the reliability range (i.e., 0.7): if so, it is determined to be an invalid detection result, and invalid data is marked and the reason is recorded; otherwise, feature correction is performed to reduce the deviation between the target features and the historical standard features of the same category, so that the corrected target features are more in line with the distribution law of effective features in rain, snow and fog scenarios, thereby increasing the probability that the dual recognition verification value falls into the reliability range and enhancing the validity of the detection result.

[0062] The specific process of feature correction is as follows: Feature deviation values ​​are input into a feature inverse correction mapping table for mapping, resulting in the output of feature correction factors, including texture correction factors, edge correction factors, and grayscale correction factors. Texture correction factors quantify the enhancement of texture details in the target area; a larger value indicates a higher level of texture clarity to be enhanced. By enhancing the response of high-contrast pixel pairs in the grayscale co-occurrence matrix, the texture hierarchy of the target surface is highlighted, compensating for texture blurring in rain, snow, and fog scenes. Edge correction factors reflect the compensation intensity of the target edge contour; a larger value indicates a more obvious edge contour gap to be filled. By enhancing the edge gradient... The degree operator enhances the detection sensitivity of weak edges, strengthens the continuity and integrity of the target contour, and reduces the occlusion interference of fog or raindrops on edge features. The grayscale correction factor is used to adjust the dynamic range of pixel grayscale values ​​in the target area. The larger the value, the greater the grayscale contrast that needs to be expanded. By stretching dark details and compressing overly bright areas, the corrected target grayscale distribution is made closer to the grayscale range of historical standard features of the same category, reducing the grayscale shift caused by rain, snow, and fog. Based on the obtained feature correction factors, the target features are reverse-corrected and verified to improve the matching degree between the features and historical reference features of the same category, thereby ensuring that the detection results meet the validity requirements.

[0063] like Figure 5 The reverse correction and verification flowchart shown is designed as follows: First, three correction factors are applied for reverse correction, followed by the generation of dual-identification verification values. The subsequent operation is determined by whether the verification value falls within the reliability range: if it is within the range, the correction result is directly retained and included in the valid set; if not, it is further determined whether the reduction in feature bias reaches a preset standard; if it does, the reverse correction process is repeated for further optimization; if it does not, it is marked as suspicious data and manual review is initiated. Finally, all valid detection results are summarized to form a complete detection report. This logic, through multi-level verification and manual intervention mechanisms, effectively improves the accuracy and reliability of target detection.

[0064] The reverse correction and verification process involves: adjusting the contrast parameters of the gray-level co-occurrence matrix using a texture correction factor, optimizing the threshold setting of the Canny operator using an edge correction factor, and dynamically stretching the gray-level range using a gray-level correction factor, ultimately outputting the corrected target features. After reverse correction, the feature deviation between the corrected target features and historical reference features of the same category is recalculated, represented by the Euclidean distance between the vector corresponding to the target features and the vector corresponding to the historical reference features of the same category. Simultaneously, a dual recognition verification value is regenerated based on the corrected target features. If the regenerated dual recognition verification value is within a preset reliability range, the reverse correction is deemed effective. If the reverse correction is effective, retain the corrected target features and corresponding detection results, and include them in the final valid detection set; otherwise, further determine whether the recalculated feature deviation value has decreased by a preset amount (usually set to 30%) or more compared to before the reverse correction: if so, it means that although the reverse correction has not fully met the standard, there is an improvement effect, and the feature correction factor needs to be readjusted (such as increasing the cause sub-value by 20%) and the reverse correction process should be repeated; if not, the reverse correction is deemed invalid, the target detection result is marked as suspicious data, and the manual review process is initiated; after all target features have completed reverse correction and verification, all valid detection results are summarized to ensure the accuracy and reliability of the output data.

[0065] Specifically, based on statistics from over 1000 valid detection samples in rain, snow, and fog scenarios, when the validation value is ≥0.7, the matching accuracy between the target category and the localization reaches over 92%. Below this value, the accuracy drops sharply, hence this is used as the threshold. In extreme scenarios (such as heavy snow / dense fog), it can be fine-tuned according to the actual detection accuracy requirements (e.g., reduced to ≥0.65) to balance the false negative and false positive rates. Based on feature correction experimental data, when the deviation value is reduced by more than 30% after correction, the probability of secondary verification meeting the standard exceeds 80%. Below this range, the effect of subsequent correction is limited. In scenarios with blurred target textures (such as nighttime fog), it can be fine-tuned to 25% to avoid feature distortion due to over-correction.

[0066] Based on 1000+ sets of target feature data under different rain, snow and fog scenarios, the data is divided into intervals (such as 0-0.3, 0.3-0.6, 0.6-1.0), and the optimal values ​​of texture, edge and grayscale correction factors that are manually verified are matched. Finally, a mapping table between the correction requirement coefficient and the three types of correction factors is established, namely the feature reverse correction mapping table.

[0067] If the texture correction factor is 0.7, first extract the gray-level co-occurrence matrix of the target area, then increase the contrast parameter in the matrix (which originally reflects the difference in pixel gray levels) by 70% according to the factor value to enhance the response of high-contrast pixel pairs (such as vehicle body texture). By enhancing the difference in texture details, the texture blur caused by rain and snow occlusion is compensated, and the optimized texture features are output.

[0068] If the edge correction factor is 0.5, first obtain the current low and high thresholds of the Canny operator (e.g., 50, 150), then reduce the low threshold by 50% (to 25) and the high threshold by 30% (to 105) according to the factor value, improve the detection sensitivity of weak edges in fog (e.g., pedestrian contours), fill edge gaps, enhance contour continuity, and output complete edge features.

[0069] If the grayscale correction factor is 0.6, first calculate the original grayscale range of the target area, such as [70, 190]. Stretch the range by 60% according to the factor value, calculate the new range, such as [50, 210]. Map the original grayscale values ​​to the new range through linear transformation, stretch the dark details and compress the overly bright areas, so that the grayscale distribution fits the historical standard features, and output the corrected target features.

[0070] In this embodiment, firstly, high-confidence targets are quickly screened using a preset reliability range to reduce redundant processing. For targets that do not meet the standards, the root cause of the deviation is accurately located to avoid blind correction. Secondly, the feature correction stage uses three types of factors—texture, edge, and grayscale—for targeted optimization to compensate for texture blurring, edge breakage, and grayscale shift caused by rain, snow, and fog, respectively. Compared with the single correction method in existing technologies, this approach can more comprehensively restore target features. Finally, a reverse verification and secondary correction mechanism ensures that the correction effect meets the standards and avoids invalid loops through preset range judgment. If the standards are not met, manual review is initiated, forming a closed loop of automatic verification, precise correction, and manual fallback. The overall process significantly improves the efficiency and accuracy of target detection results in rain, snow, and fog scenarios, thereby effectively realizing full-process control of target detection under adverse weather conditions such as rain, snow, and fog.

[0071] The following points need to be explained:

[0072] (1) The accompanying drawings of the embodiments of the present invention only involve the structures involved in the embodiments of the present invention. Other structures can refer to the general design.

[0073] (2) For clarity, the thickness of layers or regions is enlarged or reduced in the drawings used to describe embodiments of the invention, i.e., these drawings are not drawn to scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element or there may be intermediate elements.

[0074] (3) Where there is no conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.

[0075] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A target image recognition and target detection method based on video enhancement algorithms, characterized in that, Includes the following steps: Step 1: After the vision sensor receives the video stream signal of the specified scene, it converts the video stream signal into a digital image sequence to generate continuous video frames, synchronously extracts the target image in the video frame, and performs target region detection to obtain an effectively enhanced image. Step 2: Input the acquired enhanced image into the improved YOLO model to perform image feature extraction and category recognition, and output preliminary target bounding boxes and category labels to determine whether there is a target to be detected in the enhanced image and the specific category of the target; Step 3: If there is a target to be detected and its specific category, then perform validity verification to ensure the accuracy and reliability of the detection results; otherwise, perform reverse correction and verification to reduce the distortion of target features in rain, snow and fog scenarios. The specific process for outputting the initial target bounding box and category label is as follows: For a rectangular target region, if the obtained dual recognition verification value is not less than the set dual recognition verification value, and the bounding box coordinates of the candidate target do not exceed the pixel range of the effective enhanced image, then the corresponding candidate target is determined to be a valid target to be detected, and the corresponding preliminary target bounding box and category label are output. For irregular target regions, if the obtained dual recognition verification value is not less than the set dual recognition verification value, and the bounding box coordinates of the minimum bounding rectangle of the candidate target do not exceed the pixel range of the effective enhanced image, then the corresponding candidate target is determined to be a valid target to be detected, and the corresponding preliminary target bounding box and category label are output. If the obtained dual recognition verification value is less than the set dual recognition verification value, or the bounding box coordinates of the candidate target exceed the pixel range of the effective enhanced image, the corresponding candidate target is determined to be an invalid target to be detected, and the corresponding preliminary target bounding box and category label are not output. After all candidate targets have completed dual recognition verification and bounding box range verification, the preliminary target bounding boxes and category labels of all valid targets to be detected are summarized, sorted by target category, and labeled with the corresponding dual recognition verification values, finally forming a structured output result.

2. The target image recognition and target detection method based on video enhancement algorithm as described in claim 1, characterized in that, Before performing target region detection, a bidirectional video frame calibration determination is also included. The specific determination process is as follows: Based on the preset acquisition resolution and frame rate of the visual sensor, video frame data extracted in the current scene is acquired and used as input samples for bidirectional calibration data. The video frame data includes pixel grayscale matching values ​​and distortion adaptation matching values. The bidirectional calibration data input samples are judged based on the video frame standard parameter table to select video frame data that meet the expected requirements for subsequent detection. The video frame standard parameter table includes the pixel gray value distribution range and the image distortion degree reference value. The pixel grayscale matching value is used to reflect the degree of matching between the grayscale distribution of the current video frame and the standard grayscale range; The distortion adaptation matching value is used to reflect the degree of distortion of the current video frame caused by severe weather such as rain, snow, and fog, and its adaptation to the reference distortion allowable range.

3. The target image recognition and target detection method based on video enhancement algorithm as described in claim 2, characterized in that, The specific steps for judging the bidirectional calibration data input samples based on the video frame standard parameter table are as follows: If the pixel grayscale matching value is within the pixel grayscale value distribution range and the distortion matching value is not greater than the image distortion degree reference value, then the video frame bidirectional calibration judgment result is recorded as the video frame bidirectional calibration qualified, and the target area detection is performed directly. If the pixel grayscale matching value is not within the pixel grayscale value distribution range, and the distortion matching value is greater than the image distortion reference value, a step-by-step correction strategy is adopted to simultaneously optimize the video frame grayscale parameters and video frame distortion parameters. If only the pixel grayscale matching value is not within the pixel grayscale value distribution range, then only the grayscale parameters of the video frame are adjusted, and after adjustment, the bidirectional calibration data input sample is regenerated and judged. If only the distortion matching value is greater than the image distortion reference value, then only the video frame distortion parameters are adjusted, and after adjustment, bidirectional calibration data input samples are regenerated and judged. If the preset maximum number of corrections is not met, a two-way calibration warning will be issued to prompt manual intervention to check the status of the vision sensor, and the subsequent target area detection process will be suspended. The step-by-step correction strategy is specifically as follows: To improve pixel grayscale matching values, the grayscale parameters of video frames are adjusted based on pixel grayscale matching values. Specifically: The Gamma correction coefficient is increased by a preset multiple of the pixel grayscale matching value to enhance the overall brightness of the image. Adjusting video frame distortion parameters based on distortion fit matching values ​​to improve distortion fit matching values, specifically: Increase the size of the structuring element of the adaptive morphological filter by a preset multiple of the distortion fit matching value to enhance rain line / snowflake separation; After adjusting the grayscale parameters and distortion parameters of the video frames, new bidirectional calibration data input samples are generated by re-acquiring video frame data, and bidirectional calibration is performed again.

4. The target image recognition and target detection method based on video enhancement algorithm as described in claim 3, characterized in that, The specific process for target region detection is as follows: The first step is to acquire the target image after the bidirectional calibration of the video frame. If the Euclidean distance between adjacent edge pixels in the target image is greater than the set Euclidean distance, it indicates that there is an edge contour gap in the target image. Then, target feature enhancement processing is performed to fill the edge contour features of the target in the target image. The target image after target feature enhancement is divided into target region and background region to initially locate the range of target region and record the corresponding pixel coordinates and area information. The target region represents the set of pixels containing the target to be detected and whose features are effective. If the Euclidean distance between adjacent edge pixels in the target image is not greater than the set Euclidean distance, it indicates that there are no edge contour gaps in the target image, and the target area range can be preliminarily located directly. The second step is to perform feature verification on the initially located target area to identify the effective target area, specifically: Obtain the gray-level co-occurrence matrix contrast of the target area corresponding to the target area range, which is used to quantify the degree of difference in pixel gray levels within the target area; If the contrast of the gray-level co-occurrence matrix is ​​not less than the corresponding reference contrast, then the corresponding target region is recorded as the effective target region, and the pixel coordinates, area and gray-level distribution information of the region are extracted to obtain the effectively enhanced image; Otherwise, the corresponding target region is marked as an invalid target region, and the process returns to step one to perform preliminary target region localization and feature verification again until all valid target regions are selected or the preset maximum number of retries is reached, ensuring that the final valid target regions can support the generation of an effectively enhanced image. If the target area is still recorded as an invalid target area within the preset maximum number of retries, a target area recognition warning will be issued to prompt manual inspection of the visual sensor acquisition parameters, and the effective enhanced image generation process will be suspended.

5. The target image recognition and target detection method based on video enhancement algorithm as described in claim 1, characterized in that, The output of the initial target bounding box and category label includes a dual recognition and verification process involving feature matching degree and bounding box confidence, specifically: The total number of dimensions corresponding to the candidate target features output by the feature extraction network of the improved YOLO model is counted. The overlap ratio between the candidate target features and the reference target features trained by the improved YOLO model is calculated and denoted as the number of consistent dimensions. Based on the number of consistent dimensions and the total number of dimensions, the feature matching degree is obtained. Obtain the bounding box confidence score, which reflects the probability that the candidate target is the real target to be detected, from the output of the improved YOLO model. The bounding box confidence score is used to reflect the degree of confidence of the improved YOLO model in the accuracy of the candidate target bounding box localization. The obtained feature matching degree and bounding box confidence are multiplied, and a first integration interval and a second integration interval are set and combined by Cartesian to obtain a two-dimensional interval. The result of the product operation is subjected to double integration within a two-dimensional interval to obtain a double recognition verification value, which serves as a quantitative indicator of the overall reliability of the candidate target.

6. The target image recognition and target detection method based on video enhancement algorithm as described in claim 1, characterized in that, The specific steps for determining whether a target to be detected exists in the effectively enhanced image and its specific category are as follows: Obtain the preliminary target bounding boxes and category labels corresponding to the valid targets to be detected in the structured output results; If there are valid preliminary target bounding boxes and category labels for a target to be detected in the preliminary target bounding box data and category label information, then it is determined that there is a target to be detected, and the content of each label is extracted to determine the specific category of the target. Conversely, if the target is not found, it is determined that there is no target to be detected. For existing targets to be detected, the number of each category is counted according to the category label to form a conclusion on the existence of target categories and quantities.

7. The target image recognition and target detection method based on video enhancement algorithm as described in claim 1, characterized in that, The specific process for performing validity verification is as follows: Obtain the detection data of the valid target in the structured output, including bounding box coordinates, category labels, and dual recognition verification values; If the dual recognition verification value of each target is within the preset reliability range, the detection result of the target is directly determined to be valid; Otherwise, determine whether the current dual-identification verification value is less than the minimum value of the reliability interval: If so, it is determined as an invalid test result, and invalid data is marked and the reason is recorded and prompted; Otherwise, feature modification is performed to reduce the deviation between the target feature and the historical standard features of the same category.

8. The target image recognition and target detection method based on video enhancement algorithm as described in claim 7, characterized in that, The specific process for feature correction is as follows: The feature deviation value is input into the feature inverse correction mapping table for mapping, and the corresponding feature correction factors are output, including texture correction factor, edge correction factor and grayscale correction factor; The texture correction factor is used to quantify the enhancement of texture details in the target area, the edge correction factor is used to reflect the compensation intensity of the target edge contour, and the grayscale correction factor is used to adjust the dynamic range of pixel grayscale values ​​in the target area. Based on the obtained feature correction factors, the target features are reverse-corrected and verified to improve the matching degree between the features and historical reference features of the same category.

9. The target image recognition and target detection method based on video enhancement algorithm as described in claim 8, characterized in that, The reverse correction and verification process specifically involves: The contrast parameters of the gray-level co-occurrence matrix are adjusted by the texture correction factor, the threshold setting of the Canny operator is optimized by the edge correction factor, and the dynamic range is stretched by the gray-level correction factor. Finally, the corrected target features are output. After the reverse correction, the feature deviation value between the corrected target feature and the historical reference feature of the same category is recalculated, and the dual recognition verification value is regenerated based on the corrected target feature. If the regenerated dual recognition verification value is within the preset reliability range, the reverse correction is deemed valid, the corrected target features and corresponding detection results are retained, and included in the final valid detection set. Otherwise, further determine whether the recalculated characteristic deviation value has decreased by more than a preset amount compared to before the reverse correction: If so, it means that although the reverse correction did not fully meet the target, it has an improvement effect. The feature correction factor needs to be readjusted and the reverse correction process needs to be repeated. If not, the reverse correction is deemed invalid, the target detection result is marked as suspicious data, and the manual review process is initiated. After all target features have been reverse-corrected and verified, all valid detection results are summarized to ensure the accuracy and reliability of the output data.

Citation Information

Patent Citations

  • Target tracking control method and device based on image recognition and camera

    CN119338857A

  • Manipulator grabbing method based on deep learning target detection and image segmentation

    CN120563819A

  • Image feature calibration method based on motion offset

    CN120655549A

  • Target tracking method and system based on AI vision

    CN120976875A