Target detection method, related device, equipment and storage medium

By globally pooling the first feature map of the detected image, determining the target area, reducing interference from non-target areas, solving the problem of high hardware resource occupancy of traditional target detection methods, and improving computing efficiency and detection accuracy.

CN119963891APending Publication Date: 2025-05-09JIANGSU YIXING ZHILIAN AUTOMOBILE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411971140.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Traditional object detection methods have a high occupation of hardware resources, especially in the multi-objective detection scenarios at night, where there are problems such as complex lighting conditions and excessive hardware resources.

Method used

By globally pooling the first feature map of the detected image, the global light intensity map and the local light change map are obtained. The target area is determined based on these images, and the interference of the non-target area to the second feature map is reduced, and the redundant calculation and the number of parameters are reduced.

Benefits of technology

It improves computing efficiency, reduces the occupation of hardware resources, and enhances detection accuracy and adaptability in vehicle multi-object detection scenarios at night.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963891A_ABST
    Figure CN119963891A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method, a related device, equipment and a storage medium, and the method comprises the steps: carrying out the feature extraction of a to-be-detected image based on the shooting of a target object, and obtaining a first feature graph of the to-be-detected image; performing global pooling in the channel dimension based on the first feature map to obtain a global illumination intensity map, and obtaining a local illumination change map based on the feature value of each pixel in the first feature map; based on the global illumination intensity graph and the local illumination change graph, determining a target area in the first feature graph; performing convolution based on a target area in the first feature map to obtain a second feature map; performing prediction based on the second feature map to obtain a detection result of the target object in the to-be-detected image; wherein the detection result at least comprises the target position of the target object. Therefore, redundancy calculation and the number of parameters are reduced, the calculation efficiency is improved, and then occupation of hardware resources is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a target detection method and related devices, equipment, and storage media. Background Art

[0002] Object detection refers to the rapid and accurate recognition of objects (such as vehicles) in the acquired images. Traditional object detection methods mainly rely on image processing and machine learning techniques. These methods usually require image preprocessing, such as denoising and enhancement, and then extract features from the image. Finally, classifiers are used to classify the features to achieve object detection.

[0003] However, traditional target detection methods have the problem of occupying too much hardware resources. Summary of the invention

[0004] The present application at least provides a target detection method and related devices, equipment, and storage media to reduce the occupancy of hardware resources.

[0005] In a first aspect, the present application provides a target detection method, comprising extracting features of an image to be detected taken of a target object to obtain a first feature map of the image to be detected; performing global pooling in a channel dimension based on the first feature map to obtain a global illumination intensity map, and obtaining a local illumination change map based on the characteristic values ​​of each pixel in the first feature map; determining a target area in the first feature map based on the global illumination intensity map and the local illumination change map; performing convolution based on the target area in the first feature map to obtain a second feature map; performing prediction based on the second feature map to obtain a detection result of the target object in the image to be detected; wherein the detection result includes at least a target position of the target object.

[0006] Therefore, a global illumination intensity map is obtained by globally pooling the first feature map in the channel dimension, and a local illumination change map is obtained by using the feature values ​​of each pixel in the second feature map. Based on the global illumination intensity map and the local illumination change map, the target area in the first feature map is determined; that is, the first feature map is screened according to the local illumination change map and the global illumination intensity map to determine the target area, and subsequently only the target area needs to be convolved to obtain the second feature map, thereby reducing the interference of other information in the first feature map except the target area on the second feature map, reducing redundant calculations and the number of parameters, improving calculation efficiency, and thus reducing the occupancy of hardware resources.

[0007] The second aspect of the present application provides a target detection device, including a first feature map extraction module, an illumination information extraction module, a target area determination module, a second feature map extraction module, and a prediction module, wherein the first feature map extraction module is used to perform feature extraction based on an image to be detected taken of a target object to obtain a first feature map of the image to be detected; the illumination information extraction module is used to perform global pooling in the channel dimension based on the first feature map to obtain a global illumination intensity map, and obtain a local illumination change map based on the characteristic values ​​of each pixel in the second feature map; the target area determination module is used to determine a target area in the first feature map based on the global illumination intensity map and the local illumination change map; the second feature map extraction module is used to perform convolution based on the target area in the first feature map to obtain a second feature map; the prediction module is used to perform prediction based on the second feature map to obtain a detection result of the target object in the image to be detected; wherein the detection result at least includes the target position of the target object.

[0008] A third aspect of the present application provides an electronic device, comprising a memory and a processor coupled to each other, wherein the processor is used to execute program instructions stored in the memory to implement the above-mentioned target detection method.

[0009] A fourth aspect of the present application provides a computer-readable storage medium having program instructions stored thereon, which implement the target detection method in the above-mentioned first aspect when the program instructions are executed by a processor.

[0010] The above scheme obtains a global illumination intensity map by globally pooling the first feature map in the channel dimension, and obtains a local illumination change map by the feature values ​​of each pixel in the second feature map, and determines the target area in the first feature map based on the global illumination intensity map and the local illumination change map; that is, the first feature map is screened according to the local illumination change map and the global illumination intensity map to determine the target area therein, and subsequently only the target area needs to be convolved to obtain the second feature map, thereby reducing the interference of other information in the first feature map except the target area on the second feature map, reducing redundant calculations and the number of parameters, improving calculation efficiency, and thus reducing the occupancy of hardware resources.

[0011] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application.

[0013] Figure 1 It is a flowchart of an embodiment of the target detection method of the present application;

[0014] Figure 2 is a schematic diagram of a framework of a multi-scale global context enhancement module in some embodiments of the present application;

[0015] Figure 3 It is a schematic diagram of the framework of an embodiment of the electronic device of the present application;

[0016] Figure 4 A schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application;

[0017] Figure 5 It is a schematic diagram of the framework of an embodiment of the target detection device of the present application;

[0018] Figure 6 A summary table of experimental results of applying the target detection method of this application and other algorithms in the prior art to detect nighttime vehicle datasets. DETAILED DESCRIPTION

[0019] The scheme of the embodiment of the present application is described in detail below in conjunction with the drawings of the specification.

[0020] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0021] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the objects associated before and after are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of, for example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C.

[0022] The target detection method provided in this application can be applied to application scenarios such as intelligent traffic management systems and automatic driving assistance systems. For example, night vehicle detection can be performed in application scenarios such as intelligent traffic management systems and automatic driving assistance systems. With the rapid development of urban traffic and the continuous improvement of the degree of intelligence, night vehicle detection has become an important part of the intelligent transportation system. The goal of night vehicle detection is to achieve rapid and accurate identification of vehicles on the road in a dark environment at night, and provide real-time and accurate traffic information for the intelligent transportation system. Traditional night vehicle detection methods mainly rely on image processing and machine learning techniques. These methods usually require preprocessing of the image first, such as denoising, enhancement and other operations, and then extracting features from the image, and finally using a classifier to classify the features to achieve vehicle detection. However, these methods have some problems in night environments: first, the light is dark at night and the image quality is poor, which makes it difficult to extract features; second, there are a lot of shadows and noise interference in the night environment, which easily affects the accuracy of the detection results; finally, traditional vehicle detection methods can often only detect a single target, and the effect of multi-target detection is not good. It should be noted that the above application scenarios are only exemplary scenarios in which the target detection method provided by this application can be applied. The target detection method of this application can also be applied to other scenarios, such as human body detection, etc., and this application does not limit it. For the convenience of description, the principle of the target detection method is exemplified below by taking the target object as a vehicle as an example.

[0023] The target detection algorithm based on deep learning can automatically learn features from images without manually designing feature extractors, and has strong robustness and generalization capabilities. However, traditional algorithms still have some shortcomings in multi-target detection of vehicles at night, such as complex lighting conditions at night, difficulty in detecting target vehicles, and excessive occupation of hardware resources.

[0024] Some embodiments of the target detection method proposed in this application are intended to alleviate at least one of the above-mentioned problems existing in traditional algorithms.

[0025] See also Figure 1 , the target detection method provided in this application includes:

[0026] Step S100: extracting features of an image to be detected taken from a target object to obtain a first feature map of the image to be detected.

[0027] Step S200: Perform global pooling in the channel dimension based on the first feature map to obtain a global illumination intensity map, and obtain a local illumination change map based on the characteristic values ​​of each pixel in the first feature map.

[0028] Step S300: Determine a target area in the first feature map based on the global illumination intensity map and the local illumination variation map.

[0029] Step S400: performing convolution based on the target area in the first feature map to obtain a second feature map.

[0030] Step S500: Predicting based on the second feature map to obtain a detection result of the target object in the image to be detected; wherein the detection result at least includes a target position of the target object.

[0031] The above scheme obtains a global illumination intensity map by globally pooling the first feature map in the channel dimension, and obtains a local illumination change map by the feature value of each pixel in the second feature map. Based on the global illumination intensity map and the local illumination change map, the target area in the first feature map is determined; the local illumination change map can reflect the area with drastic illumination changes in the image to be detected, such as the area near the light source in the image to be detected. In other words, the first feature map is screened according to the local illumination change map and the global illumination intensity map to determine the target area therein, and subsequently only the target area needs to be convolved to obtain the second feature map, thereby reducing the interference of other information in the first feature map except the target area on the second feature map, reducing redundant calculations and the number of parameters, improving calculation efficiency, and thus reducing the occupation of hardware resources.

[0032] In some embodiments, convolution is performed based on the target area in the first feature map to obtain the second feature map according to the following formula:

[0033]

[0034] Among them, F o is the second feature map; Conv is the standard convolution operation; F i is the first feature map; Represents element-wise multiplication.

[0035] In some embodiments, the image to be detected taken by the target object can be input into the target detection model for feature extraction, such as the YOLO series network. In some embodiments, the image to be detected taken by the target object can be input into the YOLOv7 network for feature extraction. The backbone network of YOLOv7 is responsible for extracting features from the image to be detected, and multiple feature extraction layers are used in the backbone network of YOLOv7 to successively extract features to improve feature expression capabilities. For example, in step S100, the image to be detected can be feature extracted by the CBS module to obtain the first feature map of the image to be detected. The definition of the CBS module is well known in the art and will not be repeated here. And it can be understood that the feature extraction layer used for feature extraction of the image to be detected to obtain the first feature map in step S100 can be other modules, such as stem modules, etc., in addition to the CBS module, and this application is not limited.

[0036] In some embodiments, determining the target area in the first feature map based on the global illumination intensity map and the local illumination variation map includes:

[0037] Step S310: Fusing the global illumination intensity map and the local illumination change map to obtain a fused illumination weight map; wherein the fused illumination weight map includes the illumination weights of each pixel in the first feature map.

[0038] In some embodiments, the global illumination intensity map and the local illumination change map may be fused according to the following formula to obtain a fused illumination weight map:

[0039] W light =σ(α·L avg +β·L std + b);

[0040] Among them, W light is the illumination weight of each pixel in the first feature map; α and β are learnable parameters, σ is the Sigmoid activation function; L avg is the global illumination intensity vector in the global illumination intensity map; L std is the local illumination change map, and b is a constant (which can be adjusted according to actual conditions).

[0041] Step S320: Determine a target area in the first feature map based on the illumination weights in the fused illumination weight map.

[0042] In some embodiments, since the illumination weights of each pixel in the first feature map are determined in step S310, step S320 can take the set of pixels whose illumination weights are greater than a predetermined weight threshold as the target area in the first feature map. Thus, more attention can be paid to the target area (such as a strong halo area and / or an area with uneven illumination) in the future, while reducing the interference of useless information, thereby improving the convolution efficiency and the accuracy of feature extraction.

[0043] In some embodiments, obtaining a local illumination change map based on the feature values ​​of each pixel in the first feature map includes:

[0044] Step S210: Obtain several local areas in the first feature map.

[0045] Step S220: Calculate the standard deviation based on the characteristic values ​​of the pixels in each of the local areas to obtain the local illumination change map.

[0046] Exemplarily, obtaining a local illumination change map based on the feature values ​​of each pixel in the first feature map may be to divide the first feature map into a number of local regions, and calculate the standard deviation of the feature values ​​of each pixel in each local region to form a local illumination change map. Since the standard deviation can reflect the illumination difference in each local region, for example, it can reflect the brightness difference, therefore, the region with drastic illumination change can be determined according to the local illumination change map.

[0047] Dividing the first feature map into a number of local areas may be to evenly divide the first feature map into a number of local areas of the same size (i.e., each local area contains the same number of pixels); or, the image to be detected may be analyzed first, for example, the image to be detected may be converted into a YUV format, and the brightness of each pixel may be determined according to the YUV value of each pixel, and the first feature map may be divided into a number of local areas according to the brightness of the pixel, etc. The specific method of dividing the first feature map into a number of local areas is not limited in this application.

[0048] In some embodiments, the predicting based on the second feature map to obtain the detection result of the target object in the image to be detected includes:

[0049] Step S510: performing multi-scale feature extraction based on the second feature map to obtain a plurality of third feature maps of different scales.

[0050] In some embodiments, the multi-scale feature extraction includes at least one of one-dimensional convolution feature extraction and adaptive feature extraction, and the adaptive feature extraction is used to regard the second feature map as the first feature map to perform the steps based on the global illumination intensity map and the local illumination change map until a new second feature map is extracted, and the new second feature map serves as the third feature map of the adaptive feature extraction.

[0051] For example, please refer to Figure 2 , the second feature map can be input into four convolution branches (branch 1-branch 4) to extract information of different scales. For example, branch 1 and branch 2 are 1×1 convolution kernels for processing small targets or complex scenes; branch 3 focuses on adaptive feature extraction under illumination changes; branch 4 focuses on feature extraction under the combination of multi-scale and illumination adaptation. Among them, branch 3 and branch 4 can be light adaptive convolution (LAC) kernels.

[0052] Step S520: performing fusion based on the plurality of third feature maps to obtain a fused feature map.

[0053] Step S530: Predict based on the fused feature map to obtain the attention weight of each pixel in the fused feature map.

[0054] Step S540: Based on the attention weight of each pixel in the fused feature map, the feature value of each pixel in the fused feature map is weighted accordingly to obtain a weighted feature map.

[0055] Step S550: Predicting based on the weighted feature map to obtain a detection result of the target object in the image to be detected.

[0056] In some embodiments, the performing of prediction based on the fused feature map to obtain the attention weight of each pixel in the fused feature map includes:

[0057] Step S531: Perform global average pooling based on the fused feature map to obtain a global context vector.

[0058] Based on the fused feature map, global average pooling is performed to obtain a global context vector which can be expressed as:

[0059] G context =GAP(F ms ).

[0060] Among them, G context is the global context vector; GAP() represents the global average pooling operation; F ms Represents the fused feature map.

[0061] Step S532: Perform full connection processing based on the global context vector to obtain the attention weight.

[0062] For example, in step S532, please refer to Figure 2 , the global context vector G context After two fully connected layers, the attention weights are generated:

[0063] W global =σ(λ1·ReLU(λ2·G contexy +b2)+b3);

[0064] Among them, λ1 and λ2 are learnable parameters, σ is the Sigmoid activation function; b2 and b3 are constants; W global is the attention weight.

[0065] In some embodiments, the feature values ​​of each pixel in the fused feature map are weighted accordingly based on the attention weights of each pixel in the fused feature map, and the weighted feature map obtained can be obtained by applying the attention weights to the feature values ​​of each pixel in the fused feature map, which can be specifically expressed as the following formula:

[0066]

[0067] Among them, F out is the weighted feature map, F ms is the fusion feature map, Represents element-by-element multiplication, that is, multiplying the feature value of each pixel in the fusion feature map by the attention weight W global .

[0068] The above scheme extracts multi-scale features based on the second feature map to obtain several third feature maps of different scales, fuses the several third feature maps to obtain a fused feature map, predicts based on the fused feature map to obtain the attention weight of each pixel in the fused feature map, weights the feature values ​​of each pixel in the fused feature map based on the attention weight of each pixel in the fused feature map to obtain a weighted feature map, predicts based on the weighted feature map to obtain the detection result of the target object in the image to be detected, and enhances the perception of global context information at each feature position. The weighted feature map finally outputted fuses multi-scale features and global context information, and has higher feature expression capability and adaptability to complex lighting scenes.

[0069] In some embodiments, the image to be detected includes at least a nighttime vehicle image, and the detection result is obtained by detecting the image to be detected by a target detection model. Deep learning technology is one of the key solutions to the problem of multi-target detection of nighttime vehicles. Research in the field of deep learning target detection can basically be divided into two directions: one is a two-stage detector based on candidate regions, such as Faster-RCNN; the other is a single-stage detector based on regression calculation, such as YOLO. In the embodiments provided in this application, a YOLO series model (e.g., YOLOv7) is used as a target detection model. Compared with Faster-RCNN, the detection speed of the YOLO algorithm is faster. In other embodiments, Faster-RCNN can also be used as a target detection model, which is not limited by this application.

[0070] In some embodiments, since the backbone network of the original YOLOv7 adopts the CSPDarkNet53 structure, which is a dual-branch structure, although it has strong feature extraction capabilities, in the scene of multi-target detection at night, light source interference causes a large number of uneven light and dark areas to appear in the feature map. Standard convolution treats all areas equally and is easily interfered by the halo effect, thereby weakening the feature extraction ability of key target areas. In addition, CSPDarkNet53 focuses more on the extraction of local features, ignores the fusion of global context information, and is difficult to fully capture the overall impact of the lighting environment on target features. To this end, the light adaptive convolution module (LAC) and the multi-scale global context enhancement module (MGAM) are used to replace the 3×3 standard convolution layer in the ELAN module in the original YOLOv7 backbone network CSPDarkNet53. Among them, the LAC module can execute steps S200-S400 in the target detection method. Through the LAC module, the focus area of ​​the convolution process can be dynamically adjusted, and more attention can be paid to the strong halo area or the uneven illumination area, while reducing the interference of useless information, thereby improving the convolution efficiency and the accuracy of feature extraction; the MGAM module can execute step S500 in the target detection method, please refer to Figure 2 , Figure 2 MGAM module is a schematic diagram of the framework of the MGAM module in some embodiments of the present application. The operations performed by the MGAM module can refer to the description of step S500, step S510-step S550, step S531-step S532 in the aforementioned embodiments, which will not be repeated here. The MGAM module combines multi-scale features with global context modeling, and improves the sensitivity of the target detection model to nighttime vehicle features through the attention mechanism, thereby enhancing the perception of global context information at each feature position. The feature F finally output by the LAC module and the MGAM moduleoit It integrates multi-scale features and global context information, and has higher feature expression ability and adaptability to complex lighting scenes.

[0071] In some embodiments, the target detection model is trained based on a first sample image in a sample image set, and the first sample image is obtained by subjecting a second sample image in an initial image set to at least one data enhancement of halo simulation and exposure adjustment.

[0072] In some embodiments, the step of simulating the halo includes: determining a highlight area in the second sample image whose brightness is greater than a preset brightness threshold; and performing a Gaussian blur operation on the highlight area.

[0073] In some embodiments, the exposure adjustment step includes: adjusting the brightness and / or contrast of the second sample image.

[0074] In some embodiments, the highlight area in the image can be detected first, and the halo effect can be simulated using Gaussian blur, and the intensity, range and distribution of the halo can be randomly adjusted. Then, the exposure adjustment function in the Albumentations framework can be used to adjust the local brightness and contrast of the image to simulate overexposure or underexposure and enhance the dynamic range of the image. For example, the halo effect generated by the light source in the workshop scene can be simulated first to increase the realism of the image and provide a more suitable light source simulation for subsequent exposure adjustment. The highlight area in the image is detected by threshold segmentation (for example, a brightness threshold is set to identify all areas with brightness greater than the brightness threshold). Threshold segmentation helps to locate the area where the light source is located and provides regional information for subsequent processing. Gaussian blur is further applied to the highlight area to simulate the halo effect of the light source. By randomly adjusting the intensity, range and distribution of the halo, data diversity can be increased. Simulate different lighting conditions (overexposure, underexposure) to enhance the dynamic range of the image, so that the image has better performance and stronger visual effects under different lighting changes. Randomly adjust the local brightness and contrast of the image to simulate overexposure or underexposure. After simulating halo and exposure adjustment in highlight areas, the dynamic range of the image is expanded. The highlight areas of the image (such as lights and reflective surfaces) will be more prominent, and the details of the shadows will be preserved. This helps the target detection model to better handle night scenes with large brightness differences. The above scheme combines different enhancement methods to improve the diversity of the data set and the variability of the image, ensure that the training data can cover more actual scenes, enhance the robustness of the model under complex lighting conditions, and generate training data sets covering a variety of lighting conditions, thereby improving the detection accuracy of the target detection model in night or low-light scenes.

[0075] In some embodiments, the sample images in the sample image set may include various types of vehicles (such as cars, buses, trucks, motorcycles, etc.) to improve the detection performance of the target detection model trained based on the sample image set for various types of vehicles, thereby improving the accuracy of target detection.

[0076] In order to comprehensively evaluate the performance of the improved algorithm (i.e., the target detection method proposed in this application), the improved algorithm was compared and analyzed with the current mainstream target detection algorithm. In the experiment, multiple evaluation indicators were used to quantify the performance of the algorithm, including the average precision (AP) of a single category, the average precision (mAP) of all categories, the model volume, and the number of frames per second (FPS). These indicators can comprehensively reflect the performance of the algorithm in terms of accuracy, efficiency, and model complexity. The experimental results are as follows: Figure 6 shown.

[0077] from Figure 6 It can be seen that in dim scenes, compared with the benchmark YOLOv7, the algorithm proposed in this case has improved mAP by 2.28 percentage points, significantly improving the detection accuracy. The optimized network not only enhances the adaptability to complex scenes, but also optimizes the feature extraction process, thereby improving the inference speed and increasing the frame rate from 61 to 67. In addition, compared with traditional single-stage and two-stage detectors, the improved algorithm has also been significantly improved in comprehensive capabilities.

[0078] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.

[0079] See also Figure 3 , Figure 3 : is a schematic diagram of a framework of an embodiment of an electronic device of the present application. The electronic device 30 includes a memory 31 and a processor 32 coupled to each other, and the processor 32 is used to execute program instructions stored in the memory 31 to implement the steps in any of the above target detection method embodiments. In a specific implementation scenario, the electronic device 30 may include but is not limited to: a microcomputer, a server, and in addition, the electronic device 30 may also include a mobile device such as a laptop computer and a tablet computer, which is not limited here.

[0080] Specifically, the processor 32 is used to control itself and the memory 31 to implement the steps in any of the above-mentioned target detection method embodiments. The processor 32 can also be called a CPU (Central Processing Unit). The processor 32 may be an integrated circuit chip with signal processing capabilities. The processor 32 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 32 can be implemented by an integrated circuit chip.

[0081] The above scheme obtains a global illumination intensity map by globally pooling the first feature map in the channel dimension, and obtains a local illumination change map by the feature values ​​of each pixel in the second feature map, and determines the target area in the first feature map based on the global illumination intensity map and the local illumination change map; that is, the first feature map is screened according to the local illumination change map and the global illumination intensity map to determine the target area therein, and subsequently only the target area needs to be convolved to obtain the second feature map, thereby reducing the interference of other information in the first feature map except the target area on the second feature map, reducing redundant calculations and the number of parameters, improving calculation efficiency, and thus reducing the occupancy of hardware resources.

[0082] See also Figure 4 , Figure 4 The computer-readable storage medium 40 stores program instructions 401 that can be executed by a processor, and the program instructions 401 are used to implement the steps in any of the above-mentioned target detection method embodiments.

[0083] The above scheme obtains a global illumination intensity map by globally pooling the first feature map in the channel dimension, and obtains a local illumination change map by the feature values ​​of each pixel in the second feature map, and determines the target area in the first feature map based on the global illumination intensity map and the local illumination change map; that is, the first feature map is screened according to the local illumination change map and the global illumination intensity map to determine the target area therein, and subsequently only the target area needs to be convolved to obtain the second feature map, thereby reducing the interference of other information in the first feature map except the target area on the second feature map, reducing redundant calculations and the number of parameters, improving calculation efficiency, and thus reducing the occupancy of hardware resources.

[0084] See also Figure 5 , Figure 5 It is a schematic diagram of the framework of an embodiment of the target detection device of the present application. The target detection device 50 includes: a first feature map extraction module 51, which is used to extract features based on the image to be detected taken of the target object, and obtain the first feature map of the image to be detected; an illumination information extraction module 52, which is used to perform global pooling in the channel dimension based on the first feature map to obtain a global illumination intensity map, and obtain a local illumination change map based on the characteristic values ​​of each pixel in the second feature map; a target area determination module 53, which is used to determine the target area in the first feature map based on the global illumination intensity map and the local illumination change map; a second feature map extraction module 54, which is used to perform convolution based on the target area in the first feature map to obtain a second feature map; a prediction module 55, which is used to perform prediction based on the second feature map to obtain the detection result of the target object in the image to be detected; wherein the detection result at least includes the target position of the target object.

[0085] The above scheme obtains a global illumination intensity map by globally pooling the first feature map in the channel dimension, and obtains a local illumination change map by the feature values ​​of each pixel in the second feature map, and determines the target area in the first feature map based on the global illumination intensity map and the local illumination change map; that is, the first feature map is screened according to the local illumination change map and the global illumination intensity map to determine the target area therein, and subsequently only the target area needs to be convolved to obtain the second feature map, thereby reducing the interference of other information in the first feature map except the target area on the second feature map, reducing redundant calculations and the number of parameters, improving calculation efficiency, and thus reducing the occupancy of hardware resources.

[0086] In some embodiments, the target area determination module 53 includes: an illumination weight calculation unit, used to fuse the global illumination intensity map and the local illumination change map to obtain a fused illumination weight map; wherein the fused illumination weight map contains the illumination weights of each pixel in the first feature map; and a target area determination unit, used to determine the target area in the first feature map based on the illumination weights in the fused illumination weight map.

[0087] In some embodiments, the illumination information extraction module 52 includes: an area acquisition unit, used to acquire several local areas in the first feature map; a standard deviation calculation unit, used to calculate the standard deviation based on the characteristic values ​​of pixels in each of the local areas to obtain the local illumination change map.

[0088] In some embodiments, the prediction module 55 includes: a third feature map extraction unit, used to perform multi-scale feature extraction based on the second feature map to obtain several third feature maps of different scales; a fusion unit, used to fuse based on the several third feature maps to obtain a fused feature map; an attention weight prediction unit, used to predict based on the fused feature map to obtain the attention weight of each pixel in the fused feature map; a weighted feature calculation unit, used to perform corresponding weighting on the feature values ​​of each pixel in the fused feature map based on the attention weight of each pixel in the fused feature map to obtain a weighted feature map; a prediction unit, used to predict based on the weighted feature map to obtain the detection result of the target object in the image to be detected.

[0089] In some embodiments, the attention weight prediction unit includes: a global average pooling subunit, used to perform global average pooling based on the fused feature map to obtain a global context vector; a fully connected subunit, used to perform fully connected processing based on the global context vector to obtain the attention weight.

[0090] In some embodiments, the multi-scale feature extraction includes at least one of one-dimensional convolution feature extraction and adaptive feature extraction, and the adaptive feature extraction is used to regard the second feature map as the first feature map to perform the steps based on the global illumination intensity map and the local illumination change map until a new second feature map is extracted, and the new second feature map serves as the third feature map of the adaptive feature extraction.

[0091] In some embodiments, the image to be detected includes at least a night vehicle image, and the detection result is obtained by detecting the image to be detected by a target detection model, and the target detection model is trained based on a first sample image in a sample image set, and the first sample image is obtained by enhancing the second sample image in the initial image set through at least one of halo simulation and exposure adjustment.

[0092] In some embodiments, the step of simulating the halo includes: determining a highlight area in the second sample image whose brightness is greater than a preset brightness threshold; and performing a Gaussian blur operation on the highlight area.

[0093] In some embodiments, the exposure adjustment step includes: adjusting the brightness and / or contrast of the second sample image.

[0094] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0095] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0096] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0097] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0098] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

[0099] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

Claims

1. A target detection method, characterized in that: include: Based on extracting features of the image to be detected shot from the target object, a first feature map of the image to be detected is obtained; Performing global pooling in the channel dimension based on the first feature map to obtain a global illumination intensity map, and obtaining a local illumination change map based on the feature value of each pixel in the first feature map; Determine a target area in the first feature map based on the global illumination intensity map and the local illumination variation map; Performing convolution based on the target area in the first feature map to obtain a second feature map; Prediction is performed based on the second feature map to obtain a detection result of the target object in the image to be detected; wherein the detection result at least includes a target position of the target object.

2. The method according to claim 1, characterized in that: The determining the target area in the first feature map based on the global illumination intensity map and the local illumination variation map includes: Based on the global illumination intensity map and the local illumination change map, a fused illumination weight map is obtained; wherein the fused illumination weight map includes the illumination weight of each pixel in the first feature map; Based on the illumination weights in the fused illumination weight map, a target area in the first feature map is determined.

3. The target detection method according to claim 1, characterized in that: The obtaining of a local illumination change map based on the characteristic values ​​of each pixel in the first characteristic map includes: Acquire several local areas in the first feature map; The standard deviation is calculated based on the characteristic values ​​of the pixels in each of the local areas to obtain the local illumination change map.

4. The target detection method according to claim 1, characterized in that: The performing prediction based on the second feature map to obtain a detection result of the target object in the image to be detected includes: Performing multi-scale feature extraction based on the second feature map to obtain a plurality of third feature maps of different scales; Perform fusion based on the plurality of third feature maps to obtain a fused feature map; Perform prediction based on the fused feature map to obtain the attention weight of each pixel in the fused feature map; Based on the attention weight of each pixel in the fused feature map, the feature value of each pixel in the fused feature map is weighted accordingly to obtain a weighted feature map; Prediction is performed based on the weighted feature map to obtain a detection result of the target object in the image to be detected.

5. The method according to claim 4, characterized in that The step of performing prediction based on the fused feature map to obtain the attention weight of each pixel in the fused feature map includes: Performing global average pooling based on the fused feature map to obtain a global context vector; Fully connected processing is performed based on the global context vector to obtain the attention weight.

6. The method according to claim 4, characterized in that The multi-scale feature extraction includes at least one of one-dimensional convolution feature extraction and adaptive feature extraction, and the adaptive feature extraction is used to regard the second feature map as the first feature map to perform the steps based on the global illumination intensity map and the local illumination change map until a new second feature map is extracted, and the new second feature map is used as the third feature map of the adaptive feature extraction.

7. The method according to claim 1, characterized in that The image to be detected includes at least a night vehicle image. The detection result is obtained by detecting the image to be detected by a target detection model. The target detection model is trained based on a first sample image in a sample image set. The first sample image is obtained by enhancing the second sample image in an initial image set by at least one of halo simulation and exposure adjustment.

8. The method according to claim 7, characterized in that The steps of halo simulation include: Determine a highlight area in the second sample image whose brightness is greater than a preset brightness threshold; A Gaussian blur operation is performed on the highlight area.

9. The method according to claim 7, characterized in that: The exposure adjustment step comprises: The brightness and / or contrast of the second sample image is adjusted.

10. A target detection device, characterized in that: include: A first feature map extraction module, used to extract features based on the image to be detected taken of the target object, to obtain a first feature map of the image to be detected; an illumination information extraction module, configured to perform global pooling in a channel dimension based on the first feature map to obtain a global illumination intensity map, and to obtain a local illumination change map based on a feature value of each pixel in the second feature map; A target region determination module, configured to determine a target region in the first feature map based on the global illumination intensity map and the local illumination variation map; A second feature map extraction module, used for performing convolution based on the target area in the first feature map to obtain a second feature map; A prediction module is used to make a prediction based on the second feature map to obtain a detection result of the target object in the image to be detected; wherein the detection result at least includes a target position of the target object.

11. An electronic device, characterized in that: It comprises a memory and a processor coupled to each other, wherein the processor is used to execute program instructions stored in the memory to implement the target detection method according to any one of claims 1 to 9.

12. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the target detection method according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Target detection method, target detection network training method and related device

    CN121236357A