Target detection method

By combining infrared and visible light images into a detection model, and leveraging the complementary nature of their information, target detection is achieved, solving the problem of low detection accuracy under adverse conditions and realizing higher precision and efficiency in detection.

CN120912933BActive Publication Date: 2026-08-25CHINA RAILWAY FIRST SURVEY & DESIGN INST GRP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510741762.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2026-08-25
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

Existing target detection methods have low accuracy and large errors under adverse environments (such as rain, fog, insufficient light or complex background).

Method used

By combining infrared and visible light images, target detection is performed separately using a detection model. Multiple local features of the same object are associated and fused for verification, and the final detection result is output.

Benefits of technology

It improves the accuracy and comprehensiveness of target detection, reduces the impact of complex environments on detection results, improves detection efficiency, and reduces energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912933B_ABST
    Figure CN120912933B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a target detection method, relates to the technical field of image processing, and comprises the following steps: obtaining an infrared image and a visible light image of a first region; detecting objects in a detection region in the infrared image and a detection region in the visible light image respectively to obtain a first detection result and a second detection result; associating the same objects in the first detection result and the second detection result to obtain an association result; wherein the association result at least indicates associated objects and / or unassociated objects in the first detection result and the second detection result; based on a plurality of local features of each object in the association result, each object in the association result is verified by fusion, and a detection result is output according to a fusion verification result. The target detection method provided by the embodiment of the present application can improve detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and more specifically, to a target detection method. Background Technology

[0002] Object detection is one of the core tasks in computer vision, aiming to detect objects of interest in images or videos and determine their locations. It has wide applications in security, transportation, and healthcare. With the continuous development of technology, techniques such as machine learning and deep learning are widely used in object detection to improve its accuracy.

[0003] Most current object detection methods are developed using visible light images as their data foundation. However, in adverse environments such as rain, fog, insufficient lighting, or complex backgrounds, the quality of the acquired visible light images is significantly reduced, leading to substantial detection errors when these images are processed. Therefore, there is an urgent need to provide an object detection method that can improve detection accuracy. Summary of the Invention

[0004] This application provides a target detection method that can improve the accuracy of target detection.

[0005] In a first aspect, embodiments of this application provide a target detection method, comprising: obtaining an infrared image and a visible light image of a first region; detecting objects in a detection region in the infrared image and a detection region in the visible light image, respectively, to obtain a first detection result and a second detection result; wherein the first detection result indicates an object detected in the infrared image, and the second detection result indicates an object detected in the visible light image; associating the same object in the first detection result and the second detection result to obtain an association result; wherein the association result at least indicates associated objects and / or unassociated objects in the first detection result and the second detection result; performing fusion verification on each object in the association result based on multiple local features of each object in the association result; and outputting a detection result based on the fusion verification result.

[0006] Optionally, the association result includes at least one of a first object, a second object, and a third object, wherein the first object is an unassociated object in the first detection result, the second object is an unassociated object in the second detection result, and the third object is an associated object in both the first and second detection results;

[0007] The step of fusing and verifying each object in the association result based on multiple local features of each object in the association result, and outputting a detection result based on the fusing and verification result, includes: extracting multiple local features of each object in the first detection result from the infrared image; extracting multiple local features of each object in the second detection result from the visible light image; for the first object, determining the probability that the first object is an object, background, and uncertain region based on the multiple local features of the first object based on the infrared image; for the second object, determining the probability that the second object is an object, background, and uncertain region based on the multiple local features of the second object based on the visible light image; for the third object, determining the probability that the third object is an object, background, and uncertain region based on the multiple local features of the third object based on both the infrared image and the visible light image; and outputting a detection result based on the probabilities that the first object, the second object, and the third object are objects, background, and uncertain regions respectively; wherein, the background is the region in the infrared image and the visible light image located outside the detection region.

[0008] Optionally, the plurality of local features includes four local features, namely local entropy, local contrast, local gradient, and local texture;

[0009] The step of determining the probability that the first object is an object, background, or uncertain region based on multiple local features of the infrared image includes: determining four sets of classification probabilities for the first object based on the local entropy, local contrast, local gradient, and local texture of the infrared image; wherein, the four sets of classification probabilities of the first object correspond one-to-one with the four local features; each of the four sets of classification probabilities of the first object includes the probability that the first object is an object, background, or uncertain region under the corresponding local feature; and determining the probability that the first object is an object, background, or uncertain region based on the four sets of classification probabilities of the first object.

[0010] Optionally, the plurality of local features includes four local features, namely local entropy, local contrast, local gradient, and local texture;

[0011] The step of determining the probability that the second object is an object, background, or uncertain region based on multiple local features of the visible light image includes: determining four sets of classification probabilities for the second object based on the local entropy, local contrast, local gradient, and local texture of the visible light image; wherein, the four sets of classification probabilities of the second object correspond one-to-one with the four local features; each of the four sets of classification probabilities of the second object includes the probability that the second object is an object, background, or uncertain region under the corresponding local feature; and determining the probability that the second object is an object, background, or uncertain region based on the four sets of classification probabilities of the second object.

[0012] Optionally, the plurality of local features includes four local features, namely local entropy, local contrast, local gradient, and local texture;

[0013] The step of determining the probability that the third object is an object, background, and uncertain region based on multiple local features of the infrared image and the visible light image respectively includes: determining four sets of classification probabilities of the third object based on the infrared image based on the local entropy, local contrast, local gradient, and local texture of the third object based on the infrared image; wherein, the four sets of classification probabilities of the third object based on the infrared image correspond one-to-one with the four local features; each set of classification probabilities of the third object based on the infrared image includes the probability that the third object is an object, background, and uncertain region based on the infrared image and the corresponding local features; determining a first probability that the third object is an object, background, and uncertain region based on the four sets of classification probabilities of the third object based on the infrared image; and further determining the probability that the third object is an object, background, and uncertain region based on the four sets of classification probabilities of the third object based on the infrared image. The third object is classified into four sets of probabilities based on the local entropy, local contrast, local gradient, and local texture of the visible light image. Each set of probabilities corresponds one-to-one with one of the four local features. Each set of probabilities includes the probability that the third object, based on the visible light image and under the corresponding local features, is an object, background, or an uncertain region. Based on the four sets of probabilities, a second probability is determined that the third object is an object, background, or an uncertain region. Finally, based on the first and second probabilities, the probability that the third object is an object, background, or an uncertain region is determined.

[0014] Optionally, the step of outputting the detection result based on the probabilities that the first object, the second object, and the third object are respectively an object, background, and uncertain region includes: selecting an object from the first object, the second object, and the third object whose probability of being the object is higher than that of the background and the uncertain region, as a fourth object; if the fourth object is the associated object, outputting the position and category of the fourth object based on the position and category of the fourth object in the infrared image and the visible light image; if the fourth object is the unassociated object, outputting the position and category of the fourth object based on the position and category of the fourth object in the infrared image or the visible light image.

[0015] Optionally, associating the same object in the first detection result and the second detection result to obtain the association result includes: constructing a first matrix of I rows and J columns based on the category and position of each object in the first detection result and the category and position of each object in the second detection result; wherein I indicates the number of objects in the first detection result, J indicates the number of objects in the second detection result, and the element located at the intersection of the i-th row and the j-th column indicates the distance between the object corresponding to the i-th row and the object corresponding to the j-th column; based on the first matrix, subtracting the minimum element in the corresponding row from each row element, and then subtracting the minimum element in the corresponding column from each column element to obtain a second matrix; performing a covering step: covering the zero elements in the second matrix with the minimum number of straight lines; If the number of lines equals the number of rows in the second matrix, the second matrix is ​​determined as the third matrix; if the number of lines is less than the number of rows in the second matrix, each non-zero element in the second matrix is ​​subtracted from the smallest non-zero element of the second matrix, and the zero elements in the second matrix covered by line intersections are added to the smallest non-zero element to obtain a fourth matrix; after using the fourth matrix as the second matrix, the covering step is executed from the beginning until the number of lines equals the number of rows in the second matrix; the row objects and column objects corresponding to the zero elements in the rows with only one zero element in the third matrix are determined as the associated objects, and the objects in the first detection result and the second detection result other than the associated objects are determined as the unassociated objects.

[0016] Optionally, detecting objects in the detection areas of the infrared image and the visible light image respectively to obtain a first detection result and a second detection result includes: identifying the detection areas from the infrared image and the visible light image respectively; performing background processing on the areas located outside the detection areas in the infrared image and the visible light image respectively; and detecting objects in the background-processed infrared image and the visible light image respectively to obtain the first detection result and the second detection result.

[0017] Optionally, the detection of objects in the detection areas of the infrared image and the visible light image respectively includes: detecting objects in the detection areas of the infrared image and the visible light image respectively using a detection model.

[0018] Optionally, the detection model is trained, and the training process includes: obtaining multiple visible light training samples and multiple infrared training samples; performing background processing on the regions outside the detection area in each of the visible light training samples and each of the infrared training samples; labeling the position and category of objects in each of the background-processed visible light training samples and each of the infrared training samples; amplifying the visible light training samples and the infrared training samples to which objects of the category with a smaller proportion belong; and training the neural network using the labeled and amplified visible light training samples and the infrared training samples to obtain the detection model.

[0019] In a second aspect, embodiments of this application provide an electronic device, including a memory and a processor; the memory is used to store a computer program; the processor is used to implement the target detection method described in any one of the first aspects when the computer program is executed.

[0020] The beneficial effects of the target detection method in the embodiments of this application are:

[0021] Because infrared images have the advantage of penetrating in complex environments such as low light and smoke, and visible light images have advantages in terms of detail, resolution, and color information, compared with the method of target detection using only visible light images, this embodiment of the application can improve the accuracy of target detection and reduce the impact of complex environments on the detection results by detecting objects in the detection areas of the infrared image and the detection areas of the visible light image of the first region.

[0022] Furthermore, since the detection is performed on objects within the detection area, users can set the area to be detected according to their needs, and detection is only performed within the detection area, which improves detection efficiency and reduces detection energy consumption.

[0023] Furthermore, by associating the same object in the first and second detection results, and based on multiple local features of each object in the association results, a fusion verification is performed on each object in the association results to output the detection result. In other words, by associating the detection results of infrared and visible light images, the complementary information of different image modalities is fully utilized, enabling more comprehensive and accurate object detection. Moreover, performing fusion verification based on multiple local features on each object in the association results avoids misjudgments that may occur due to single-feature judgments, greatly improving the accuracy of verification and making the output detection results more reliable and precise. Attached Figure Description

[0024] Figure 1 A schematic flowchart of a target detection method provided in an embodiment of this application;

[0025] Figure 2 This is a schematic diagram illustrating the process of detecting objects in a detection area, provided in an embodiment of this application.

[0026] Figure 3 A flowchart illustrating the process of determining the association result provided in an embodiment of this application;

[0027] Figure 4 A schematic diagram of the fusion verification process provided for embodiments of this application;

[0028] Figure 5 A flowchart illustrating the process of determining the probabilities of a third object being an object, a background, and an uncertain region, as provided in an embodiment of this application;

[0029] Figure 6 A schematic diagram illustrating the training process of the detection model provided in this application embodiment;

[0030] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0031] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0032] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0033] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0034] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0035] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0036] Visible light images provide rich color and texture information, which aligns with human visual habits. However, their performance can be significantly compromised in complex environments such as low light or occlusion. Therefore, target detection methods that rely solely on visible light images suffer from large detection errors.

[0037] Infrared images, due to their ability to capture thermal radiation, exhibit significant advantages in low-light or no-light environments, playing an irreplaceable role, especially in night vision surveillance and target detection. Therefore, to improve the accuracy of target detection, this application utilizes the complementary information of infrared and visible light images for target detection, thereby enhancing the accuracy and comprehensiveness of target detection.

[0038] Figure 1 This is a flowchart illustrating the target detection method provided in the embodiments of this application. The target detection method can be executed by an electronic device, including but not limited to desktop computers, laptops, mobile phones, and other electronic devices with data processing capabilities.

[0039] The target detection method can be applied to the detection of intrusive foreign objects in railway tracks, traffic bridges, tunnels, open-air tracks, roads, and green belts at different times or throughout the day.

[0040] For example, taking railway tracks as an example, intruding foreign objects include, but are not limited to, objects that could cause train operation safety incidents, such as stones, water bottles, animals, branches, and workers.

[0041] like Figure 1 As shown, the target detection method may include the following steps:

[0042] 101. Obtain the infrared and visible light images of the first region.

[0043] For example, in a railway track scenario, the first region can be the area where the railway tracks are located. In a road scenario, the first region can be the area where the road is located. In a green belt scenario, the first region can be the area where the green belt is located.

[0044] An infrared image of the first region can be obtained using an infrared camera, and a visible light image of the first region can be obtained using a visible light camera. It should be noted that the infrared and visible light images of the first region are acquired simultaneously.

[0045] 102. Detect objects in the detection areas of the infrared image and the detection areas of the visible light image respectively to obtain a first detection result and a second detection result.

[0046] The detection area is the region of interest in both infrared and visible light images.

[0047] For example, in a railway track scenario, the detection area can be the area defined by the railway. In a road scenario, the detection area can be the area defined by the road. In a green belt scenario, the detection area can be the area defined by the green belt.

[0048] An object can refer to a foreign object that intrudes into the detection area. For example, in a railway track scenario, an object is an object that intrudes into or on the track and affects the safe operation of the train, such as falling rocks, people, or animals.

[0049] The first detection result indicates the object detected from the infrared image. For example, the first detection result may indicate the location and category of the object detected from the infrared image.

[0050] The second detection result indicates the object detected in the visible light image. For example, the second detection result may indicate the location and category of the object detected in the visible light image.

[0051] The categories of objects include, but are not limited to, stones, people, and animals. The position of an object can be represented by a bounding box that encloses the object; the bounding box can be square, circle, etc.

[0052] The specific implementation process of 102 can be shown below:

[0053] The detection model is used to detect objects in the detection areas of infrared images and visible light images, respectively.

[0054] The specific implementation methods for detecting objects include, but are not limited to, the following two:

[0055] The first method involves inputting an infrared image into a detection model. This model detects objects within the infrared image and outputs the location and category of the detected objects. Then, based on the location of the detection area and the location of the detected objects in the infrared image, objects located within the detection area are filtered out. The set of locations and categories of these objects is then used as the first detection result.

[0056] The principle of detecting objects in a detection region of a visible light image using a detection model is the same as above, and will not be repeated here.

[0057] The second type, such as Figure 2 As shown, it includes the following steps:

[0058] 201. Identify the detection area from infrared and visible light images respectively.

[0059] 202. Perform background processing on the regions located outside the detection area in the infrared and visible light images respectively.

[0060] For example, backgrounding can be achieved by setting the areas outside the detection area in the infrared and visible light images to white, black, or mosaic colors.

[0061] 203. Detect objects in the background-processed infrared image and visible light image respectively to obtain the first detection result and the second detection result.

[0062] For example, the infrared image and the visible light image after background processing can be input into the detection model respectively, so that the detection model outputs the first detection result and the second detection result.

[0063] Obviously, there are no objects outside the detection area in the backgrounded infrared and visible light images. Therefore, we directly detect the objects in the backgrounded infrared and visible light images, and the result is the object detected from the detection area of ​​the infrared and visible light images.

[0064] It should be noted that the specific structure and training process of the detection model will be explained below, and will not be repeated here.

[0065] 103. Associate the same object in the first and second detection results to obtain the association result;

[0066] The association results indicate at least the associated objects and / or unassociated objects in the first and second detection results.

[0067] Associated objects indicate objects that are shared between infrared and visible light images. Unassociated objects indicate objects detected only from infrared or visible light images.

[0068] The specific implementation methods of 103 include, but are not limited to, the following two:

[0069] The first method calculates the distance between two objects of the same category in the first and second detection results. The distance can be the distance between the centers of the detection boxes of the two objects, or it can be the GIoU distance between the detection boxes of the two objects.

[0070] When the distance between two objects of the same category is less than a preset distance, the two objects of the same category are determined to be the same object (i.e., related objects). The preset distance can be set according to experimental requirements.

[0071] Objects in the first and second detection results that are not associated with other objects are identified as unassociated objects. Thus, the set of associated and unassociated objects constitutes the associated result.

[0072] The second type, such as Figure 3 As shown, the following steps may be included:

[0073] 301. Construct a first matrix with I rows and J columns based on the category and location of each object in the first detection result and the category and location of each object in the second detection result.

[0074] Where I indicates the number of objects in the first detection result (the number of objects detected in the detection area of ​​the infrared image), J indicates the number of objects in the second detection result (the number of objects detected in the detection area of ​​the visible light image), and the element at the intersection of the i-th row and the j-th column indicates the distance between the object corresponding to the i-th row and the object corresponding to the j-th column. Each row corresponds to one object detected from the detection area of ​​the infrared image, and each column corresponds to one object detected from the detection area of ​​the visible light image. i is an integer greater than zero and less than or equal to I, and j is an integer greater than zero and less than or equal to J.

[0075] For example, the first matrix It can be as follows:

[0076] ;

[0077] The distance between the object in the i-th row and the object in the j-th column The calculation method is as follows:

[0078] When the category of the object in row i is different from that of the object in column j, the distance... Set to infinity;

[0079] When the object in row i and the object in column j have the same category, the distance is... The calculation formula is as follows:

[0080]

[0081] In the formula, Let represent the area of ​​the minimum bounding rectangle of the detection box corresponding to the object in the i-th row and the detection box corresponding to the object in the j-th column. Indicates intersection, union, and ratio. This represents the area of ​​the intersection of the bounding box of the object in the i-th row and the bounding box of the object in the j-th column. U represents the area of ​​the union of the detection boxes of the objects in the i-th row and the objects in the j-th column.

[0082] 302. Based on the first matrix, subtract the smallest element in the corresponding row from each row element, and then subtract the smallest element in the corresponding column element from each column element to obtain the second matrix.

[0083] Specifically, in the first matrix, after subtracting the smallest element in the corresponding row from each row element, the smallest element in the corresponding column of the calculated matrix is ​​then subtracted from each column element of the calculated matrix to obtain the second matrix.

[0084] 303. Perform the covering step: Cover the zero elements in the second matrix with the minimum number of straight lines. These straight lines include both horizontal and vertical lines.

[0085] 304. If the number of lines is equal to the number of rows in the second matrix, then the second matrix is ​​determined to be the third matrix.

[0086] 305. If the number of lines is less than the number of rows in the second matrix, subtract the smallest non-zero element of the second matrix from each non-zero element in the second matrix, and add the smallest non-zero element to the zero elements in the second matrix that are covered by line intersections (i.e., intersections of horizontal and vertical lines) to obtain the fourth matrix.

[0087] 306. After taking the fourth matrix as the second matrix, start from the covering step (303) and continue until the number of lines equals the number of rows in the second matrix.

[0088] 307. The row objects (objects detected in the detection area of ​​the infrared image) and column objects (objects detected in the detection area of ​​the visible light image) corresponding to the zero element in the row with only one zero element in the third matrix are identified as associated objects, and the objects in the first detection result and the second detection result other than associated objects are identified as unassociated objects.

[0089] 104. Based on multiple local features of each object in the association results, perform fusion verification on each object in the association results, and output the detection results based on the fusion verification results.

[0090] Multiple local features include, but are not limited to, at least two of local entropy, local contrast, local gradient, and local texture. It should be noted that the following explanation uses an example of multiple local features including the four local features mentioned above.

[0091] For example, the association result may include at least one of a first object, a second object, and a third object, wherein the first object is an unassociated object in the first detection result, the second object is an unassociated object in the second detection result, and the third object is an associated object in the first and second detection results.

[0092] It can be understood that the first and second objects are the unrelated objects in the association results mentioned above, and the third object is the related object in the association results mentioned above.

[0093] Based on this, such as Figure 4 As shown, the fusion verification can be implemented as follows:

[0094] 401. Extract multiple local features of each object from the first detection result from the infrared image.

[0095] Local entropy reflects the degree of gray-level dispersion in an image region. The local entropy of an object can be determined based on the height and width of the region occupied by the object in the infrared image, as well as the gray-level value of each position in the region occupied by the object in the infrared image.

[0096] For example, the local entropy of an object It can be obtained through the following formula:

[0097]

[0098] Where M and N are the width and height of the area occupied by the object in the infrared image, respectively. The area occupied by the object in the infrared image The grayscale value at that location.

[0099] Local contrast reflects the detailed information of an image region. The local contrast of an object can be determined based on the object's inner window, outer window, the number and distribution of pixels in the inner window, the number and distribution of pixels in the outer window, and the pixel value at the corresponding position in the infrared image.

[0100] For example, the local contrast of an object It can be obtained through the following formula:

[0101]

[0102] in, This is the inner window of the object in the infrared image. This is the outer window of the object in the infrared image. This represents the number of pixels in the inner window of the object in the infrared image. This represents the number of pixels in the outer window of the object in the infrared image. This represents the mean pixel distribution of the object's inner window in the infrared image. This represents the mean pixel distribution of the outer window of the object in the infrared image. This represents the pixel value at the corresponding position.

[0103] Local gradients reflect the edge strength of an image region. The local gradient of an object can be determined based on the gradients of the area occupied by the object in the infrared image in the horizontal and vertical directions, as well as the area occupied by the object in the infrared image.

[0104] For example, the local gradient of an object It can be obtained through the following formula:

[0105]

[0106] in, This represents the gradient of the area occupied by the object in the infrared image in the horizontal direction. Let I be the gradient of the region occupied by the object in the infrared image in the vertical direction, where I is the region occupied by the object in the infrared image, and * indicates convolution calculation.

[0107] Local texture reflects the texture difference between an object and the background. The local texture of an object can be determined based on the object's inner window, the pixel distribution of the inner window, and the pixel values ​​at the corresponding positions in the infrared image.

[0108] For example, a local texture of an object It can be obtained through the following formula:

[0109]

[0110] in, This is the inner window of the object in the infrared image. This represents the mean of the pixel distribution within the inner window of the object in the infrared image. This represents the pixel value at the corresponding position.

[0111] Clearly, the above formula can be used to extract the local entropy, local contrast, local gradient, and local texture of each object in the first detection result from the infrared image.

[0112] 402. Extract multiple local features of each object from the second detection result from the visible light image.

[0113] The specific implementation process of 402 can be referred to 401. The difference is that the relevant parameters of the infrared image in the formula are replaced with the relevant parameters of the visible light image.

[0114] 403. For the first object, based on multiple local features of the first object in the infrared image, determine the probabilities of the first object being an object, background, and uncertain region.

[0115] For example, the implementation process can be as follows:

[0116] First, based on the local entropy, local contrast, local gradient, and local texture of the first object in the infrared image, four sets of classification probabilities for the first object are determined.

[0117] The four classification probabilities of the first object correspond one-to-one with the four local features. Each of the four classification probabilities of the first object includes the probability that the first object is an object, background, or an uncertain region under the corresponding local feature.

[0118] For example, based on any local feature of the first object in the infrared image, and combined with the following formula, the object (A), background (B), and uncertain region (A) of the first object under any local feature can be calculated. The probability of ).

[0119]

[0120] in, , , Let be the probability that the first object is an object under any local feature. Let be the probability that the first object is the background under any local feature. Let be the probability that the first object is in an uncertain region under any local feature. The first object is any local feature based on the infrared image.

[0121] The above formula is used to determine the probabilities of the first object as an object, background, and uncertain region under local entropy, which are then used as a set of classification probabilities for the first object.

[0122] Similarly, the probabilities of the first object being an object, background, and uncertain region under local contrast are determined by the above formula, which serve as another set of classification probabilities for the first object.

[0123] The above formula is used to determine the probabilities of the first object as object, background, and uncertain region under the local gradient, which are then used as another set of classification probabilities for the first object.

[0124] The above formula is used to determine the probability that the first object is an object, background, or uncertain region under local texture, which serves as another set of classification probabilities for the first object.

[0125] Thus, the four classification probabilities of the first object were finally obtained.

[0126] Then, based on the four sets of classification probabilities of the first object, the probabilities of the first object being object, background, and uncertain region are determined.

[0127] For example, firstly, the local entropy of the first object is used as... The local contrast of the first object is used as Calculate the conflict factor K. K can be calculated based on the probability that the first object is an object under local entropy and local contrast, and the probability that the first object is background under local entropy and local contrast, respectively.

[0128] For example, the formula for calculating K is as follows:

[0129]

[0130] in, Let be the probability that the first object is an object under the local entropy. Let be the probability that the first object is the background under local entropy. Let be the probability that the first object is an object under local contrast. This represents the probability that the first object is the background under local contrast.

[0131] Then, based on K, the probabilities of the first object as object, background, and uncertain region under local entropy, and the probabilities of the first object as object, background, and uncertain region under local contrast, probability fusion is performed to obtain the fused probabilities. For example, fusion is performed using the following formula:

[0132]

[0133] in, , and These represent the probabilities that the first object after fusion is the target, the background, and the uncertain object, respectively. Let be the probability that the first object is in an uncertain region under local entropy. Let be the probability that the first object is an uncertain region under local contrast.

[0134] Next, , and Reset to The local gradient of the first object is used as Repeat the above process to obtain the fused result.

[0135] Finally, the result of the above re-fusion is reset to The local texture of the first object is used as Repeat the above process and determine the fused result as the final result, that is, obtain the probabilities of the first object, the background and the uncertain region respectively.

[0136] 404. For the second object, based on multiple local features of the second object in the visible light image, determine the probabilities of the second object being an object, background, and uncertain region.

[0137] For example, the implementation process can be as follows:

[0138] First, based on the local entropy, local contrast, local gradient, and local texture of the second object in the visible light image, four sets of classification probabilities for the second object are determined.

[0139] The four classification probabilities of the second object correspond one-to-one with the four local features. Each of the four classification probabilities of the second object includes the probability that the second object is an object, background, or an uncertain region under the corresponding local feature.

[0140] It should be noted that the implementation process of this step can be found in the relevant content of step 403, the difference being that the parameters corresponding to the infrared image in the formula are replaced with the parameters corresponding to the visible light image.

[0141] Then, based on the four sets of classification probabilities of the second object, the probabilities of the second object being object, background, and uncertain region are determined.

[0142] It should be noted that the implementation process of this step can be found in the relevant content of step 403.

[0143] 405. For the third object, based on multiple local features of the third object in the infrared image and the visible light image respectively, determine the probability that the third object is an object, background and uncertain region respectively.

[0144] For example, such as Figure 5 As shown, the implementation process of this step can be described as follows:

[0145] 501. Based on the local entropy, local contrast, local gradient, and local texture of the third object in the infrared image, determine the four sets of classification probabilities of the third object in the infrared image.

[0146] Among them, the four classification probabilities of the third object based on the infrared image correspond one-to-one with the four local features.

[0147] Each of the four classification probabilities for the third object based on the infrared image includes the probability that the third object is an object, background, or uncertain region based on the infrared image and the corresponding local features.

[0148] 502. Based on the four sets of classification probabilities of the third object based on the infrared image, determine the first probability of the third object as object, background and uncertain region.

[0149] It should be noted that the specific implementation process of 501 and 502 can be found in the relevant content of 403, and will not be repeated here.

[0150] 503. Based on the local entropy, local contrast, local gradient, and local texture of the third object in the visible light image, determine the four sets of classification probabilities of the third object in the visible light image.

[0151] Among them, the four sets of classification probabilities of the third object based on the visible light image correspond one-to-one with the four local features; each set of classification probabilities of the third object based on the visible light image includes the probability that the third object is an object, background and uncertain region respectively under the corresponding local features.

[0152] 504. Based on the four sets of classification probabilities of the third object in the visible light image, determine the second probability of the third object as object, background and uncertain region.

[0153] It should be noted that the specific implementation process of 503 and 504 can be found in the relevant content of 404, and will not be repeated here.

[0154] 505. Based on the first and second probabilities that the third object is the object, the background, and the uncertain region, determine the probability that the third object is the object, the background, and the uncertain region.

[0155] For example, the first probability of the third object being the object, the background, and the uncertain region is used as... The relevant parameters are used to determine the second probability of the third object, which is the object, the background, and the uncertain region. The relevant parameters are used to execute the above formulas (13) to (16) to determine the probability of the third object, which is the object, background and uncertain region respectively.

[0156] 406. Output the detection results based on the probabilities of the first object, the second object, and the third object as the object, the background, and the uncertain region, respectively.

[0157] The background is the area outside the detection area in both the infrared and visible light images.

[0158] The specific implementation process can be described as follows:

[0159] The object with a higher probability of being an object than the background and the uncertain region is selected from the first, second, and third objects and designated as the fourth object.

[0160] If the fourth object is an associated object, output the position and category of the fourth object based on its position and category in the infrared and visible light images.

[0161] For example, the position of the fourth object is obtained by averaging the position of the detection box in the infrared image and the position of the detection box in the visible light image.

[0162] If the fourth object is an unassociated object, output the position and category of the fourth object based on its position and category in the infrared or visible light image.

[0163] It should be noted that after identifying the fourth object, the location and category of the fourth object are output, and the detection boxes and categories of all objects in the first, second, and third objects, except for the fourth object, are deleted. Clearly, further filtering based on probability to identify the fourth object further ensures the accuracy and reliability of the detection.

[0164] In some possible embodiments, after outputting the location and category of the fourth object, the system can also locate the corresponding intruder based on the location and category of the fourth object. Furthermore, it can determine whether the intruder will cause a safety incident; for example, in the rail transit field, it can determine whether the intruder will affect the safe operation of the train. If it will cause a safety incident, information is sent to the dispatch center and maintenance center for safety inspection and maintenance. If it will not cause a safety incident, feedback can be given to the maintenance center so that safety inspection and maintenance can be carried out after the emergency request is processed.

[0165] As can be seen from the above, infrared images have the advantage of penetrating in complex environments such as low light and smoke, while visible light images have advantages in terms of detail, resolution, and color information. Therefore, compared with the method of target detection using only visible light images, the embodiments of this application can improve the accuracy of target detection and reduce the impact of complex environments on the detection results by detecting objects in the detection areas of the infrared image and the detection areas of the visible light image of the first region.

[0166] Furthermore, since the detection is performed on objects within the detection area, users can set the area to be detected according to their needs, and detection is only performed within the detection area, which improves detection efficiency and reduces detection energy consumption.

[0167] Furthermore, by associating the same object in the first and second detection results, and based on multiple local features of each object in the association results, a fusion verification is performed on each object in the association results to output the detection result. In other words, by associating the detection results of infrared and visible light images, the complementary information of different image modalities is fully utilized, enabling more comprehensive and accurate object detection. Moreover, performing fusion verification based on multiple local features on each object in the association results avoids misjudgments that may occur due to single-feature judgments, greatly improving the accuracy of verification and making the output detection results more reliable and precise.

[0168] like Figure 6 As shown, the training process of the detection model may include the following steps:

[0169] 601. Obtain multiple visible light training samples and multiple infrared training samples.

[0170] Multiple infrared training samples can be obtained through an infrared camera, and multiple visible light training samples can be obtained through a visible light camera.

[0171] It should be noted that the visible light camera and the infrared camera are installed in the application scenario of this target detection method to ensure that the obtained visible light training samples and infrared training samples are consistent with the application scenario of the target detection method.

[0172] For example, in a rail transit scenario, a visible light camera and an infrared camera can be installed above the area where the track is located. The visible light camera and the infrared camera can be used to obtain infrared training samples and visible light training samples of the object being detected under different times (daytime, nighttime, etc.) and different weather conditions (rain, snow, and sunny days, etc.) to improve the diversity of the samples.

[0173] 602. Perform background processing on the areas outside the detection area in each visible light training sample and each infrared training sample respectively.

[0174] Background processing can be done manually or by equipment. The methods for using equipment include the following:

[0175] The network model identifies the detection region in each visible light training sample and each infrared training sample. The regions outside the detection region in each visible light training sample and each infrared training sample are then filled with background material, for example, by changing it to white, black, or blurring it.

[0176] This network model employs edge recognition of the detection region based on Unet semantic segmentation. Its basic idea is to classify each pixel in the image. The network structure consists of an encoder and an inverse decoder. The encoder is responsible for feature extraction and downsampling of the input image. In this stage, the input image size is reduced by a factor of 16 through four rounds of max pooling. During decoding, the feature map size is restored to its original size through four rounds of bilinear upsampling (2x). Feature maps of the same size are connected across layers. Finally, the classification result for each pixel is output, filling the scene outside the detection region edge as the background.

[0177] 603. Label the location and category of objects in each visible light training sample and each infrared training sample after background processing.

[0178] Because of the background processing, the amount of labeling required for object position and category is reduced, thus improving labeling efficiency.

[0179] 604. Amplify the visible light training samples and infrared training samples belonging to objects with a small category proportion.

[0180] For the visible light training samples, calculate the proportion of objects in each category. If the proportion is small, expand the visible light training samples to which that object belongs.

[0181] Specifically, visible light training samples can be augmented using image coordinate transformation, which is achieved by the following formula:

[0182]

[0183] in, These are the original pixel coordinates. These are the transformed coordinates. By changing the form of T, operations such as rotation, translation, and scaling of the image can be achieved, ultimately expanding the dataset.

[0184] The amplification method for infrared training samples is the same as above, and will not be repeated here.

[0185] By expanding the sample size, the diversity of the sample can be further improved, thereby increasing the accuracy of the detection model training.

[0186] It should be noted that the training samples amplified here are training samples after background processing.

[0187] 605. The neural network is trained using labeled and amplified visible light training samples and infrared training samples to obtain a detection model.

[0188] Neural networks, for example, can employ object detection based on YOLOv10n. The YOLOv10n algorithm consists of four parts: Input, Backbone, Neck, and Head. The Input part handles various operations on the input image to ensure it meets the model's input requirements. The Backbone is the core of feature extraction in YOLOv10n. It extracts rich feature information from the input image through a series of complex convolutional and pooling layers, enabling the model to better capture key information. The Neck effectively integrates feature information from different levels. While retaining the classic structures of FPN and PANet, the Neck cleverly integrates shallow and deep features. The Head employs an innovative decoupled head structure, separating the regression and prediction branches, making the model more efficient and accurate in object detection tasks.

[0189] The labeled and amplified visible light training samples and infrared training samples are input into the neural network, and the neural network is trained by adjusting the weights of the parameters in the neural network, so as to determine the trained neural network as the detection model.

[0190] like Figure 7 As shown, an electronic device 700 provided in this embodiment of the invention may include a processor 710 and a memory 720; the memory 720 is used to store a computer program; the processor 710 is used to implement the target detection method as described above when the computer program is executed.

[0191] This invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the target detection method described above.

[0192] The present invention will now be described an electronic device 700 that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. Electronic device 700 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 700 can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0193] Electronic device 700 includes a computing unit that can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) or loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The computing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0194] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by hardware related to computer program instructions. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0195] The above embodiments are merely illustrative of the technical solutions of the present invention and are not intended to limit it. Those skilled in the art, within the scope of the technical essence of the present invention, can modify the technical solutions described in the foregoing embodiments or replace technical features, and such modifications or replacements still fall within the protection scope of the present invention.

Claims

1. A target detection method, characterized in that, include: Obtain infrared and visible light images of the first region; The objects in the detection area of ​​the infrared image and the detection area of ​​the visible light image are detected respectively to obtain a first detection result and a second detection result; Wherein, the first detection result indicates an object detected from the infrared image, and the second detection result indicates an object detected from the visible light image; Associate the same object in the first detection result and the second detection result to obtain the association result; The association results at least indicate the associated objects and / or unassociated objects in the first detection results and the second detection results; Based on multiple local features of each object in the association results, a fusion verification is performed on each object in the association results, and a detection result is output based on the fusion verification result; The step of associating the same object in the first detection result and the second detection result to obtain the association result includes: Based on the category and location of each object in the first detection result and the category and location of each object in the second detection result, construct a first matrix with I rows and J columns; Where I indicates the number of objects in the first detection result, J indicates the number of objects in the second detection result, and the element located at the intersection of the i-th row and the j-th column indicates the distance between the object corresponding to the i-th row and the object corresponding to the j-th column; Based on the first matrix, the minimum element in the corresponding row is subtracted from each row element, and the minimum element in the corresponding column element is subtracted from each column element to obtain the second matrix; Perform the covering step: cover the zero elements in the second matrix with the minimum number of straight lines; If the number of lines is equal to the number of rows in the second matrix, the second matrix is ​​determined as the third matrix; If the number of lines is less than the number of rows in the second matrix, subtract the smallest non-zero element of the second matrix from each non-zero element in the second matrix, and add the smallest non-zero element to the zero elements in the second matrix that are covered by line intersections to obtain the fourth matrix. After using the fourth matrix as the second matrix, the process begins from the covering step and continues until the number of lines equals the number of rows in the second matrix. The row and column objects corresponding to the zero element in the row with only one zero element in the third matrix are identified as the associated objects, and the objects in the first and second detection results other than the associated objects are identified as the unassociated objects.

2. The method according to claim 1, characterized in that, The association result includes at least one of a first object, a second object, and a third object, wherein the first object is an unassociated object in the first detection result, the second object is an unassociated object in the second detection result, and the third object is an associated object in both the first and second detection results; The step of fusing and verifying each object in the association result based on multiple local features of each object, and outputting a detection result based on the fusing and verification result, includes: Extract multiple local features of each object from the first detection result from the infrared image; Extract multiple local features of each object from the second detection result from the visible light image; For the first object, based on multiple local features of the first object in the infrared image, the probabilities of the first object being an object, background, and uncertain region are determined; For the second object, based on multiple local features of the second object in the visible light image, the probabilities of the second object being an object, background, and an uncertain region are determined; For the third object, based on multiple local features of the third object in the infrared image and the visible light image respectively, the probabilities of the third object being an object, background and uncertain region are determined; Based on the probabilities of the first object, the second object, and the third object being the object, the background, and the uncertain region, respectively, the detection result is output. The background is the area outside the detection area in the infrared image and the visible light image.

3. The method according to claim 2, characterized in that, The multiple local features include four local features, namely local entropy, local contrast, local gradient, and local texture; The step of determining the probability that the first object is an object, background, and an uncertain region based on multiple local features of the infrared image includes: Based on the local entropy, local contrast, local gradient, and local texture of the first object in the infrared image, four sets of classification probabilities for the first object are determined respectively. Among them, the four sets of classification probabilities of the first object correspond one-to-one with the four local features; Each of the four classification probabilities of the first object includes the probability that the first object is an object, background, and uncertain region under the corresponding local features; Based on the four sets of classification probabilities of the first object, the probabilities of the first object being object, background, and uncertain region are determined.

4. The method according to claim 2, characterized in that, The multiple local features include four local features, namely local entropy, local contrast, local gradient, and local texture; The step of determining the probability that the second object is an object, background, and an uncertain region based on multiple local features of the visible light image includes: Based on the local entropy, local contrast, local gradient, and local texture of the second object in the visible light image, four sets of classification probabilities for the second object are determined respectively. Among them, the four sets of classification probabilities of the second object correspond one-to-one with the four local features; Each of the four classification probabilities of the second object includes the probability that the second object is an object, background, and uncertain region under the corresponding local features; Based on the four sets of classification probabilities of the second object, the probabilities of the second object being object, background, and uncertain region are determined.

5. The method according to claim 2, characterized in that, The multiple local features include four local features, namely local entropy, local contrast, local gradient, and local texture; The step of determining the probability that the third object is an object, background, and uncertain region based on multiple local features of the infrared image and the visible light image respectively includes: Based on the local entropy, local contrast, local gradient, and local texture of the third object in the infrared image, four sets of classification probabilities of the third object based on the infrared image are determined respectively. The third object corresponds one-to-one with the four local features based on the four sets of classification probabilities of the infrared image. Each of the four sets of classification probabilities for the third object based on the infrared image includes the probability that the third object is an object, background, and uncertain region based on the infrared image and under the corresponding local features. Based on the four sets of classification probabilities of the third object based on the infrared image, the first probability of the third object being an object, background, and uncertain region is determined; Based on the local entropy, local contrast, local gradient, and local texture of the third object in the visible light image, four sets of classification probabilities of the third object in the visible light image are determined respectively. The third object corresponds one-to-one with the four local features based on the four sets of classification probabilities of the visible light image; Each of the four sets of classification probabilities for the third object based on the visible light image includes the probability that the third object is an object, background, and uncertain region based on the visible light image and under the corresponding local features. Based on the four sets of classification probabilities of the third object in the visible light image, a second probability is determined for the third object to be an object, background, and an uncertain region. Based on the first probability and the second probability that the third object is an object, a background, and an uncertain region, the probability that the third object is an object, a background, and an uncertain region is determined.

6. The method according to claim 2, characterized in that, The output detection result, based on the probabilities of the first object, the second object, and the third object being the object, the background, and the uncertain region, respectively, includes: The object with a higher probability of being the object than the background and the uncertain region is selected from the first object, the second object, and the third object and designated as the fourth object; If the fourth object is the associated object, the position and category of the fourth object are output according to the position and category of the fourth object in the infrared image and the visible light image; If the fourth object is the unassociated object, the position and category of the fourth object are output according to the position and category of the fourth object in the infrared image or the visible light image.

7. The method according to claim 1, characterized in that, The step of detecting objects in the detection areas of the infrared image and the detection areas of the visible light image respectively to obtain a first detection result and a second detection result includes: The detection area is identified from the infrared image and the visible light image, respectively; The regions located outside the detection area in the infrared image and the visible light image are respectively subjected to background processing; Objects in the infrared image and the visible light image after background processing are detected respectively to obtain the first detection result and the second detection result.

8. The method according to claim 1, characterized in that, The detection of objects in the detection areas of the infrared image and the detection areas of the visible light image respectively includes: The detection model is used to detect objects in the detection areas of the infrared image and the detection areas of the visible light image, respectively.

9. The method according to claim 8, characterized in that, The training process of the detection model includes: Multiple visible light training samples and multiple infrared training samples were obtained; Background processing is performed on the areas outside the detection area in each of the visible light training samples and each of the infrared training samples, respectively; The location and category of objects in each of the visible light training samples and each of the infrared training samples after background processing are labeled; Amplify the visible light training samples and the infrared training samples to which the objects of the category with a smaller proportion belong; The neural network is trained using the labeled and amplified visible light training samples and the infrared training samples to obtain the detection model.

Citation Information

Patent Citations

  • Visible light and infrared light fused target recognition method

    CN111611905A

  • Target detection method and system based on infrared and visible light image fusion

    CN117392496A