Image detection method and device and terminal equipment

Through sample annotation and training for application scenarios, combined with the C2F module and the CBAM module, a lightweight infrared image object detection model is built, which solves the problem that infrared image object detection is difficult to achieve real-time and high accuracy in the prior art, and realizes real-time and high accuracy of infrared image object detection.

CN120219889APending Publication Date: 2025-06-27BEIJING KANKAN INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311796446.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing infrared image object detection methods are difficult to achieve real-time detection, and the lack of targeted data sets in application scenarios, resulting in insufficient detection accuracy and robustness.

Method used

Sample annotation and training for application scenarios are adopted, combined with C2F module and CBAM module, a lightweight target image detection model is built to improve the performance and generalization capabilities of the model.

Benefits of technology

Real-time and high accuracy of infrared image object detection are achieved, the robustness of detection and the generalization ability of the model are improved, and the fast real-time detection needs in application scenarios are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219889A_ABST
    Figure CN120219889A_ABST
Patent Text Reader

Abstract

The invention provides an image detection method and device and terminal equipment, and the method comprises the steps: training a preset image detection training model through a test image set, and obtaining a target image detection model. Each to-be-tested image in the to-be-tested image set is an image of a preset scene, each to-be-tested image comprises a category label and a position label, the category label is used for representing the category of a target test object in the to-be-tested image, and the position label represents the position information of the test object. The preset image detection training model and the target image detection model respectively comprise a feature acquisition module, a feature fusion module and a detection module. Each of the feature acquisition module, the feature fusion module and the detection module comprises three outputs. And inputting the to-be-recognized image sample into the target image detection model for recognition to obtain detection target object information. In addition, the invention also provides an image detection device and terminal equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence image detection, and in particular to an image detection method, device and terminal device. Background Art

[0002] At present, infrared image target detection based on infrared sensors has received extensive attention in the scientific field. The infrared images collected by infrared sensors have strong penetration ability, strong anti-interference ability, high measurement accuracy, and can work day and night. However, most of the current methods are proposed based on general platforms and cannot achieve real-time detection. In some application fields, most of the situations can only be deployed on embedded platforms and have high requirements for real-time performance. Therefore, the lightweight of infrared image target detection has become the main goal at present.

[0003] Currently, target detection on infrared images is mainly carried out based on publicly available data. For infrared image target detection without relevant publicly available data sets, it is necessary to collect infrared images of the application scenario and manually annotate them, as well as model training, etc., so as to improve the accuracy, accuracy and robustness of target detection in specific application scenarios. Summary of the Invention

[0004] The present invention provides an image detection method, device and terminal device, which can achieve real-time performance of targets on infrared images, improve the performance and generalization ability of the model, maximize the algorithm, the accuracy and robustness of target detection, and realize the lightweight of infrared image target detection.

[0005] In a first aspect, an embodiment of the present invention provides an image detection method, and the image detection method includes: Training a preset image detection training model with a test image set to obtain a target image detection model. Each test image in the test image set is an image of a preset scene. Each test image includes a class label and a location label. The class label is used to represent the class of the target test object in the test image, and the location label represents the location information where the test object is located. The preset image detection training model and the target image detection model respectively include a feature acquisition module, a feature fusion module, and a detection module. Among them, the image feature sampling module includes several cascaded image feature acquisition layers for performing image feature acquisition on the input image data. The image features collected by each layer are input to the next image feature acquisition layer as the input image data of the next image feature acquisition layer. The feature acquisition module includes a CBAM module, a CBM module, a C2F module, and an SPPF module. The feature fusion module includes a CBAM module, a CBM module, and a C2F module. The detection module includes a CBM module; and Input the image sample to be recognized into the target image detection model for recognition to obtain the detected target object information. The detected target object includes the type label and position label of the detected target object. Each image to be recognized in the image sample to be recognized is an image of a preset scene, and each image sample to be recognized includes the category label and the position label.

[0006] In a second aspect, an embodiment of the present invention provides an image detection device, including: A training device that trains a preset image detection training model using a test image set to obtain a target image detection model. Each test image in the test image set is an image of a preset scene, and each test image includes a category label and a position label. The category label is used to represent the category of the target test object in the test image, and the position label represents the position information where the test object is located. The preset image detection training model and the target image detection model respectively include a feature acquisition module, a feature fusion module, and a detection module. The feature acquisition module, the feature fusion module, and the detection module all have three outputs. Among them, the image feature sampling module includes several cascaded image feature acquisition layers for performing image feature acquisition on the input image data. The image features collected by each layer are input to the next image feature acquisition layer as the input image data of the next image feature acquisition layer. The feature acquisition module includes a CBAM module, a CBM module, a C2F module, and an SPPF module. The feature fusion module includes a CBAM module and a CBM module. The detection module includes a CBM module; and A recognition device that inputs the image sample to be recognized into the target image detection model for recognition to obtain the detected target object information. The detected target object includes the type label and position label of the detected target object. Each image to be recognized in the image sample to be recognized is an image of a preset scene, and each image sample to be recognized includes the category label and the position label.

[0007] In a third aspect, an embodiment of the present invention provides a terminal device, including: A memory for storing computer-executable programs; A processor for executing the computer-executable program to implement the image detection method according to any one of claims 1 to 8.

[0008] The above-mentioned image detection method, device, and terminal device mainly improve the detection accuracy and model robustness by using sample annotation and training for the application scenario. By adopting the C2F module and the CBAM module, since the C2F module deletes two upsampling convolutions and the CBAM module has fewer internal convolution structures and pooling layers, this structure reduces the large amount of calculations brought by convolution multiplication, reducing the module complexity and computational load. Therefore, using the C2F module and the CBAM module on lightweight models can improve performance and reduce the corresponding computational load, which is beneficial to achieving the lightweight and real-time detection of the model. By adopting the trained target image detection model, the risk of overfitting of the model during training is reduced, thereby reducing the probability of model training failure, improving the training efficiency and the generalization of the model. Thus, it meets the requirements of current infrared image target detection accuracy, model lightweight, and fast real-time target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0010] Figure 1 It is a schematic flowchart of an image detection method provided by an embodiment of the present application.

[0011] Figure 2 It is a schematic diagram of an image detection device provided by an embodiment of the present application.

[0012] Figure 3 It is a schematic diagram of the detection model processing process of an image detection device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the following further details the present invention in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0014] In the description, claims, and the above-mentioned drawings of this application, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar planning objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances. In other words, the described embodiments are implemented in an order other than that illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof may also include other content. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to only those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0015] It should be noted that in the present invention, the descriptions involving "first", "second", etc. are only for descriptive purposes, and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. Additionally, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0016] Please refer to Figures 1 to 3 , Figure 1 which is a schematic flowchart of an image detection method provided for an embodiment of this application. Figure 2 which is a schematic diagram of an image detection device provided for an embodiment of this application. Figure 3 which is a schematic diagram of the detection model processing process of an image detection device provided for an embodiment of this application. An embodiment of the present invention discloses an image detection device. The image detection device 2 includes a training device 21 and an identification device 22. The training device 21 trains a preset image detection training model using a test image set to obtain a target image detection model. The preset image detection training model and the target image detection model respectively include a feature acquisition module, a feature fusion module, and a detection module. The identification device 22 inputs a to-be-tested image sample into the target image detection model for identification to obtain detection target object information. The to-be-tested image 1 is an image of a preset scenario. The to-be-tested image 1 includes a category label 11 and a position label 12. The category label 11 is used to represent the category of the target test object in the to-be-tested image 1, and the position label 12 represents the position information where the test object is located.

[0017] The training device 21 includes a feature acquisition module 211, a feature fusion module 212, and a detection module 213. The feature acquisition module 211, the feature fusion module 212, and the detection module 213 all have three outputs. Among them, the image feature sampling module 211 includes several cascaded image feature acquisition layers for acquiring image features from the input image data. The image features acquired by each layer are input into the next image feature acquisition layer as the input image data of the next image feature acquisition layer. The feature acquisition module 211 includes a CBAM module 2111, a CBM module 2112, a C2F module 2113, and an SPPF module 2114. The feature fusion module 212 includes a CBAM module 2121, a CBM module 2122, and a C2F module 2123. The detection module 231 includes a CBM module 2131.

[0018] The recognition device 22 inputs the image sample to be tested into the target image detection model for recognition to obtain the detection target object information. Among them, each image to be tested in the image sample to be tested is an image of a preset scene, and each image sample 1 to be tested includes a category label 11 and a location label 12. Therefore, after being detected and recognized by the target detection model, the detection target object also includes the type label and location label of the detection target object.

[0019] The feature acquisition module 211 includes a CBAM module 2111, a CBM module 2112, a C2F module 2113, and an SPPF module 2114. The feature acquisition module 211 has a total of fifteen layers. The CBM module 2112 is located in the first, second, fifth, eighth, eleventh, and twelfth image feature acquisition layers. The C2F module 2113 is located in the third, sixth, and ninth image feature acquisition layers. The CBAM module 2111 is located in the fourth, seventh, tenth, and fourteenth image feature acquisition layers. The SPPF module 2114 is located in the fifteenth image acquisition layer. The C2F module 2113 is arranged in multiple image feature acquisition layers to obtain C2F convolutional layers. The CBAM module 2111 and the CBM module 2112 are arranged between the C2F convolutional layers. The SPPF module 2114 is arranged in the last image feature acquisition layer to obtain the SPPF convolutional layer. The feature acquisition module 211 outputs three extracted features. Among them, two of the extracted feature outputs are output from two C2F convolutional layers respectively, and the other extracted feature output is output from the SPPF convolutional layer.

[0020] Specifically, the processing process of the feature acquisition module 211 is as follows: First, an infrared image sample to be tested is input. This image sample 1 to be tested includes a labeled class label 11 and a position label 12. The input image passes through the CBM module 2112, the CBM module 2112, the C2F module 2113, the CBAM module 2111, and the CBM module 2112 in sequence and reaches the C2F convolutional layer of the sixth layer of the feature acquisition module, and the extracted feature F1 is output. Then, it passes through the CBAM module 2111 and the CBM module 2112 and reaches the C2F convolutional layer of the ninth layer of the feature acquisition module, and the extracted feature F2 is output. Finally, it passes through the CBAM module 2111, the CBM module 2112, the CBM module 2112, the C2F module 2113, and the CBAM module 2111 and reaches the SPPF convolutional layer of the fifteenth layer of the feature acquisition module, and the extracted feature F3 is output.

[0021] The CBAM module 2111 has fewer convolutional structures and pooling layers inside the module, which avoids a large amount of calculations brought by convolutional multiplication and reduces the module complexity. It also further realizes the strong quantization of the model. The CBAM module has strong versatility and high portability, and the acceleration algorithm is implemented on an embedded platform. The CBAM module starts from two scopes of spatial attention and channel attention. Among them, spatial attention pays more attention to the pixel regions that play a decisive role in classification in the image and ignores other regions, while channel attention is used to process the allocation relationship of the feature map channels. At the same time, the attention allocation for both dimensions enhances the improvement effect of the attention mechanism on the model performance.

[0022] The CBM module 2112 consists of a Conv convolutional layer, a BN layer, and a sigmoid activation function layer, which can extract features from data, enable the training parameters to propagate forward better and faster, and increase the non-linearity of the neural network model.

[0023] The C2F module 2113 makes the gradient flow of features richer, deletes the number of blocks in the largest stage of the backbone network, thus greatly reducing the number of convolutional parameters and the amount of calculation, and realizing the lightweight of the model.

[0024] The SPPF module 2114 can effectively avoid problems such as image distortion caused by operations such as cropping and scaling of image regions, and solve the problem of repeated feature extraction of the convolutional neural network for images. The SPPF layer captures the multi-scale information of the input image by using pooling kernels of different sizes and reduces the number of channels, which helps to improve the receptive field and feature expression ability of the model.

[0025] The feature fusion module 212 includes a CBAM module 2121, a CBM module 2122, and a C2F module 2123. Among them, the CBM module 2122 includes a CBM module 2122a and a CBM module 2122b; the C2F module 2123 includes a C2F module 2123a, a C2F module 2123b, a C2F module 2123c, and a C2F module 2123d. The feature fusion module 212 has three layers. The first layer includes an upsampling module, the CBAM module 2121, and a fusion module. The second layer includes the C2F module 2123a, the CBAM module 2121, the upsampling module, and a fusion module. The third layer includes the C2F module 2123b, the CBAM module 2121, the CBM module 2122a, the C2F module 2123c, the CBM module 2122b, the C2F module 2123d, and a fusion module. The three enhanced feature outputs of the feature fusion module are output by the C2F module, and the corresponding C2F modules of the three enhanced feature outputs are respectively located in the third layer of the feature fusion layer.

[0026] Specifically, the inputs of the feature fusion module are the extracted features F1, F2, and F3 extracted from the sixth, ninth, and fifteenth layers of the feature acquisition module respectively. The processing process of feature fusion is that the extracted feature F3 passes through the upsampling module and the CBAM module 2121 in sequence and then fuses with the extracted feature F2 to obtain the enhanced feature F6. The enhanced feature F6 passes through the C2F module 2123a to obtain the enhanced feature F5. The enhanced feature F6 passes through the C2F module 2123a, the upsampling module, and the CBAM module 2121 in sequence and then fuses with the extracted feature F1 to obtain the enhanced feature F7. The enhanced feature F7 passes through the C2F module 2123b to obtain the enhanced feature F4. The enhanced feature F7 passes through the C2F module 2123b, the CBM module 2122a, and the CBAM module 2121 in sequence and then fuses with the enhanced feature F5 to obtain the enhanced feature F11. The enhanced feature F11 passes through the C2F module 2123c to obtain the enhanced feature F10. The enhanced feature F11 passes through the C2F module 2123b, the CBM module 2122b, and the CBAM module 2121 in sequence and then fuses with the extracted feature F3 to obtain the enhanced feature F8. The enhanced feature F8 passes through the C2F module 2123d to obtain the enhanced feature F9. Finally, the enhanced feature F4, the enhanced feature F9, and the enhanced feature F10 are used as the outputs of the feature fusion module. The CBM module 2121 is composed of a Conv convolution, a BN layer, and a sigmoid activation function in sequence. By using the feature fusion module, the accuracy of classification can be improved, the robustness of the model can be enhanced, and the risk of overfitting of the model during training can be reduced by fusing different feature data. Feature fusion can help the model better understand the data, thereby improving the performance and generalization ability of the model.

[0027] The detection module 213 includes the CBM module 2131. The detection module 213 includes three identically structured CBM modules 2131 located in the first layer of the detection module, a convolution module in the second layer of the detection module, and a class loss module and a location loss module in the third layer of the detection module.

[0028] Specifically, the inputs to the detection module 213 are the enhanced features F4, F9, and F10 extracted from the third layer of the feature fusion module 211 respectively. The detection module includes three identical decoupled heads, and the enhanced features F4, F9, and F10 are respectively input into the three decoupled heads. When the enhanced features F4, F9, and F10 are respectively input into the three identical decoupled heads, each decoupled head has two identical branches, and each branch consists of a CBM module 2131 and a common convolution (Conv). One branch is responsible for class prediction, and the other branch is responsible for detection box prediction. The class loss uses the cross-entropy loss function, and the regression of the bounding box uses the DFL Loss + CIOU Loss method. Among them, the CBM module 2131 can extract features in the image, normalize the input data, and increase the non-linearity of the network, thereby improving the expression ability and training ability of the network.

[0029] The recognition device 22 inputs the image sample to be tested into the target image detection model for recognition to obtain the detection target object information. Among them, each image to be tested in the image sample to be tested is an image of a preset scene, and each image sample 1 to be tested includes a class label 11 and a location label 12. Therefore, after being detected and recognized by the target detection model, the detection target object also includes the type label and location label of the detection target object. Before training the target image detection model, the detection effect of the preset image detection training model has problems such as the detection position standard and the missed detection of vehicles and people in the detection of infrared images. After using the trained target image detection training model, vehicles and cyclists can be accurately detected.

[0030] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an image detection method. The image detection method includes the following steps S101 - S103.

[0031] Step S101: Use a test image set to train a preset image detection training model to obtain a target image detection model 2. Each test image 1 in the test image set is an image of a preset scenario. Each test image 1 includes a class label 11 and a location label 12. The class label 11 is used to represent the class of the target test object in the test image, and the location label 12 represents the location information of the test object. The preset image detection training model and the target image detection model 2 each include a feature acquisition module 211, a feature fusion module 212, and a detection module 213. The feature acquisition module 211, the feature fusion module 212, and the detection module 213 each have three outputs. Among them, the image feature sampling module includes several cascaded image feature acquisition layers for performing image feature acquisition on the input image data. The image features collected by each layer are input to the next image feature acquisition layer as the input image data of the next image feature acquisition layer. The depth of the feature acquisition module is fifteen layers, the feature fusion module includes three feature fusion layers, and the detection module includes three detection layers. The feature acquisition module 211 includes a CBAM module 2111, a CBM module 2112, a C2F module 2113, and an SPPF module 2114; the feature fusion module 212 includes a CBAM module 2121, a CBM module 2122, and a C2F module 2123; the detection module 213 includes a CBM module 2131.

[0032] Understandably, each test image 1 in the test image set is an image of a preset scenario. Each test image includes a class label 11 and a location label 12. The class label 11 is used to represent the class of the target test object in the test image, which can distinguish different types of objects, such as people, vehicles, and other obstacles; the location label 12 represents the location information of the test object, which can accurately describe the location where the test object is located.

[0033] Since the current target detection on infrared images is mainly trained and applied based on publicly available data, and there is no relevant publicly available dataset for current infrared image target detection, infrared image samples need to be collected from real scenes through infrared cameras. The collected infrared image samples are manually labeled for the application scenario, and the labeled samples are divided into a training set and a test set according to a set ratio. Each infrared sample image contains a class label and a location label, which can effectively identify the type and location information of the object to be tested. The above method for receiving infrared image samples realizes the use of sample collection for the application scenario, greatly improving the detection rate and accuracy in the application scenario.

[0034] The feature acquisition module 211 includes a CBAM module 2111, a CBM module 2112, a C2F module 2113, and an SPPF module 2114. The feature acquisition module 211 has a total of fifteen layers. The CBM module 2112 is located in the first, second, fifth, eighth, eleventh, and twelfth image feature acquisition layers. The C2F module 2113 is located in the third, sixth, and ninth image feature acquisition layers. The CBAM module 2111 is located in the fourth, seventh, tenth, and fourteenth image feature acquisition layers. The SPPF module 2114 is located in the fifteenth image acquisition layer. The C2F module 2113 is arranged in multiple image feature acquisition layers to obtain C2F convolutional layers. The CBAM module 2111 and the CBM module 2112 are arranged between the C2F convolutional layers. The SPPF module 2114 is arranged in the last image feature acquisition layer to obtain an SPPF convolutional layer. The feature acquisition module 211 outputs three extracted features. Among them, two extracted feature outputs are respectively output from two C2F convolutional layers, and the other extracted feature output is output from the SPPF convolutional layer.

[0035] Specifically, the processing process of the feature acquisition module 211 is as follows: First, an infrared image sample to be tested is input. This image sample 1 to be tested includes a labeled class label 11 and a position label 12. The input image sequentially passes through the CBM module 2112, the CBM module 2112, the C2F module 2113, the CBAM module 2111, and the CBM module 2112 to reach the C2F convolutional layer of the sixth layer of the feature acquisition module, and the extracted feature F1 is output. Then, it passes through the CBAM module 2111 and the CBM module 2112 to reach the C2F convolutional layer of the ninth layer of the feature acquisition module, and the extracted feature F2 is output. Finally, it passes through the CBAM module 2111, the CBM module 2112, the CBM module 2112, the C2F module 2113, and the CBAM module 2111 to reach the SPPF convolutional layer of the fifteenth layer of the feature acquisition module, and the extracted feature F3 is output.

[0036] The CBAM module 2111 has fewer convolutional structures and pooling layers inside the module. This avoids a large amount of calculations brought by convolutional multiplication and reduces the module complexity. It also further realizes the strong quantization of the model. The CBAM module has strong versatility and high portability, and the acceleration algorithm is implemented on an embedded platform. The CBAM module starts from two scopes of spatial attention and channel attention. Among them, spatial attention pays more attention to the pixel regions that play a decisive role in classification in the image while ignoring other regions, and channel attention is used to process the allocation relationship of the feature map channels. At the same time, the attention allocation in both dimensions enhances the improvement effect of the attention mechanism on the model performance.

[0037] The CBM module 2112 consists of a Conv convolutional layer, a BN layer, and a sigmoid activation function layer, which can extract features from data, enable the training parameters to propagate forward better and faster, and increase the non-linearity of the neural network model.

[0038] The C2F module 2113 makes the gradient flow of features richer, reduces the number of blocks in the largest stage of the backbone network, thus greatly reducing the number of convolutional parameters and the computational amount, and realizing the lightweight of the model.

[0039] The SPPF module 2114 can effectively avoid problems such as image distortion caused by operations such as cropping and scaling of image regions, solve the problem of repeated feature extraction of the convolutional neural network for images. The SPPF layer captures multi-scale information of the input image by using pooling kernels of different sizes and reduces the number of channels, which helps to improve the receptive field and feature expression ability of the model.

[0040] The feature fusion module 212 includes a CBAM module 2121, a CBM module 2122, and a C2F module 2123; among them, the CBM module 2122 includes a CBM module 2122a and a CBM module 2122b; the C2F module 2123 includes a C2F module 2123a, a C2F module 2123b, a C2F module 2123c, and a C2F module 2123d. The feature fusion module 212 has three layers. The first layer includes an upsampling module, a CBAM module 2121, and a fusion module. The second layer includes a C2F module 2123a, a CBAM module 2121, an upsampling module, and a fusion module. The third layer includes a C2F module 2123b, a CBAM module 2121, a CBM module 2122a, a C2F module 2123c, a CBM module 2122b, a C2F module 2123d, and a fusion module. The three enhanced feature outputs of the feature fusion module are output by the C2F module, and the C2F modules corresponding to the three enhanced feature outputs are respectively located in the third layer of the feature fusion layer.

[0041] Specifically, the inputs of the feature fusion module are the extracted features F1, F2, and F3 extracted from the sixth, ninth, and fifteenth layers of the feature acquisition module respectively. The process of feature fusion is as follows: the extracted feature F3 passes through the upsampling module and the CBAM module 2121 in sequence and then fuses with the extracted feature F2 to obtain the enhanced feature F6. The enhanced feature F6 passes through the C2F module 2123a to obtain the enhanced feature F5. The enhanced feature F6 passes through the C2F module 2123a, the upsampling module, and the CBAM module 2121 in sequence and then fuses with the extracted feature F1 to obtain the enhanced feature F7. The enhanced feature F7 passes through the C2F module 2123b to obtain the enhanced feature F4. The enhanced feature F7 passes through the C2F module 2123b, the CBM module 2122a, and the CBAM module 2121 in sequence and then fuses with the enhanced feature F5 to obtain the enhanced feature F11. The enhanced feature F11 passes through the C2F module 2123c to obtain the enhanced feature F10. The enhanced feature F11 passes through the C2F module 2123c, the CBM module 2122b, and the CBAM module 2121 in sequence and then fuses with the extracted feature F3 to obtain the enhanced feature F8. The enhanced feature F8 passes through the C2F module 2123d to obtain the enhanced feature F9. Finally, the enhanced feature F4, the enhanced feature F9, and the enhanced feature F10 are used as the outputs of the feature fusion module. The CBM module 2121 consists of a Conv convolution, a BN layer, and a sigmoid activation function in sequence. By using the feature fusion module, the classification accuracy can be improved, the robustness of the model can be enhanced, and the risk of overfitting of the model during training can be reduced by fusing different feature data. Feature fusion can help the model better understand the data, thereby improving the performance and generalization ability of the model.

[0042] The detection module 213 includes a CBM module 2131. The detection module 213 includes three CBM modules 2131 with the same structure located in the first layer of the detection module, a convolution module located in the second layer of the detection module, and a class loss module and a location loss module located in the third layer of the detection module.

[0043] Specifically, the inputs of the detection module are the enhanced features F4, F9, and F10 extracted from the third layer of the feature fusion module respectively. The detection module includes three identical decoupled heads, and the enhanced features F4, F9, and F10 are respectively input into the three decoupled heads. The enhanced features F4, F9, and F10 are respectively input into the three identical decoupled heads. Each decoupled head has two identical branches, and each branch consists of a CBM module 2131 and a common convolution (Conv). One branch is responsible for class prediction, and the other branch is responsible for detection box prediction. The category loss uses the cross-entropy loss function, and the regression of the bounding box uses the DFL Loss + CIOU Loss method. Among them, the CBM module 2131 can extract features in the image, normalize the input data, and increase the non-linearity of the network, thereby improving the expression ability and training ability of the network.

[0044] Use the above-mentioned labeled samples to train a preset image detection training model, set training parameters according to the actual needs of the project, and then use the same parameters as those during training for model testing after completing model training. Repeat the training and testing processes to achieve the upgrade and iteration of the model. Finally, train an improved target image detection model that meets the actual application scenario. Use the trained target image detection model to achieve the lightweight of the model, and realize the real-time detection of infrared targets on the application platform, improving the accuracy, robustness, and generalization of the detection. For target detection on infrared images, infrared light images have the characteristics of strong penetration and little influence from the external environment, which can ensure the accuracy and stability of the collected images.

[0045] Step S103, input the image sample to be recognized into the target image detection model for recognition to obtain the detection target object information. It can be understood that the recognition device 22 inputs the image sample to be tested into the target image detection model for recognition to obtain the detection target object information. Among them, each image to be tested in the image sample to be tested is an image of a preset scenario, and each image sample 1 includes a category label 11 and a position label 12. Therefore, after being detected and recognized by the target detection model, the detection target object also includes the type label and position label of the detection target object. Before training the target image detection model, the detection effect of the preset image detection training model has problems with the detection position standard and the missed detection of vehicles and people in the detection of infrared images. After using the trained target image detection training model, vehicles and cyclists can be accurately detected.

[0046] Specifically, there is an infrared image sample to be recognized, in which both vehicles and cyclists appear simultaneously. The preset image detection training model can only detect the vehicles and their positions, but the specific positions of the vehicles are not very accurate and there will be certain deviations, which cannot well reflect the positional relationship between the vehicles and the people. However, after using the trained target image detection model, the image sample to be recognized is input into the target image detection model for recognition. Since the image to be recognized contains class labels and position labels and the detection module in the model has a class loss module and a position loss module, the vehicles and the cyclists can be accurately recognized, and the positions of the vehicles, the cyclists, and their previous positional relationship can be accurately read out. Target detection is carried out on the infrared light image. The infrared image collected by the infrared sensor has strong penetration ability, a relatively long visible distance, little influence from the external weather, strong anti-interference ability, and characteristics such as high measurement accuracy and the ability to work day and night. Therefore, target detection on the infrared image can greatly improve the detection accuracy and reduce the influence of the external environment, ensuring the accuracy and reliability of the target.

[0047] The above-mentioned image detection method mainly improves the detection accuracy and the robustness of the model by using sample annotation and training for the application scenario. By adopting the C2F module and the CBAM module, since the C2F module deletes two upsampling convolutions, and the CBAM module has fewer internal convolution structures, fewer pooling layers, and operations such as feature fusion, this structure reduces the large amount of calculations brought by convolutional multiplication, reducing the module complexity and the amount of calculations. Thus, using the C2F module and the CBAM module on lightweight models can improve performance and reduce the corresponding amount of calculations, which is beneficial to achieving the lightweight and real-time detection of the model. By adopting the trained target image detection model, the risk of overfitting of the model during training is reduced, thereby reducing the probability of model training failure, improving the training efficiency and the generalization ability of the model. Thus, it meets the requirements of the current infrared image target detection accuracy, the lightweight of the model, and the fast real-time detection of the target.

[0048] The embodiment of the present invention also provides a terminal infrared image target detection device, including a memory and a processor. The memory is used to store computer-executable programs, and the processor is used to execute the computer-executable programs to implement the above-mentioned image detection method. At least one control processor and a memory used to communicate with at least one control processor. The memory stores instructions executable by at least one control processor, and the instructions are executed by at least one control processor so that at least one control processor can execute the above-mentioned target detection method for infrared light.

[0049] An embodiment of the present invention also provides a computer-readable storage medium, including computer-executable instructions stored in the computer-readable storage medium, and the computer-executable instructions are used to cause the computer to execute the above infrared light target detection method.

[0050] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, provided that these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application also intends to include these changes and modifications.

[0051] The above are only the preferred embodiments of this application, and of course, the scope of the rights of this application cannot be limited thereby. Therefore, equivalent changes made according to the claims of this application still fall within the scope covered by this application.

Claims

1. An image detection method, characterized in that, The described image detection method includes: Training a preset image detection training model using a test image set to obtain a target image detection model. Each test image in the test image set is an image of a preset scene. Each test image includes a class label and a location label. The class label is used to represent the class of the target test object in the test image, and the location label represents the location information of the test object. The preset image detection training model and the target image detection model each include a feature acquisition module, a feature fusion module, and a detection module. Among them, the image feature sampling module includes several cascaded image feature acquisition layers for acquiring image features from the input image data. The image features acquired by each layer are input to the next image feature acquisition layer as the input image data of the next image feature acquisition layer. The feature acquisition module includes a CBAM module, a CBM module, a C2F module, and an SPPF module. The feature fusion module includes a CBAM module, a CBM module, and a C2F module. The detection module includes a CBM module; and Inputting an image sample to be recognized into the target image detection model for recognition to obtain detection target object information. The detection target object includes a type label and a location label of the detection target object. Each image to be recognized in the image sample to be recognized is an image of a preset scene. Each image sample to be recognized includes the class label and the location label.

2. The image detection method according to claim 1, wherein The C2F module is set in multiple image feature acquisition layers to obtain C2F convolutional layers. A CBAM module and a CBM module are set between the C2F convolutional layers. The SPPF module is set in the last image feature acquisition layer to obtain an SPPF convolutional layer. The feature acquisition module outputs three extracted features. Among them, two extracted feature outputs are respectively output from two C2F convolutional layers, and the other extracted feature output is output from the SPPF convolutional layer.

3. The image detection method according to claim 2, wherein The depth of the feature acquisition module is fifteen layers. The feature fusion module includes three feature fusion layers. The detection module includes three detection layers.

4. The image detection method according to claim 3, wherein The C2F convolutional layers are located in the third, sixth, and ninth image feature acquisition layers.

5. The image detection method according to claim 4, wherein The CBM module is located in the first, second, fifth, eighth, eleventh, and twelfth image feature acquisition layers.

6. The image detection method according to claim 5, wherein The CBAM module is located in the fourth, seventh, tenth, and fourteenth image feature acquisition layers.

7. The image detection method according to claim 4, characterized in that The CBAM module is located in the first layer, second layer, and third layer of the feature fusion layer. The CBM module is located in the third layer of the feature fusion layer. The three enhanced feature outputs of the feature fusion module are output by the C2F module. One C2F module is located in the second layer of the feature fusion layer. The C2F module corresponding to the three enhanced feature outputs is located in the third layer of the feature fusion layer.

8. The image detection method according to claim 4, wherein, The detection module includes three CBM modules with the same structure located in the first layer of the detection module.

9. An image detection device, characterized in that, The image detection device includes: A training device uses a test image set to train a preset image detection training model to obtain a target image detection model. Each test image in the test image set is an image of a preset scene. Each test image includes a class label and a location label. The class label is used to represent the class of the target test object in the test image, and the location label represents the location information of the test object. The preset image detection training model and the target image detection model each include a feature acquisition module, a feature fusion module, and a detection module. The feature acquisition module, the feature fusion module, and the detection module each have three outputs. Among them, the image feature sampling module includes several cascaded image feature acquisition layers for acquiring image features from the input image data. The image features acquired by each layer are input to the next image feature acquisition layer as the input image data of the next image feature acquisition layer. The feature acquisition module includes a CBAM module, a CBM module, a C2F module, and an SPPF module. The feature fusion module includes a CBAM module, a CBM module, and a C2F module. The detection module includes a CBM module; and An identification device inputs an image sample to be identified into the target image detection model for identification to obtain detection target object information. The detection target object includes a type label and a location label of the detection target object. Each image to be identified in the image sample to be identified is an image of a preset scene. Each image sample to be identified includes the class label and the location label.

10. A terminal device, characterized in that, Comprising: A memory for storing computer-executable programs; A processor for executing the computer-executable programs to implement the image detection method according to any one of claims 1 to 8.