Target detection method and device and vehicle

By using a lightweight convolutional neural network with cross-stage local CSP network in the object detection model, the problem of large computing resources occupied by the object detection model on resource-limited devices is solved, and efficient fire and smoke detection in charging station monitoring scenarios is achieved.

CN120298649APending Publication Date: 2025-07-11BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410039538.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When existing object detection models are deployed on hardware devices with limited resources, the computing resources are large, resulting in long-term detection and low efficiency, especially in monitoring scenarios of mobile terminals and charging stations.

Method used

A lightweight convolutional neural network (CSP network) with cross-stage local CSP network is used as the backbone structure, combining the neck and head structures to perform feature extraction of target images and fire and smoke detection, reducing model volume and computing resource requirements.

Benefits of technology

The target detection model is lightly deployed on hardware devices with limited resources, improving the efficiency of fire and smoke detection, and is suitable for monitoring scenarios of charging stations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298649A_ABST
    Figure CN120298649A_ABST
Patent Text Reader

Abstract

The invention provides a target detection method and device and a vehicle. The method comprises: obtaining a target image; the target image is input into a target detection model, the target detection model comprises a backbone structure, a neck structure and a head structure, the backbone structure comprises a CSP network, and the CSP network comprises a lightweight CNN; performing feature extraction on the target image through a lightweight CNN in a CSP network in the backbone structure to obtain a target feature of the target image; and performing fire detection and smoke detection based on the target features through the neck structure and the head structure to obtain a target detection result of the target image. Therefore, the CSP network in the scheme comprises the lightweight CNN, compared with a CSP network in related technologies, the CSP network comprises a residual network, the target detection model is lighter, model deployment is facilitated, feature extraction is performed by using the lightweight CNN so as to perform fire detection and smoke detection, and the target detection efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of object detection, and in particular, to an object detection method, apparatus, electronic device, storage medium, and vehicle. Background Art

[0002] Currently, with the continuous development of artificial intelligence technology, object detection models have been widely used in fields such as autonomous driving and station monitoring, and have advantages such as high automation and high intelligence. For example, object detection models can perform obstacle detection, fire detection, etc.

[0003] In related technologies, most object detection models adopt YOLO (You Only Look Once), SSD (SingleShot MultiBox Detector), Faster R-CNN (RegionConvolutional Neural Networks, an object detection model based on convolutional neural networks), etc., and the accuracy of the object detection models is relatively good. However, the volume of the object detection models and the required computing resources are relatively large, and the storage space and computing resources occupied on hardware devices are also relatively large. Especially in the application scenario where the object detection model is deployed to a hardware device with limited resource capacity, for example, in the application scenario where the object detection model is deployed to a mobile terminal, the computing power of the mobile terminal is insufficient, which may lead to a long time-consuming object detection and low object detection efficiency. Summary of the Invention

[0004] The present disclosure aims to at least partly solve one of the technical problems in the above technologies.

[0005] To this end, the first objective of the present disclosure is to propose an object detection method.

[0006] The second objective of the present disclosure is to propose an object detection apparatus.

[0007] The third objective of the present disclosure is to propose an electronic device.

[0008] The fourth objective of the present disclosure is to propose a computer-readable storage medium.

[0009] The fifth objective of the present disclosure is to propose a vehicle.

[0010] The first aspect embodiment of the present disclosure proposes an object detection method, including: obtaining a target image; inputting the target image into an object detection model, where the object detection model includes a backbone structure, a neck structure, and a head structure, the backbone structure includes a cross-stage local CSP network, and the CSP network includes a lightweight convolutional neural network CNN; extracting features of the target image through the lightweight CNN in the CSP network of the backbone structure to obtain target features of the target image, where the target features include at least one of color features, texture features, shape features, and spatial relationship features of the target image; performing fire detection and smoke detection based on the target features through the neck structure and the head structure to obtain the object detection result of the target image.

[0011] In addition, the object detection method proposed according to the above embodiment of the present disclosure may further have the following additional technical features:

[0012] In an embodiment of the present disclosure, the CSP network further includes a CBL network, a convolutional network, and a splicing network; the extracting features of the target image through the lightweight CNN in the CSP network of the backbone structure to obtain target features of the target image includes: extracting features of the target image through the CBL network to obtain first features of the target image; extracting features of the first features through the lightweight CNN to obtain second features; extracting features of the target image through the convolutional network to obtain third features of the target image; splicing the second features and the third features according to the channel dimension through the splicing network to obtain fourth features; performing downsampling processing on the fourth features through the CBL network to obtain the target features.

[0013] In an embodiment of the present disclosure, multiple lightweight CNNs in the CSP network form a target CNN; the extracting features of the target image through the lightweight CNN in the CSP network of the backbone structure to obtain target features of the target image includes: extracting features of the target image through the target CNN in the CSP network of the backbone structure to obtain the target features.

[0014] In one embodiment of the present disclosure, a plurality of lightweight CNNs in the CSP network are connected in series to form the target CNN; the step of extracting features of the target image through the target CNN in the CSP network of the backbone structure to obtain the target features includes: obtaining the serial order of the plurality of lightweight CNNs in the target CNN, where two adjacent lightweight CNNs in the order are connected in series; extracting features of the target image through the lightweight CNN ranked at the i-th position to obtain the i-th feature, where i is a positive integer; extracting features of the i-th extracted feature through the lightweight CNN ranked at the (i + 1)-th position to obtain the (i + 1)-th feature; and obtaining the feature obtained by the lightweight CNN ranked last in the target CNN as the target feature.

[0015] In one embodiment of the present disclosure, the step of performing fire detection and smoke detection based on the target features through the neck structure and the head structure to obtain the target detection result of the target image includes: performing fire detection and smoke detection based on the target features through the neck structure and the head structure to obtain candidate features of the target image, where the candidate features include the confidence of a set object and the candidate detection box of the set object, and the set object includes fire and smoke; screening out the fifth feature from the plurality of candidate features based on the confidence in the plurality of candidate features; and using the set object and the candidate detection box of the set object in the fifth feature as the target detection result.

[0016] In one embodiment of the present disclosure, the step of performing fire detection and smoke detection based on the target features through the neck structure and the head structure to obtain candidate features of the target image includes: performing sampling processing on the target features at multiple scales through the neck structure to obtain features at multiple scales; and performing fire detection and smoke detection based on the features at multiple scales through the head structure to obtain the candidate features.

[0017] In one embodiment of the present disclosure, the neck structure includes a Feature Pyramid Network (FPN), where the FPN includes the CSP network and the CBL network; the multi-scale sampling process of the target feature by the neck structure to obtain features of multiple scales includes: performing a 1x downsampling process on the target feature of the original scale through the CBL network to obtain a sixth feature of the first scale, where the first scale is half of the original scale; performing feature extraction on the sixth feature of the first scale through the CSP network to obtain a seventh feature of the first scale; performing a 1x downsampling process on the seventh feature of the first scale through the CBL network to obtain an eighth feature of the second scale, where the second scale is half of the first scale; performing feature fusion on the seventh feature of the first scale and the eighth feature of the second scale to obtain the features of multiple scales.

[0018] In one embodiment of the present disclosure, the FPN further includes an upsampling network; the feature fusion of the seventh feature of the first scale and the eighth feature of the second scale to obtain the features of multiple scales includes: performing a 1x upsampling process on the eighth feature of the second scale through the upsampling network to obtain a ninth feature of the first scale; concatenating the seventh feature of the first scale and the ninth feature of the first scale according to the channel dimension to obtain a tenth feature of the first scale; performing a 1x upsampling process on the tenth feature of the first scale through the upsampling network to obtain an eleventh feature of the original scale; using the eleventh feature of the original scale, the tenth feature of the first scale, and the eighth feature of the second scale as the features of multiple scales.

[0019] In one embodiment of the present disclosure, the head structure includes an output head network, where the output head network includes a CSP network and a convolutional network; the fire detection and smoke detection based on the features of multiple scales by the head structure to obtain the candidate features includes: concatenating the target feature of the original scale and the eleventh feature of the original scale according to the channel dimension to obtain a concatenated feature of the original scale; performing feature extraction on the concatenated feature of the original scale through the CSP network to obtain a twelfth feature of the original scale; performing feature extraction on the twelfth feature of the original scale through the convolutional network to obtain the candidate feature of the original scale.

[0020] In one embodiment of the present disclosure, the output head network further includes a CBL network; after the CSP network extracts features from the concatenated features of the original scale to obtain the twelfth feature of the original scale, the method further includes: performing 1x downsampling on the twelfth feature of the original scale through the CBL network to obtain the thirteenth feature of the first scale; concatenating the tenth feature of the first scale and the thirteenth feature of the first scale along the channel dimension to obtain the concatenated feature of the first scale; extracting features from the concatenated feature of the first scale through the CSP network to obtain the fourteenth feature of the first scale; and extracting features from the fourteenth feature of the first scale through the convolutional network to obtain the candidate feature of the first scale.

[0021] In one embodiment of the present disclosure, after the CSP network extracts features from the concatenated feature of the first scale to obtain the fourteenth feature of the first scale, the method further includes: performing 1x downsampling on the fourteenth feature of the first scale through the CBL network to obtain the fifteenth feature of the second scale; concatenating the eighth feature of the second scale and the fifteenth feature of the second scale along the channel dimension to obtain the concatenated feature of the second scale; extracting features from the concatenated feature of the second scale through the CSP network to obtain the sixteenth feature of the second scale; and extracting features from the sixteenth feature of the second scale through the convolutional network to obtain the candidate feature of the second scale.

[0022] An embodiment of the second aspect of the present disclosure provides an object detection device, including: an acquisition module, configured to acquire a target image; an input module, configured to input the target image into a target detection model, where the target detection model includes a backbone structure, a neck structure, and a head structure, the backbone structure includes a cross-stage local CSP network, and the CSP network includes a lightweight convolutional neural network CNN; an extraction module, configured to extract features from the target image through the lightweight CNN in the CSP network of the backbone structure to obtain target features of the target image, where the target features include at least one of a color feature, a texture feature, a shape feature, and a spatial relationship feature of the target image; and a detection module, configured to perform fire detection and smoke detection on the target features through the neck structure and the head structure to obtain a target detection result of the target image.

[0023] An embodiment of the third aspect of the present disclosure provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the object detection method described in the embodiment of the first aspect of the present disclosure is implemented.

[0024] In a fourth aspect embodiment of the present application, a computer-readable storage medium is proposed, on which a computer program is stored. When the program is executed by a processor, the target detection method described in the first aspect embodiment of the present disclosure is implemented.

[0025] In a fifth aspect embodiment of the present application, a vehicle is proposed, including the target detection device described in the second aspect embodiment of the present disclosure; or the electronic device described in the third aspect embodiment of the present disclosure; or the computer-readable storage medium described in the fourth aspect embodiment of the present disclosure.

[0026] The technical solution provided by the embodiments of the present disclosure at least brings the following beneficial effects: The CSP network in this solution includes a lightweight CNN. Compared with the CSP network in the related art that includes a residual network, the target detection model is more lightweight, the volume of the target detection model and the required computing resources are smaller, which can reduce the storage space and computing resources occupied by the target detection model on the hardware device, facilitate the deployment of the target detection model in the hardware device with limited resource capacity, and use the lightweight CNN in the CSP network in the backbone structure for feature extraction to perform fire detection and smoke detection, which helps to improve the target detection efficiency and is applicable to the monitoring scenario of the charging station.

[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The above and / or additional aspects and advantages of the present disclosure will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0029] Figure 1 It is a flowchart of a target detection method according to an embodiment of the present disclosure;

[0030] Figure 2 It is a schematic diagram of a target detection model according to an embodiment of the present disclosure;

[0031] Figure 3 It is a schematic diagram of a CSP network according to an embodiment of the present disclosure;

[0032] Figure 4 It is a flowchart of a target detection method according to another embodiment of the present disclosure;

[0033] Figure 5 It is a flowchart of a target detection method according to another embodiment of the present disclosure;

[0034] Figure 6 It is a flowchart of obtaining candidate features in a target detection method according to an embodiment of the present disclosure;

[0035] Figure 7 Schematic structural diagram of an object detection device according to an embodiment of the present disclosure;

[0036] Figure 8 Schematic structural diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners

[0037] Embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, and should not be construed as a limitation to the present disclosure.

[0038] The object detection method, device, vehicle, electronic device, and storage medium according to embodiments of the present disclosure will be described below with reference to the accompanying drawings.

[0039] Figure 1 Flow schematic diagram of an object detection method according to an embodiment of the present disclosure.

[0040] As Figure 1 shown, the object detection method according to an embodiment of the present disclosure includes:

[0041] S101, obtaining an object image.

[0042] It should be noted that the execution subject of the object detection method according to an embodiment of the present disclosure is an electronic device, and the electronic device includes mobile phones, notebooks, desktop computers, vehicle-mounted terminals, smart home appliances, etc. The object detection method according to an embodiment of the present disclosure can be executed by the object detection device according to an embodiment of the present disclosure, and the object detection device according to an embodiment of the present disclosure can be configured in any electronic device to execute the object detection method according to an embodiment of the present disclosure.

[0043] It should be noted that there are no excessive limitations on the object image. For example, it may include two-dimensional images, three-dimensional images, etc.

[0044] In one implementation manner, taking the monitoring scenario of a charging station as an example, the object image may include an image of any area within the charging station. For example, it may include images of charging piles and vehicles within the charging station. An image acquisition device is provided within the charging station, and the above-mentioned images can be acquired through the image acquisition device. Among them, the image acquisition device may include a camera, a radar, etc.

[0045] S102, inputting the object image into an object detection model, where the object detection model includes a backbone structure, a neck structure, and a head structure, and the backbone structure includes a cross-stage local CSP network, and the CSP network includes a lightweight convolutional neural network CNN.

[0046] In an embodiment of the present disclosure, the object detection model includes a backbone structure, a neck structure, and a head structure. The backbone structure includes a CSP (Cross Stage Partial) network, and the CSP network includes a lightweight CNN (Convolutional Neural Networks). It should be noted that in the related art, the CSP network includes a residual network, and in this solution, the residual network in the CSP network is replaced with a lightweight CNN.

[0047] For example, the common file of the object detection model includes the network structure of the CSP network. The residual network in the CSP network in the common file can be modified to ShuffleNetV3, and the remaining networks in the CSP network in the common file except the residual network can be retained. Among them, the remaining networks can include an activation network and a max pooling network to update the common file of the object detection model.

[0048] The yaml file of ShuffleNetV3 can be added to the yaml file of the object detection model, and the output head network in the yaml file can be modified to fire detection and smoke detection to update the yaml file of the object detection model. It should be noted that both the common file and the yaml file are a kind of configuration file.

[0049] It should be noted that the object detection model is not overly limited. For example, it can include the YOLOv5 model. It should be noted that YOLOv5 is the fifth generation version of the YOLO (You Only Look Once) model, and the YOLO model is an object detection model.

[0050] It should be noted that the backbone structure, the CSP network, and the lightweight CNN are not overly limited. For example, in addition to the CSP network, the backbone structure can also include other networks. The backbone structure can include at least one CSP network. In addition to the lightweight CNN, the CSP network can also include other networks. The CSP network can include at least one lightweight CNN. The lightweight CNN can include a ShuffleNet network (such as ShuffleNetV3), a MobileNet network, etc.

[0051] In one implementation, taking the monitoring scenario of a charging station as an example, edge devices are set in the charging station, and the object detection model can be pre-deployed on the edge devices. Among them, the edge devices can include routers, routing switches, etc. For example, an image acquisition device can send a target image to the edge device, and the edge device can input the target image into the object detection model.

[0052] In one implementation, such asFigure 2 As shown in Figure 2 , the backbone structure further includes a Focus network, which includes a slicing network, a splicing network, and a CBL network. Among them, the CBL network includes a convolutional network, a normalization network, and an activation network. For example, the normalization network includes a BN (Batch Normalization) network, and the activation network includes a Relu (Rectified Linear Unit) network.

[0053] Before performing feature extraction on the target image through the lightweight CNN in the CSP network of the backbone structure, it also includes slicing the target image through the slicing network to obtain multiple sliced images, splicing the multiple sliced images through the splicing network to obtain a spliced image, and performing convolutional processing on the spliced image through the CBL network to obtain an updated target image.

[0054] It should be noted that the slicing process can be implemented by any slicing method in related technologies, and no excessive limitation is imposed here.

[0055] For example, the pixel points of the odd rows and odd columns of the target image can be combined through the slicing network to obtain sliced image 1, the pixel points of the odd rows and even columns of the target image can be combined through the slicing network to obtain sliced image 2, the pixel points of the even rows and odd columns of the target image can be combined through the slicing network to obtain sliced image 3, the pixel points of the even rows and even columns of the target image can be combined through the slicing network to obtain sliced image 4. The sliced images 1 to 4 can be spliced according to the channel dimension through the splicing network to obtain a spliced image, and the spliced image is subjected to convolutional processing through the CBL network to obtain an updated target image.

[0056] For example, if the height, width, and number of channels of the target image are 608, 608, and 3 respectively, the height, width, and number of channels of the sliced image are 304, 304, and 3 respectively, the height, width, and number of channels of the spliced image are 304, 304, and 12 respectively, and the height, width, and number of channels of the updated target image are 304, 304, and 32 respectively. The units of the height and width of the image are both px (pixel).

[0057] S103, perform feature extraction on the target image through the lightweight CNN in the CSP network of the backbone structure to obtain the target features of the target image, where the target features include at least one of the color feature, texture feature, shape feature, and spatial relationship feature of the target image.

[0058] It should be noted that the target features can be represented by feature maps, feature matrices, etc.

[0059] In one embodiment, the CSP network includes N feature extraction networks, where at least one feature extraction network includes a lightweight CNN, N is a positive integer, and the target features of the target image are obtained by performing feature extraction on the target image through the lightweight CNN in the CSP network of the backbone structure, including performing feature extraction on the target image through the i-th feature extraction network to obtain the i-th feature of the target image, and performing feature fusion on the N features of the target image to obtain the target features. Here, 1 ≤ i ≤ N, and i is a positive integer.

[0060] It should be noted that different feature extraction networks may be the same or different. For example, the CSP network includes Feature Extraction Network 1 and 2. Feature Extraction Network 1 includes X lightweight CNNs, and Feature Extraction Network 2 includes a convolutional network. Here, X is a positive integer.

[0061] S104, perform fire detection and smoke detection based on the target features through the neck structure and the head structure to obtain the target detection result of the target image.

[0062] It should be noted that the target detection result is not overly limited. For example, the target detection result includes the detection frame of the set object, the first image includes the set object, etc., and the set object includes fire and smoke. The detection frame can carry information such as the position and size of the detection frame.

[0063] In one embodiment, as Figure 2 shown, the neck structure includes FPN (Feature Pyramid Networks), where FPN includes a CSP network, a CBL network, and an upsampling network. It should be noted that the neck structure may include other networks in addition to FPN, and FPN may include other networks in addition to the CSP network, the CBL network, and the upsampling network. For example, FPN may also include a splicing network.

[0064] In one embodiment, as Figure 2 shown, the head structure includes an output head network, where the output head network includes a CSP network and a convolutional network. It should be noted that the head structure may include other networks in addition to the output head network. For example, the head structure may also include a CBL network and a splicing network. The head structure may include at least one output head network. For example, the head structure may include 3 output head networks, and each output head network includes 1 CSP network and 1 convolutional network.

[0065] In summary, according to the object detection method of the present disclosure embodiment, an object image is obtained and input into an object detection model. The object detection model includes a backbone structure, a neck structure, and a head structure. The backbone structure includes a CSP network, and the CSP network includes a lightweight CNN. The object image is feature-extracted through the lightweight CNN in the CSP network of the backbone structure to obtain the object features of the object image. Fire detection and smoke detection are performed based on the object features through the neck structure and the head structure to obtain the object detection result of the object image. Thus, the CSP network in this solution includes a lightweight CNN. Compared with the CSP network including a residual network in the related art, the object detection model is more lightweight, the volume of the object detection model and the required computing resources are smaller, the storage space and computing resources occupied by the object detection model on the hardware device can be reduced, and it is convenient to deploy the object detection model in a hardware device with limited resource capacity. The lightweight CNN in the CSP network of the backbone structure is used for feature extraction to perform fire detection and smoke detection, which helps to improve the object detection efficiency and is applicable to the monitoring scenario of charging stations.

[0066] Based on any of the above embodiments, the CSP network further includes a CBL network, a convolutional network, and a splicing network, where the CBL network includes a convolutional network, a normalization network, and an activation network. It can be understood that the CSP network may include at least one CBL network.

[0067] Based on any of the above embodiments, as Figure 3 shown, multiple lightweight CNNs in the CSP network form an object CNN. It should be noted that the connection manner between the multiple lightweight CNNs is not overly limited. For example, it may include at least one of series connection and parallel connection.

[0068] In step S103, the object image is feature-extracted through the lightweight CNN in the CSP network of the backbone structure to obtain the object features of the object image, including feature-extracting the object image through the object CNN in the CSP network of the backbone structure to obtain the object features. Thus, multiple lightweight CNNs can be used to form an object CNN in this method to perform feature extraction on the object image.

[0069] In one implementation, as Figure 3 shown, multiple lightweight CNNs in the CSP network are connected in series to form an object CNN.

[0070] Extract features from the target image through the target CNN in the CSP network in the backbone structure to obtain target features, including obtaining the concatenation order of multiple lightweight CNNs in the target CNN. Among them, two adjacent lightweight CNNs in the sorting are concatenated. Extract features from the target image through the i-th sorted lightweight CNN to obtain the i-th feature, where i is a positive integer. Extract features from the i-th extracted feature through the (i + 1)-th sorted lightweight CNN to obtain the (i + 1)-th feature. Obtain the feature obtained by the last sorted lightweight CNN in the target CNN as the target feature. Thus, when multiple lightweight CNNs are concatenated to form the target CNN, the target image can be gradually feature-extracted through the concatenated multiple lightweight CNNs to obtain the target feature.

[0071] In one implementation, as Figure 3 shown, the CSP network includes a first CBL network, a second CBL network, a target CNN, a convolutional network, and a splicing network. That is, the CSP network includes 2 CBL networks, a target CNN, a convolutional network, and a splicing network.

[0072] Figure 4 It is a schematic flowchart of a target detection method according to another embodiment of the present disclosure.

[0073] As Figure 4 shown, the target detection method of the embodiment of the present disclosure includes:

[0074] S401, obtain the target image.

[0075] S402, input the target image into the target detection model. Among them, the target detection model includes a backbone structure, a neck structure, and a head structure. The backbone structure includes a cross-stage local CSP network, and the CSP network includes a lightweight convolutional neural network CNN.

[0076] For the relevant content of steps S401 - S402, reference can be made to the above embodiment and will not be elaborated here.

[0077] S403, extract features from the target image through the CBL network to obtain the first feature of the target image.

[0078] S404, extract features from the first feature through the lightweight CNN to obtain the second feature.

[0079] S405, extract features from the target image through the convolutional network to obtain the third feature of the target image.

[0080] S406, splice the second feature and the third feature according to the channel dimension through the splicing network to obtain the fourth feature.

[0081] S407. Downsample the fourth feature through the CBL network to obtain the target feature.

[0082] It should be noted that, except for the target feature, the remaining image features in the embodiments of the present disclosure, such as the first to fourth features, all include at least one of the color feature, texture feature, shape feature, and spatial relationship feature of the target image, which will not be elaborated here.

[0083] It should be noted that the second feature and the third feature are concatenated according to the channel dimension to obtain the fourth feature. Any feature concatenation method in related technologies can be used to implement this. Concatenating according to the channel dimension will change the number of channels of the feature, but will not change the scale of the feature, that is, the scales of the second feature, the third feature, and the fourth feature are the same. For example, if the image feature includes the features of multiple sampling points with pixel height H and pixel width W, the scale of the image feature is H*W.

[0084] It should be noted that the subsampling process will reduce the number of sampling points of the feature, that is, reduce the scale of the feature, that is, the scale of the target feature is smaller than the scale of the fourth feature.

[0085] It should be noted that the CBL network and the lightweight CNN constitute the first feature extraction network for obtaining the second feature, and the convolutional network constitutes the second feature extraction network for obtaining the third feature. For example, continuing with Figure 3 as an example, the first CBL network and the target CNN constitute the first feature extraction network for obtaining the second feature, and the convolutional network constitutes the second feature extraction network.

[0086] For example, continuing with Figure 3 as an example, the target image can be feature-extracted through the first CBL network to obtain the first feature of the target image, the first feature can be feature-extracted through the target CNN to obtain the second feature, the target image can be feature-extracted through the convolutional network to obtain the third feature, the second feature and the third feature can be concatenated according to the channel dimension through the concatenation network to obtain the fourth feature, and the fourth feature can be downsampled through the second CBL network to obtain the target feature.

[0087] S408. Perform fire detection and smoke detection based on the target feature through the neck structure and the head structure to obtain the target detection result of the target image.

[0088] For the relevant content of step S408, reference can be made to the above embodiments, which will not be elaborated here.

[0089] In summary, according to the object detection method of the present disclosure embodiment, the second feature can be obtained through the CBL network and the lightweight CNN, the third feature can be obtained through the convolutional network, the second feature and the third feature are concatenated according to the channel dimension through the concatenation network to obtain the fourth feature, and the fourth feature is downsampled through the CBL network to obtain the target feature, so as to realize the acquisition of the target feature.

[0090] Figure 5 It is a schematic flow chart of an object detection method according to another embodiment of the present disclosure.

[0091] As Figure 5 shown, the object detection method of the present disclosure embodiment includes:

[0092] S501, obtain the target image.

[0093] S502, input the target image into the object detection model, where the object detection model includes a backbone structure, a neck structure, and a head structure, the backbone structure includes a cross-stage local CSP network, and the CSP network includes a lightweight convolutional neural network CNN.

[0094] S503, extract features from the target image through the lightweight CNN in the CSP network of the backbone structure to obtain the target features of the target image, where the target features include at least one of the color feature, texture feature, shape feature, and spatial relationship feature of the target image.

[0095] For the relevant content of steps S501-S503, reference can be made to the above embodiments and will not be elaborated here.

[0096] S504, based on the target features, perform fire detection and smoke detection through the neck structure and the head structure to obtain candidate features of the target image, where the candidate features include the confidence of the set object and the candidate detection box of the set object, and the set object includes fire and smoke.

[0097] It should be noted that there is no excessive limitation on the number of candidate features. For example, it may include 9.

[0098] In one implementation, performing fire detection and smoke detection based on the target features through the neck structure and the head structure to obtain candidate features of the target image includes sampling the target features through the neck structure to obtain sampling features, and performing fire detection and smoke detection based on the sampling features through the head structure to obtain candidate features.

[0099] It should be noted that there is no excessive limitation on the sampling process. For example, it may include upsampling, downsampling, etc. Upsampling will increase the number of sampling points of the feature, that is, increase the scale of the feature.

[0100] S505. Select the fifth feature from multiple candidate features based on the confidence levels among the multiple candidate features.

[0101] In one implementation, selecting the fifth feature from multiple candidate features based on the confidence levels among the multiple candidate features includes determining the candidate feature with the maximum confidence level as the fifth feature.

[0102] In one implementation, to select the fifth feature from multiple candidate features based on the confidence levels among the multiple candidate features, the NMS (Non-Maximum Suppression) algorithm in related technologies can be used for implementation.

[0103] S506. Use the set object in the fifth feature and the candidate detection box of the set object as the target detection result.

[0104] In one implementation, the candidate feature further includes the probability that the target image includes the set object. Using the set object in the fifth feature and the candidate detection box of the set object as the target detection result includes: if the probability of fire in the fifth feature is greater than or equal to the probability of smoke, using the fire and the candidate detection box of fire in the fifth feature as the target detection result; or, if the probability of fire in the fifth feature is less than the probability of smoke, using the smoke and the candidate detection box of smoke in the fifth feature as the target detection result.

[0105] For example, if the fifth feature includes a fire probability of 0.6, a smoke probability of 0.9, a fire detection box 1, and a smoke detection box 2, then the smoke and the smoke detection box 2 can be added to the target detection result.

[0106] In one implementation, the candidate feature further includes the probability that the target image includes the set object. Using the set object in the fifth feature and the candidate detection box of the set object as the target detection result further includes: if the probability of the set object in the fifth feature is greater than the first set threshold, using the set object in the fifth feature and the candidate detection box of the set object as the target detection result.

[0107] For example, taking the first set threshold as 0.5.

[0108] If the fifth feature includes a fire probability of 0.3, a smoke probability of 0.9, a fire detection box 1, and a smoke detection box 2, then the smoke and the smoke detection box 2 can be used as the target detection result.

[0109] If the fifth feature includes a fire probability of 0.6, a smoke probability of 0.9, a fire detection box 1, and a smoke detection box 2, then the fire, smoke, fire detection box 1, and smoke detection box 2 can be used as the target detection result.

[0110] In one embodiment, the method further includes generating an alarm message for reminding of a fire or smoke if the object detection result of each consecutive target image includes at least one of fire and smoke, and / or if the size of at least one detection box in the object detection result of each consecutive target image, including the detection box of fire and the detection box of smoke, is greater than a second set threshold.

[0111] It should be noted that there is no excessive limitation on the number of consecutive multiple target images. For example, it may include the number of target images obtained within a set time period, where the set time period can be 3 seconds.

[0112] It should be noted that there is no excessive limitation on the second set threshold. For example, the size of at least one detection box being greater than the second set threshold includes that the width of at least one detection box is greater than 5% of the width of the target image, and / or the height of at least one detection box is greater than 5% of the height of the target image.

[0113] It should be noted that there is no excessive limitation on the alarm message. For example, taking the monitoring scenario of a charging station as an example, the alarm message may include the identifier of the charging station, the parking space identifier, the charging pile identifier, the location of the flame or smoke, the target image, etc.

[0114] In some examples, taking the monitoring scenario of a charging station as an example, the edge device may send the alarm message to the control platform of the charging station to timely inform the user of the control platform that there is a fire or smoke in the charging station.

[0115] In summary, according to the object detection method of the present disclosure, candidate features of the target image can be obtained through the neck structure and the head structure, where the candidate features include the confidence of a set object and the candidate detection box of the set object, the set object includes fire and smoke, and considering the confidence in multiple candidate features, the fifth feature is selected from multiple candidate features, and the set object and the candidate detection box of the set object in the fifth feature are used as the object detection result.

[0116] Based on any of the above embodiments, as Figure 6 shown, in step S504, fire detection and smoke detection are performed on the target features through the neck structure and the head structure to obtain candidate features of the target image, including:

[0117] S601, performing sampling processing on the target features at multiple scales through the neck structure to obtain features at multiple scales.

[0118] It should be noted that there is no excessive limitation on the multiple scales. For example, it may include half of the original scale of the target features, 1 / 4, 2 times, etc.

[0119] In one implementation, the neck structure includes an FPN, where the FPN includes a CSP network and a CBL network. The neck structure performs multi-scale sampling processing on the target features to obtain features of multiple scales, including performing 1x downsampling processing on the target features of the original scale through the CBL network to obtain the sixth feature of the first scale, where the first scale is half of the original scale, performing feature extraction on the sixth feature of the first scale through the CSP network to obtain the seventh feature of the first scale, performing 1x downsampling processing on the seventh feature of the first scale through the CBL network to obtain the eighth feature of the second scale, where the second scale is half of the first scale, and performing feature fusion on the seventh feature of the first scale and the eighth feature of the second scale to obtain features of multiple scales.

[0120] For example, taking the original scale of 76*76 as an example, the first scale is 38*38, and the second scale is 19*19.

[0121] It should be noted that the feature fusion of the seventh feature of the first scale and the eighth feature of the second scale can be implemented by any feature fusion method in the related art. For example, it may include feature splicing, weighted average, etc.

[0122] In some examples, the FPN further includes an upsampling network. Based on the seventh feature of the first scale and the eighth feature of the second scale, features of multiple scales are obtained, including performing 1x upsampling processing on the eighth feature of the second scale through the upsampling network to obtain the ninth feature of the first scale, splicing the seventh feature of the first scale and the ninth feature of the first scale according to the channel dimension to obtain the tenth feature of the first scale, performing 1x upsampling processing on the tenth feature of the first scale through the upsampling network to obtain the eleventh feature of the original scale, and using the eleventh feature of the original scale, the tenth feature of the first scale, and the eighth feature of the second scale as features of multiple scales. It can be understood that in this embodiment, the multiple scales include the original scale, the first scale, and the second scale.

[0123] S602, based on the features of multiple scales through the head structure, perform fire detection and smoke detection to obtain candidate features.

[0124] In one implementation, based on the features of multiple scales through the head structure, perform fire detection and smoke detection to obtain candidate features, including performing fire detection and smoke detection based on the features of the i-th scale through the head structure to obtain the candidate features of the i-th scale.

[0125] In some examples, fire detection and smoke detection can be performed based on the eleventh feature of the original scale by the head structure to obtain candidate features of the original scale. Fire detection and smoke detection can be performed based on the tenth feature of the first scale by the head structure to obtain candidate features of the first scale. Fire detection and smoke detection can be performed based on the eighth feature of the second scale by the head structure to obtain candidate features of the second scale.

[0126] In one implementation, the head structure includes an output head network. Among them, the output head network includes a CSP network and a convolutional network. Fire detection and smoke detection are performed based on features of multiple scales by the head structure to obtain candidate features, including splicing the target feature of the original scale and the eleventh feature of the original scale according to the channel dimension to obtain the spliced feature of the original scale. Feature extraction is performed on the spliced feature of the original scale by the CSP network to obtain the twelfth feature of the original scale. Feature extraction is performed on the twelfth feature of the original scale by the convolutional network to obtain candidate features of the original scale.

[0127] In some examples, the output head network further includes a CBL network. After feature extraction is performed on the spliced feature of the original scale by the CSP network to obtain the twelfth feature of the original scale, it further includes performing 1x downsampling processing on the twelfth feature of the original scale by the CBL network to obtain the thirteenth feature of the first scale. The tenth feature of the first scale and the thirteenth feature of the first scale are spliced according to the channel dimension to obtain the spliced feature of the first scale. Feature extraction is performed on the spliced feature of the first scale by the CSP network to obtain the fourteenth feature of the first scale. Feature extraction is performed on the fourteenth feature of the first scale by the convolutional network to obtain candidate features of the first scale.

[0128] In some examples, after feature extraction is performed on the spliced feature of the first scale by the CSP network to obtain the fourteenth feature of the first scale, it further includes performing 1x downsampling processing on the fourteenth feature of the first scale by the CBL network to obtain the fifteenth feature of the second scale. The eighth feature of the second scale and the fifteenth feature of the second scale are spliced according to the channel dimension to obtain the spliced feature of the second scale. Feature extraction is performed on the spliced feature of the second scale by the CSP network to obtain the sixteenth feature of the second scale. Feature extraction is performed on the sixteenth feature of the second scale by the convolutional network to obtain candidate features of the second scale.

[0129] Thus, in this method, features of multiple scales can be obtained through the neck structure, and candidate features can be obtained based on features of multiple scales by the head structure. Fire detection and smoke detection can be performed by comprehensively considering features of multiple scales to obtain candidate features, which helps to improve the accuracy of target detection.

[0130] To implement the above embodiments, the present disclosure also proposes an object detection device.

[0131] Figure 7 It is a schematic structural diagram of an object detection device according to an embodiment of the present disclosure.

[0132] As Figure 7 shown, the object detection device 100 of the embodiment of the present disclosure includes: an acquisition module 110, an input module 120, an extraction module 130, and a detection module 140.

[0133] The acquisition module 110 is configured to acquire a target image;

[0134] The input module 120 is configured to input the target image into a target detection model, where the target detection model includes a backbone structure, a neck structure, and a head structure, and the backbone structure includes a cross-stage local CSP network, and the CSP network includes a lightweight convolutional neural network CNN;

[0135] The extraction module 130 is configured to perform feature extraction on the target image through the lightweight CNN in the CSP network of the backbone structure to obtain target features of the target image, where the target features include at least one of color features, texture features, shape features, and spatial relationship features of the target image;

[0136] The detection module 140 is configured to perform fire detection and smoke detection based on the target features through the neck structure and the head structure to obtain a target detection result of the target image.

[0137] In an embodiment of the present disclosure, the CSP network further includes a CBL network, a convolutional network, and a splicing network; the extraction module 130 is further configured to: perform feature extraction on the target image through the CBL network to obtain a first feature of the target image; perform feature extraction on the first feature through the lightweight CNN to obtain a second feature; perform feature extraction on the target image through the convolutional network to obtain a third feature of the target image; splice the second feature and the third feature according to the channel dimension through the splicing network to obtain a fourth feature; perform downsampling processing on the fourth feature through the CBL network to obtain the target features.

[0138] In an embodiment of the present disclosure, multiple lightweight CNNs in the CSP network form a target CNN; the extraction module 130 is further configured to: perform feature extraction on the target image through the target CNN in the CSP network of the backbone structure to obtain the target features.

[0139] In one embodiment of the present disclosure, a plurality of lightweight CNNs in the CSP network are connected in series to form the target CNN; the extraction module 130 is further configured to: obtain the serial order of the plurality of lightweight CNNs in the target CNN, where two adjacent lightweight CNNs in the sorting are connected in series; extract features of the target image through the i-th sorted lightweight CNN to obtain the i-th feature, where i is a positive integer; extract features of the i-th extracted feature through the (i + 1)-th sorted lightweight CNN to obtain the (i + 1)-th feature; obtain the feature obtained by the last sorted lightweight CNN in the target CNN as the target feature.

[0140] In one embodiment of the present disclosure, the detection module 140 is further configured to: perform fire detection and smoke detection based on the target feature through the neck structure and the head structure to obtain candidate features of the target image, where the candidate features include the confidence of a set object and a candidate detection box of the set object, and the set object includes fire and smoke; screen out a fifth feature from the plurality of candidate features based on the confidence in the plurality of candidate features; use the set object and the candidate detection box of the set object in the fifth feature as the target detection result.

[0141] In one embodiment of the present disclosure, the detection module 140 is further configured to: perform sampling processing of multiple scales on the target feature through the neck structure to obtain features of multiple scales; perform fire detection and smoke detection based on the features of multiple scales through the head structure to obtain the candidate features.

[0142] In one embodiment of the present disclosure, the neck structure includes a Feature Pyramid Network (FPN), where the FPN includes the CSP network and the CBL network; the detection module 140 is further configured to: perform 1x downsampling processing on the target feature of the original scale through the CBL network to obtain a sixth feature of the first scale, where the first scale is half of the original scale; extract features of the sixth feature of the first scale through the CSP network to obtain a seventh feature of the first scale; perform 1x downsampling processing on the seventh feature of the first scale through the CBL network to obtain an eighth feature of the second scale, where the second scale is half of the first scale; perform feature fusion on the seventh feature of the first scale and the eighth feature of the second scale to obtain the features of multiple scales.

[0143] In one embodiment of the present disclosure, the FPN further includes an upsampling network; the detection module 140 is further configured to: perform 1x upsampling on the eighth feature of the second scale through the upsampling network to obtain the ninth feature of the first scale; splice the seventh feature of the first scale and the ninth feature of the first scale according to the channel dimension to obtain the tenth feature of the first scale; perform 1x upsampling on the tenth feature of the first scale through the upsampling network to obtain the eleventh feature of the original scale; use the eleventh feature of the original scale, the tenth feature of the first scale, and the eighth feature of the second scale as the features of the multiple scales.

[0144] In one embodiment of the present disclosure, the head structure includes an output head network, where the output head network includes a CSP network and a convolutional network; the detection module 140 is further configured to: splice the target feature of the original scale and the eleventh feature of the original scale according to the channel dimension to obtain the spliced feature of the original scale; perform feature extraction on the spliced feature of the original scale through the CSP network to obtain the twelfth feature of the original scale; perform feature extraction on the twelfth feature of the original scale through the convolutional network to obtain the candidate feature of the original scale.

[0145] In one embodiment of the present disclosure, the output head network further includes a CBL network; after performing feature extraction on the spliced feature of the original scale through the CSP network to obtain the twelfth feature of the original scale, the detection module 140 is further configured to: perform 1x downsampling on the twelfth feature of the original scale through the CBL network to obtain the thirteenth feature of the first scale; splice the tenth feature of the first scale and the thirteenth feature of the first scale according to the channel dimension to obtain the spliced feature of the first scale; perform feature extraction on the spliced feature of the first scale through the CSP network to obtain the fourteenth feature of the first scale; perform feature extraction on the fourteenth feature of the first scale through the convolutional network to obtain the candidate feature of the first scale.

[0146] In one embodiment of the present disclosure, after extracting features of the splicing features of the first scale through the CSP network to obtain the fourteenth feature of the first scale, the detection module 140 is further configured to: perform 1x downsampling processing on the fourteenth feature of the first scale through the CBL network to obtain the fifteenth feature of the second scale; splice the eighth feature of the second scale and the fifteenth feature of the second scale according to the channel dimension to obtain the splicing feature of the second scale; extract features of the splicing feature of the second scale through the CSP network to obtain the sixteenth feature of the second scale; extract features of the sixteenth feature of the second scale through the convolutional network to obtain the candidate feature of the second scale.

[0147] It should be noted that for details not disclosed in the object detection device of the embodiments of the present disclosure, please refer to the details disclosed in the object detection method of the embodiments of the present disclosure, which will not be elaborated here.

[0148] In summary, for the object detection device of the embodiments of the present disclosure, the CSP network in this solution includes a lightweight CNN. Compared with the CSP network in the related art that includes a residual network, the object detection model is more lightweight, the volume of the object detection model and the required computing resources are smaller, the storage space and computing resources occupied by the object detection model on the hardware device can be reduced, it is convenient to deploy the object detection model in a hardware device with limited resource capacity, and the lightweight CNN in the CSP network in the backbone structure is used for feature extraction to perform fire detection and smoke detection, which helps to improve the object detection efficiency and is applicable to the monitoring scenario of charging stations.

[0149] To implement the above embodiments, as Figure 8 shown, the embodiments of the present disclosure propose an electronic device 200, including: a memory 210, a processor 220, and a computer program stored on the memory 210 and executable on the processor 220. When the processor 220 executes the program, the above object detection method is implemented.

[0150] For the electronic device of the embodiments of the present disclosure, by the processor executing the computer program stored in the memory, the CSP network in this solution includes a lightweight CNN. Compared with the CSP network in the related art that includes a residual network, the object detection model is more lightweight, the volume of the object detection model and the required computing resources are smaller, the storage space and computing resources occupied by the object detection model on the hardware device can be reduced, it is convenient to deploy the object detection model in a hardware device with limited resource capacity, and the lightweight CNN in the CSP network in the backbone structure is used for feature extraction to perform fire detection and smoke detection, which helps to improve the object detection efficiency and is applicable to the monitoring scenario of charging stations.

[0151] To implement the above embodiments, embodiments of the present disclosure propose a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above object detection method is implemented.

[0152] In the computer-readable storage medium of the embodiments of the present disclosure, by storing a computer program and executing it by a processor, the CSP network in this solution includes a lightweight CNN. Compared with the CSP network in the related art that includes a residual network, the object detection model is more lightweight, the volume of the object detection model and the required computing resources are smaller, which can reduce the storage space and computing resources occupied by the object detection model on the hardware device, facilitate the deployment of the object detection model in the hardware device with limited resource capacity, and use the lightweight CNN in the CSP network in the backbone structure to extract features for fire detection and smoke detection, which helps to improve the object detection efficiency and is applicable to the monitoring scenario of the charging station.

[0153] To implement the above embodiments, embodiments of the present disclosure propose a vehicle, including the above object detection device; or the above electronic device; or the above computer-readable storage medium.

[0154] In the vehicle of the embodiments of the present disclosure, the CSP network in this solution includes a lightweight CNN. Compared with the CSP network in the related art that includes a residual network, the object detection model is more lightweight, the volume of the object detection model and the required computing resources are smaller, which can reduce the storage space and computing resources occupied by the object detection model on the hardware device, facilitate the deployment of the object detection model in the hardware device with limited resource capacity, and use the lightweight CNN in the CSP network in the backbone structure to extract features for fire detection and smoke detection, which helps to improve the object detection efficiency and is applicable to the monitoring scenario of the charging station.

[0155] In the description of the present disclosure, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present disclosure.

[0156] In addition, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present disclosure, "a plurality" means two or more unless otherwise specifically defined.

[0157] In the present disclosure, unless otherwise clearly specified and defined, terms such as "mounted", "connected", "coupled", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present disclosure can be understood according to specific circumstances.

[0158] In the present disclosure, unless otherwise clearly specified and defined, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath" and "underneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.

[0159] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0160] Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.

Claims

1. A target detection method, characterized in that, Including: Obtain a target image; Input the target image into a target detection model, where the target detection model includes a backbone structure, a neck structure, and a head structure, and the backbone structure includes a cross-stage local CSP network, and the CSP network includes a lightweight convolutional neural network (CNN); Extract features from the target image through the lightweight CNN in the CSP network of the backbone structure to obtain target features of the target image, where the target features include at least one of color features, texture features, shape features, and spatial relationship features of the target image; Perform fire detection and smoke detection based on the target features through the neck structure and the head structure to obtain a target detection result of the target image.

2. The method according to claim 1, wherein The CSP network further includes a CBL network, a convolutional network, and a splicing network; The step of extracting features from the target image through the lightweight CNN in the CSP network of the backbone structure to obtain target features of the target image includes: Extract features from the target image through the CBL network to obtain first features of the target image; Extract features from the first features through the lightweight CNN to obtain second features; Extract features from the target image through the convolutional network to obtain third features of the target image; Splice the second features and the third features according to the channel dimension through the splicing network to obtain fourth features; Perform downsampling processing on the fourth features through the CBL network to obtain the target features.

3. The method according to claim 1, characterized in that, Multiple lightweight CNNs in the CSP network form a target CNN; The step of extracting features from the target image through the lightweight CNN in the CSP network of the backbone structure to obtain target features of the target image includes: Extract features from the target image through the target CNN in the CSP network of the backbone structure to obtain the target features.

4. The method according to claim 3, wherein Multiple lightweight CNNs in the CSP network are connected in series to form the target CNN; The step of extracting features from the target image through the target CNN in the CSP network of the backbone structure to obtain target features of the target image includes: Obtain the serial connection order of multiple lightweight CNNs in the target CNN, where two adjacent lightweight CNNs in the sorting are connected in series; Extract features from the target image through the i-th lightweight CNN in the sorting to obtain the i-th features, where i is a positive integer; Extract features from the i-th extracted features through the (i + 1)-th lightweight CNN in the sorting to obtain the (i + 1)-th features; Obtain the features obtained by the last lightweight CNN in the sorting in the target CNN as the target features.

5. The method according to any one of claims 1-4, characterized in that, The step of performing fire detection and smoke detection based on the target features through the neck structure and the head structure to obtain a target detection result of the target image includes: Fire detection and smoke detection are performed on the target feature based on the neck structure and the head structure to obtain candidate features of the target image, where the candidate features include the confidence of a set object and the candidate detection box of the set object, and the set object includes fire and smoke; Based on the confidence in multiple candidate features, a fifth feature is screened out from the multiple candidate features; The set object and the candidate detection box of the set object in the fifth feature are used as the target detection result.

6. The method according to claim 5, characterized in that The step of performing fire detection and smoke detection on the target feature based on the neck structure and the head structure to obtain candidate features of the target image includes: Performing multi-scale sampling processing on the target feature through the neck structure to obtain features of multiple scales; Performing fire detection and smoke detection based on the features of multiple scales through the head structure to obtain the candidate features.

7. The method according to claim 6, characterized in that, The neck structure includes a Feature Pyramid Network (FPN), where the FPN includes a CSP network and a CBL network; The step of performing multi-scale sampling processing on the target feature through the neck structure to obtain features of multiple scales includes: Performing 1x downsampling processing on the target feature of the original scale through the CBL network to obtain a sixth feature of the first scale, where the first scale is half of the original scale; Performing feature extraction on the sixth feature of the first scale through the CSP network to obtain a seventh feature of the first scale; Performing 1x downsampling processing on the seventh feature of the first scale through the CBL network to obtain an eighth feature of the second scale, where the second scale is half of the first scale; Performing feature fusion on the seventh feature of the first scale and the eighth feature of the second scale to obtain the features of multiple scales.

8. The method according to claim 7, wherein The FPN further includes an upsampling network; The step of performing feature fusion on the seventh feature of the first scale and the eighth feature of the second scale to obtain the features of multiple scales includes: Performing 1x upsampling processing on the eighth feature of the second scale through the upsampling network to obtain a ninth feature of the first scale; Concatenating the seventh feature of the first scale and the ninth feature of the first scale according to the channel dimension to obtain a tenth feature of the first scale; Performing 1x upsampling processing on the tenth feature of the first scale through the upsampling network to obtain an eleventh feature of the original scale; Taking the eleventh feature of the original scale, the tenth feature of the first scale, and the eighth feature of the second scale as the features of multiple scales.

9. The method according to claim 8, wherein The head structure includes an output head network, where the output head network includes a CSP network and a convolutional network; The step of performing fire detection and smoke detection based on the features of multiple scales through the head structure to obtain the candidate features includes: Concatenating the target feature of the original scale and the eleventh feature of the original scale according to the channel dimension to obtain a concatenated feature of the original scale; Feature extraction is performed on the concatenated features of the original scale through the CSP network to obtain the twelfth feature of the original scale; Feature extraction is performed on the twelfth feature of the original scale through the convolutional network to obtain the candidate features of the original scale.

10. The method according to claim 9, wherein The output head network further includes a CBL network; After performing feature extraction on the concatenated features of the original scale through the CSP network to obtain the twelfth feature of the original scale, the following steps are further included: Performing 1x downsampling on the twelfth feature of the original scale through the CBL network to obtain the thirteenth feature of the first scale; Concatenating the tenth feature of the first scale and the thirteenth feature of the first scale according to the channel dimension to obtain the concatenated features of the first scale; Performing feature extraction on the concatenated features of the first scale through the CSP network to obtain the fourteenth feature of the first scale; Performing feature extraction on the fourteenth feature of the first scale through the convolutional network to obtain the candidate features of the first scale.

11. The method according to claim 10, characterized in that, After performing feature extraction on the concatenated features of the first scale through the CSP network to obtain the fourteenth feature of the first scale, the following steps are further included: Performing 1x downsampling on the fourteenth feature of the first scale through the CBL network to obtain the fifteenth feature of the second scale; Concatenating the eighth feature of the second scale and the fifteenth feature of the second scale according to the channel dimension to obtain the concatenated features of the second scale; Performing feature extraction on the concatenated features of the second scale through the CSP network to obtain the sixteenth feature of the second scale; Performing feature extraction on the sixteenth feature of the second scale through the convolutional network to obtain the candidate features of the second scale.

12. A target detection device, characterized in that, Including: An acquisition module for acquiring a target image; An input module for inputting the target image into a target detection model, where the target detection model includes a backbone structure, a neck structure, and a head structure, the backbone structure includes a cross-stage local CSP network, and the CSP network includes a lightweight convolutional neural network CNN; An extraction module for performing feature extraction on the target image through the lightweight CNN in the CSP network of the backbone structure to obtain target features of the target image, where the target features include at least one of color features, texture features, shape features, and spatial relationship features of the target image; A detection module for performing fire detection and smoke detection on the target features through the neck structure and the head structure to obtain a target detection result of the target image.

13. An electronic device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the target detection method according to any one of claims 1-11 is implemented.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the target detection method according to any one of claims 1-11 is implemented.

15. A vehicle, characterized in that, Including: The target detection device according to claim 12; or the electronic device according to claim 13; Or a computer-readable storage medium as described in claim 14.