Distribution center flame smoke identification method, device and equipment and storage medium

Through the flame smoke recognition method based on centerNet network, the problem of slow detection speed and low accuracy in the early stage of fire in the prior art is solved, and the automatic and accurate identification of flames and smoke is achieved, which improves the reliability of the monitoring system.

CN120047898APending Publication Date: 2025-05-27SHANGHAI DONGPU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510151381.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the early stage of a fire, the existing target detection algorithm is slow and the detection accuracy is low when facing complex situations where smoke state changes and the flame size is very small.

Method used

By collecting image samples from scenes where flames and smoke appear, the image sample data set is obtained, and the data set is divided into training sets and test sets; the initial recognition model is built based on the centerNet network, the model is trained through the training set, and the parameters are adjusted to obtain the flame smoke recognition model; the images obtained in real time are input to the recognition model, the recognition results are output, and whether an alarm is issued.

Benefits of technology

It realizes automatic and accurate identification of flames and smoke, reduces dependence on manual monitoring, reduces labor costs and monitoring work intensity, reduces false alarms and missed reports, and improves the reliability of the monitoring system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047898A_ABST
    Figure CN120047898A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, and discloses a distribution center flame smoke recognition method, device and equipment and a storage medium. The method comprises the following steps: labeling and classifying collected image samples to obtain an image sample data set; dividing the image sample data set into a training set and a test set according to a predetermined proportion; constructing an initial recognition model based on a centerNet network; training the initial recognition model through a training set to obtain a flame smoke recognition model; and inputting a to-be-recognized distribution center unloading area image obtained in real time into the flame smoke recognition model, outputting a recognition result, and judging whether to give an alarm or not according to the recognition result. According to the invention, the flame smoke recognition model is constructed based on the centerNet network, and the flame and smoke can be automatically and accurately recognized through the flame smoke recognition model, so that the dependence on manual monitoring is reduced, and the labor cost is reduced; and meanwhile, false alarm and missing alarm can be reduced, and the reliability of the monitoring system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular, to a method, device, equipment and storage medium for flame and smoke recognition in a sorting center. Background Art

[0002] The sorting center is a key hub in logistics operations, with a complex environment and many potential fire hazards. Once a fire accident occurs, it is very likely to cause heavy casualties and property losses. Therefore, quickly and accurately detecting fires and smoke is a key measure to ensure the production safety of the sorting center.

[0003] However, existing object detection algorithms will have problems of slow detection speed and low detection accuracy when facing the complex situations of diverse smoke states and very small flame sizes in the initial stage of a fire.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to solve the problems of slow detection speed and low detection accuracy of existing object detection algorithms when facing the complex situations of diverse smoke states and very small flame sizes in the initial stage of a fire.

[0006] The first aspect of the present invention provides a method for flame and smoke recognition in a sorting center, including: collecting image samples from scenes with flames and smoke, performing annotation and classification processing on the collected image samples to obtain an image sample data set; dividing the image sample data set into a training set and a test set according to a predetermined ratio, where both the training set and the test set include flame image samples, smoke image samples and normal image samples; constructing an initial recognition model based on the centerNet network; training the initial recognition model with the training set, adjusting the parameters of the initial recognition model, and obtaining a flame and smoke recognition model after testing with the test set; inputting the image of the area to be recognized in the sorting center obtained in real time into the flame and smoke recognition model, outputting a recognition result, and determining whether to issue an alarm according to the recognition result.

[0007] Optionally, in the first implementation manner of the first aspect of the present invention, performing annotation and classification processing on the collected image samples to obtain an image sample data set includes the steps of: performing data augmentation processing on the image samples to obtain augmented picture samples; performing annotation and classification processing on the augmented picture samples to obtain an image sample data set.

[0008] Optionally, in the second implementation manner of the first aspect of the present invention, data augmentation processing is performed on the image sample to obtain an augmented picture sample, including the steps of: performing mirroring, rotation, scaling, cropping, translation, or Gaussian noise processing on the image sample to obtain a deformed picture sample; performing Cutout processing and / or Mix-up processing on the image sample and the deformed picture sample to obtain an augmented picture sample.

[0009] Optionally, in the third implementation manner of the first aspect of the present invention, annotation and classification processing are performed on the augmented picture sample to obtain an image sample data set, including the steps of: using the Labelme tool to perform feature annotation on the flames and smoke existing in the augmented picture sample to obtain an image sample annotated with feature information; classifying the image sample according to the feature information to obtain a flame image sample, a smoke image sample, and a normal image sample, and the flame image sample, the smoke image sample, and the normal image sample constitute the image sample data set.

[0010] Optionally, in the fourth implementation manner of the first aspect of the present invention, an initial recognition model is constructed based on the centerNet network, including the steps of: building a centerNet network framework, the centerNet network framework includes a backbone network, a feature extraction layer, and a detection head, wherein the detection head includes three branches of a heat map, an offset, and a size, which are respectively used to predict the target center point, the center point offset, and the target size; introducing a multi-scale processing mechanism into the centerNet network framework to detect targets at different scales and improve the detection ability for small-scale and large-scale targets, thereby constructing an initial recognition model.

[0011] Optionally, in the fifth implementation manner of the first aspect of the present invention, introducing a multi-scale processing mechanism into the centerNet network framework includes the steps of: using a smaller 3x3 convolutional kernel to perform convolutional processing on a higher-level feature map in the feature extraction layer of the CenterNet network framework to capture the fine features of small-scale targets; using larger 5x5 and / or 7x7 convolutional kernels to perform convolutional processing on a lower-level feature map to obtain the global features of large-scale targets; fusing the features of different levels and different scales through a top-down path and lateral connections to introduce a multi-scale processing mechanism into the centerNet network framework.

[0012] Optionally, in the sixth implementation manner of the first aspect of the present invention, a multi-scale processing mechanism is introduced into the centerNet network framework, including the steps of constructing a multi-scale feature representation by using multiple downsampling and upsampling operations through a Hourglass network structure; setting skip connections in the Hourglass network to fuse feature maps at different levels during the downsampling process with the feature maps at the corresponding positions during the upsampling process, so as to introduce a multi-scale processing mechanism into the centerNet network framework.

[0013] The second aspect of the present invention provides a sorting center flame and smoke recognition device, including: a labeling and classification module, configured to collect image samples from scenes with flames and smoke, perform labeling and classification processing on the collected image samples to obtain an image sample data set; a data division module, configured to divide the image sample data set into a training set and a test set according to a predetermined ratio, where both the training set and the test set include flame image samples, smoke image samples, and normal image samples; a model construction module, configured to construct an initial recognition model based on the centerNet network; a training module, configured to train the initial recognition model through the training set, adjust the parameters of the initial recognition model, and obtain a flame and smoke recognition model after testing by the test set; an identification module, configured to input an image of the area of the sorting center to be identified obtained in real time into the flame and smoke recognition model, output an identification result, and determine whether to issue an alarm according to the identification result.

[0014] Optionally, in the first implementation manner of the second aspect of the present invention, the labeling and classification module includes: an augmentation unit, configured to perform data augmentation processing on the image samples to obtain augmented picture samples; a classification unit, configured to perform labeling and classification processing on the augmented picture samples to obtain an image sample data set.

[0015] Optionally, in the second implementation manner of the second aspect of the present invention, the augmentation unit includes: a deformation subunit, configured to perform mirroring, rotation, scaling, cropping, translation, or Gaussian noise processing on the image samples to obtain deformed picture samples; a transformation subunit, configured to perform Cutout processing and / or Mix-up processing on the image samples and the deformed picture samples to obtain augmented picture samples.

[0016] Optionally, in the third implementation manner of the second aspect of the present invention, the classification unit includes: a labeling subunit, configured to perform feature labeling on the flames and smoke existing in the augmented picture samples through the Labelme tool to obtain image samples labeled with feature information; a classification subunit, configured to classify the image samples according to the feature information to obtain flame image samples, smoke image samples, and normal image samples, and the flame image samples, smoke image samples, and normal image samples constitute the image sample data set.

[0017] Optionally, in the fourth implementation manner of the second aspect of the present invention, the model construction module includes: a building unit for building a CenterNet network framework, which includes a backbone network, a feature extraction layer, and a detection head. Among them, the detection head includes three branches: a heat map, an offset, and a size, which are respectively used to predict the target center point, the center point offset, and the target size; an introduction unit for introducing a multi-scale processing mechanism into the CenterNet network framework to detect targets at different scales and improve the detection ability for small-scale and large-scale targets, thereby constructing an initial recognition model.

[0018] Optionally, in the fifth implementation manner of the second aspect of the present invention, the introduction unit includes: a small convolution kernel processing subunit for performing convolution processing on a higher-level feature map by using a smaller 3x3 convolution kernel in the feature extraction layer of the CenterNet network framework to capture the fine features of small-scale targets; a large convolution kernel processing subunit for performing convolution processing on a lower-level feature map by using larger 5x5 and / or 7x7 convolution kernels to obtain the global features of large-scale targets; a fusion subunit for fusing features at different levels and scales through a top-down path and lateral connections, thereby introducing a multi-scale processing mechanism into the CenterNet network framework.

[0019] Optionally, in the sixth implementation manner of the second aspect of the present invention, the introduction unit includes: a sampling subunit for constructing a multi-scale feature representation by using multiple downsampling and upsampling operations through an Hourglass network structure; a skip connection subunit for setting skip connections in the Hourglass network to fuse feature maps at different levels during the downsampling process with the corresponding position feature maps during the upsampling process, thereby introducing a multi-scale processing mechanism into the CenterNet network framework.

[0020] The third aspect of the present invention provides a sorting center flame and smoke recognition device, including: a memory and at least one processor, wherein computer-readable instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor calls the computer-readable instructions in the memory to enable the sorting center flame and smoke recognition device to execute each step of the sorting center flame and smoke recognition method as described above.

[0021] The fourth aspect of the present invention provides a computer-readable storage medium, in which computer-readable instructions are stored. When it runs on a computer, it enables the computer to execute each step of the sorting center flame and smoke recognition method as described above.

[0022] Beneficial effects: In the technical solution of the present invention, image samples are collected from scenes with flames and smoke, the collected image samples are labeled and classified to obtain an image sample data set; the image sample data set is divided into a training set and a test set according to a predetermined ratio, and both the training set and the test set include flame image samples, smoke image samples and normal image samples; an initial recognition model is constructed based on the centerNet network; the initial recognition model is trained by the training set, the parameters of the initial recognition model are adjusted, and after being tested by the test set, a flame and smoke recognition model is obtained; the image of the area of the sorting center to be recognized obtained in real time is input into the flame and smoke recognition model, and the recognition result is output, and whether to issue an alarm is judged according to the recognition result. The present invention constructs a flame and smoke recognition model based on the centerNet network, and the flame and smoke recognition model can realize automatic and accurate recognition of flames and smoke, which not only reduces the dependence on manual monitoring, reduces labor costs and the intensity of monitoring work; at the same time, it can also reduce false alarms and missed alarms and improve the reliability of the monitoring system. Description of the Drawings

[0023] Figure 1 The first flowchart of the sorting center flame and smoke recognition method provided by the embodiment of the present invention;

[0024] Figure 2 The second flowchart of the sorting center flame and smoke recognition method provided by the embodiment of the present invention;

[0025] Figure 3 The third flowchart of the sorting center flame and smoke recognition method provided by the embodiment of the present invention;

[0026] Figure 4 The fourth flowchart of the sorting center flame and smoke recognition method provided by the embodiment of the present invention;

[0027] Figure 5 The fifth flowchart of the sorting center flame and smoke recognition method provided by the embodiment of the present invention;

[0028] Figure 6 The sixth flowchart of the sorting center flame and smoke recognition method provided by the embodiment of the present invention;

[0029] Figure 7 The seventh flowchart of the sorting center flame and smoke recognition method provided by the embodiment of the present invention;

[0030] Figure 8 The eighth flowchart of the sorting center flame and smoke recognition method provided by the embodiment of the present invention;

[0031] Figure 9A structural schematic diagram of the sorting center flame and smoke recognition device provided by an embodiment of the present invention;

[0032] Figure 10 Another structural schematic diagram of the sorting center flame and smoke recognition device provided by an embodiment of the present invention;

[0033] Figure 11 A structural schematic diagram of the sorting center flame and smoke recognition equipment provided by an embodiment of the present invention. Detailed implementation manners

[0034] An embodiment of the present invention provides a sorting center flame and smoke recognition method, device, equipment and storage medium. Image samples are collected from scenes with flames and smoke, and the collected image samples are labeled and classified to obtain an image sample data set. The image sample data set is divided into a training set and a test set according to a predetermined ratio. Both the training set and the test set include flame image samples, smoke image samples and normal image samples. An initial recognition model is constructed based on the CenterNet network. The initial recognition model is trained with the training set, the parameters of the initial recognition model are adjusted, and after being tested by the test set, a flame and smoke recognition model is obtained. The image of the area of the sorting center to be recognized obtained in real time is input into the flame and smoke recognition model, and a recognition result is output. Whether to issue an alarm is judged according to the recognition result. The present invention constructs a flame and smoke recognition model based on the CenterNet network. Through the flame and smoke recognition model, automatic and accurate recognition of flames and smoke can be realized, which not only reduces the dependence on manual monitoring, reduces the labor cost and the intensity of monitoring work, but also can reduce false alarms and missed alarms and improve the reliability of the monitoring system.

[0035] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" or "having" and any deformation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0036] For ease of understanding, the specific process of the embodiment of the present invention is described below. Please refer to Figure 1 , the first embodiment of the sorting center flame and smoke recognition method in the embodiment of the present invention includes:

[0037] S100. Collect image samples from scenes with flames or smoke, annotate and classify the collected image samples to obtain an image sample dataset.

[0038] Specifically, in this embodiment, multiple scenes where flames or smoke may appear can be selected for image acquisition. Inside the sorting center, key attention is paid to the cargo storage area, the area with concentrated electrical equipment, and the loading and unloading area. For example, in the cargo storage area, collect images of flames and smoke caused by spontaneous combustion of goods or illegal hot work; in the area with concentrated electrical equipment, collect images of flames and smoke caused by electrical short circuits and overloads; in the loading and unloading area, record images of vehicle exhaust (similar to smoke) and flames and smoke generated by vehicle fires due to faults. At the same time, samples are also obtained from external scenes, such as smoke spreading from a nearby factory fire, images of flames and smoke from vehicle fires on the surrounding roads, etc. Use multiple high-definition cameras, set different shooting angles and distances to obtain diverse images. Each camera shoots continuously for several hours at a frame rate of 30 frames per second, and a total of more than 10,000 images are collected; annotate and classify the collected image samples to obtain an image sample dataset.

[0039] S200. Divide the image sample dataset into a training set and a test set according to a predetermined ratio, and both the training set and the test set include flame image samples, smoke image samples, and normal image samples.

[0040] In this embodiment, since the initially constructed initial recognition model needs to be trained with a large amount of data and verified before an accurate flame and smoke recognition model can be obtained. Therefore, in this embodiment, the image sample dataset needs to be divided into a training set and a test set according to a predetermined ratio. To ensure that there is a sufficient sample dataset for training the initial recognition model, the number of the training set should be greater than the number of the test set. Preferably, the number ratio of the training set to the test set is 8:2 - 6:4, where both the training set and the test set include flame image samples, smoke image samples, and normal image samples.

[0041] S300. Build an initial recognition model based on the CenterNet network.

[0042] Specifically, CenterNet transforms the object detection problem into a keypoint detection problem. For each object in the image, instead of predicting the bounding box as in traditional object detection algorithms, it predicts the center point of the object, that is, the central position of the object. This method infers the position and size of the object by detecting and characterizing the center point; the CenterNet network outputs a heatmap, where each peak on the heatmap represents the center position of an object. This heatmap is obtained by extracting and processing features from the input image, and the position and intensity of the peak represent the position and confidence of the object center in the image. Building a flame and smoke recognition model based on the CenterNet network has the following advantages: in the scenario of flame and smoke recognition in the sorting center, flames and smoke may present irregular shapes, and some small flames or a small amount of smoke are small objects. CenterNet can locate small objects more accurately by precisely detecting the center point. For example, at the initial stage of a fire, the center point of a small flame can be accurately found. Compared with traditional bounding box-based detection algorithms, it is not difficult to locate the boundary due to the small size of the object, reducing the problem of missed detection of small objects; flames and smoke are often slender, and the keypoint detection method of CenterNet is not limited by the shape of the object. It can more flexibly detect slender smoke strips or irregularly shaped flames without failing to detect due to irregular shapes.

[0043] Traditional object detection algorithms usually need to generate candidate regions (such as Selective Search or Region Proposal Network in the R-CNN series), and then classify and adjust the bounding boxes of the candidate regions. However, CenterNet is an end-to-end architecture that directly obtains the center points and corresponding attribute information of the objects from the image input, without the need for a complex candidate region generation and screening process, making the training and inference processes more concise. Due to avoiding the complex candidate region generation step, CenterNet has a high speed during inference, which is very important for real-time flame and smoke recognition. In environments such as sorting centers, the collected images can be processed quickly to timely detect potential fire hazards and avoid delaying the warning and handling time due to slow detection speed. CenterNet also has good adaptability to a small amount of data to a certain extent. In the scenario of flame and smoke recognition, if the sample data volume is not particularly sufficient, through some data augmentation techniques and combined with the advantages of CenterNet, the performance of the model can be guaranteed to a certain extent, enabling the model to better learn the characteristics of flames and smoke and achieve relatively accurate recognition. In short, constructing a flame and smoke recognition model based on the CenterNet network can utilize its unique key point detection method, end-to-end architecture, flexible network selection, and the convenience of data annotation and training, etc., to better meet the requirements of accuracy, real-time performance, and adaptability for flame and smoke recognition in environments such as sorting centers, providing strong support for the early warning and prevention of fires.

[0044] S400. Train the initial recognition model with the training set, adjust the parameters of the initial recognition model, and obtain a flame and smoke recognition model after testing with the test set.

[0045] S500. Input the image of the area of the sorting center to be recognized obtained in real time into the flame and smoke recognition model, output the recognition result, and determine whether to issue an alarm according to the recognition result.

[0046] In this embodiment, after training the initial recognition model with the training set and the test set and passing the test, a flame and smoke recognition model is obtained. Based on this model, the image of the area of the sorting center to be recognized can be quickly recognized in real time, so as to determine whether there is flame or smoke in the image. If flame or smoke is detected, the background system will be immediately reported. The background system can trigger alarms, notify relevant personnel, or start emergency measures such as fire extinguishing equipment. At the same time, the detection results can be recorded and stored for subsequent analysis and evaluation. If the recognition result is a normal hat image sample, no alarm will be issued. The present invention constructs a flame and smoke recognition model based on the CenterNet network. Through the flame and smoke recognition model, automatic and accurate recognition of flame and smoke can be realized, which not only reduces the dependence on manual monitoring, but also reduces the labor cost and the intensity of monitoring work. At the same time, it can also reduce false alarms and missed alarms and improve the reliability of the monitoring system.

[0047] Please refer to Figure 2 , the second embodiment of the method for recognizing flame and smoke in the sorting center according to the embodiment of the present invention includes: S110. Perform data augmentation processing on the image sample to obtain an augmented picture sample; S120. Perform annotation and classification processing on the augmented picture sample to obtain an image sample data set.

[0048] Specifically, since the number of initially obtained image samples is limited and it is not conducive to training the constructed network, it is necessary to perform data augmentation processing on the obtained image samples through data augmentation to obtain more augmented picture samples; then perform annotation processing and classification processing with special marks on these augmented picture samples, and place the processed augmented picture samples in the corresponding directories to form an image sample data set. In this embodiment, by training a machine learning or deep learning model with a large number of annotated and classified image sample data sets, the recognition accuracy of flame and smoke can be significantly improved; at the same time, the data augmentation processing enables the model to come into contact with more diverse image samples during the training process, so that it can maintain a high recognition ability when facing various complex situations in actual applications.

[0049] Please refer to Figure 3 , the third embodiment of the method for recognizing flame and smoke in the sorting center according to the embodiment of the present invention includes: S111. Perform mirroring, rotation, scaling, cropping, translation, or Gaussian noise processing on the image sample to obtain a deformed picture sample; S112. Perform Cutout processing and / or Mix-up processing on the image sample and the deformed picture sample to obtain an augmented picture sample.

[0050] Specifically, in this embodiment, first, the initial image samples are processed by common data augmentation means such as mirroring, rotation, scaling, cropping, translation, or Gaussian noise to generate new pictures. Although such processed pictures are new to the network model and can achieve the purpose of expanding the dataset, the labels of the original pictures will not change for the deformed picture samples after such processing, and the training effect is not good. Therefore, in this embodiment, Cutout processing and / or Mix-up processing are further performed on these deformed picture samples and the initial image samples. The labels of the expanded picture samples obtained in this way may be different from the labels of the original pictures, which is beneficial to the training of the network model and a network model with better prediction effect can be obtained. Specifically, the basic operation of the Cutout processing method is to randomly select a rectangular area in the image and set the pixel values in this area to 0 or other fixed values (such as the mean value). When processing flame and smoke images, a rectangular area with a size of, for example, 50x50 pixels may be randomly selected in the image (the specific size can be adjusted according to the image resolution and actual requirements), and the pixel values in this area are all changed to 0. The purpose of this is to simulate the situation of partial information loss in the image, so that the model can learn during the training process that even if some features of the target object are occluded or lost, it can still accurately detect the flame or smoke. For example, in a fire scene image, the flame may be partially occluded by some debris, and the Cutout method can enable the model to adapt to this situation and enhance its robustness. As an example, the size and position of the occlusion area can be controlled, and different degrees of information loss can be simulated by using different ratios (such as 10%, 20%, 50%, etc.). In the recognition of flame and smoke in the sorting center, if the image resolution is 512x512 pixels, when the occlusion ratio is set to 10%, an occlusion area of approximately 51x51 pixels may be generated (the actual size will vary according to the aspect ratio of the image); when the ratio is increased to 20%, the occlusion area will increase accordingly and may reach about 72x72 pixels. In addition to a single rectangular area, multiple irregularly shaped occlusion areas can also be tried to further increase the diversity of the data. For example, circular, triangular, or other polygonal areas can be used for occlusion, or multiple occlusion areas with different shapes and sizes can be set in the image at the same time, so that the model can handle more complex scenarios and avoid overfitting to the flame and smoke characteristics of specific shapes and positions.

[0051] The basic principle of the Mix-up processing method is to linearly combine two different sample images to generate a new sample image. The specific operation is to randomly select two samples and then mix their pixel values according to a certain ratio. In flame and smoke recognition, for example, there is a flame image A and a smoke image B. When performing the Mix-up operation according to the ratio α (such as α = 0.4), the pixel value calculation formula for the newly generated image is: new_image = α * image_A + (1 - α) * image_B. At the same time, the labels of the two samples are also mixed according to the same ratio. Assuming the label of the flame image A is [1, 0, 0] (indicating belonging to the flame category) and the label of the smoke image B is [0, 1, 0] (indicating belonging to the smoke category), then the label of the new image is [0.4, 0.6, 0]. This can prevent the model from overly relying on specific image features and improve its generalization ability. The method of this embodiment can be extended to more samples (such as three or four) to generate more complex synthetic images, further enhancing the robustness of the model. For example, select the flame image A, the smoke image B, and the normal scene image C, and mix them according to different ratios (such as α = 0.3, β = 0.3, γ = 0.4). The pixel value of the new image is: new_image = α * image_A + β * image_B + γ * image_C, and the label is also mixed accordingly. Randomly flip the image with a certain probability (such as 50%) to increase the data diversity. Randomly rotate (such as ±30 degrees) and scale (such as 0.8 to 1.2 times) the image to simulate different shooting angles and distance changes, enabling the model to adapt to the morphological changes of flames and smoke in various actual scenarios and improving the accuracy and reliability of flame and smoke recognition in the complex environment of the sorting center.

[0052] Compared with a single data augmentation method or the case without using data augmentation, the data augmentation strategy provided in this embodiment can improve the generalization ability of the model, reduce the risk of overfitting, and thus obtain more stable and accurate recognition results in practical applications.

[0053] Please refer to Figure 4 , the fourth embodiment of the method for recognizing flames and smoke in the sorting center of express parcels in the embodiment of the present invention includes: S121. Use the Labelme tool to perform feature annotation on the flames and smoke existing in the augmented picture samples to obtain image samples annotated with feature information; S122. Classify the image samples according to the feature information to obtain flame image samples, smoke image samples, and normal image samples, and the flame image samples, smoke image samples, and normal image samples constitute the image sample dataset.

[0054] In this embodiment, Labelme is an open-source image annotation tool that supports various annotation methods such as polygons, rectangles, and circles, and is suitable for complex image annotation tasks. In this embodiment, the Labelmg annotation tool is used to annotate the collected images, defining three main categories: "flame", "smoke", and "normal", and assigning a unique label to each category. For example, "flame" is 1, "smoke" is 2, and "normal" is 0. For flames and smoke, the polygon tool is used to precisely outline the target area. During the annotation process, if an image with multiple situations coexists, such as both flame and smoke, the flame and smoke areas are annotated separately. When annotating an image with both flame and smoke, the category division is determined based on the hazard level, area ratio, or impact on the monitoring focus of the flame and smoke in the image to determine the main target. In the sorting center scenario, if the flame area is large and in the critical initial stage of combustion, which poses a great threat to the safety of goods, even if there is a lot of smoke, the image will be classified into the "flame" category; if the smoke spreads widely, seriously affecting visibility and air quality and hindering personnel evacuation and rescue work, even if the flame is large, it can be classified into the "smoke" category. This classification method enables the model to focus on key risks, prioritize the identification and early warning of major hazards, quickly lock in the core danger in the fire monitoring of the sorting center, and provide a clear direction for subsequent emergency measures such as fire extinguishing and evacuation. To further refine the classification, sub-categories are added, such as "open flame", "smoldering flame", "white smoke", "black smoke", etc. After the annotation is completed, the images are classified and sorted. All images annotated as "flame" and related to its sub-categories are classified into the flame-class dataset, images annotated as "smoke" and related to its sub-categories are classified into the smoke-class dataset, and the "normal" class images are stored separately. Finally, 3000 flame-class images, 4000 smoke-class images, and 3000 normal-class images are obtained. The image sample dataset constructed in this embodiment has rich diversity, covering images of flames, smoke, and normal situations in various scenarios around the sorting center, providing sufficient and comprehensive data for model training. The accuracy of annotation and classification ensures that the model can accurately learn the features of different categories during training.

[0055] Please refer to Figure 5 , the fifth embodiment of the flame and smoke recognition in the sorting center in the embodiment of the present invention includes: S310. Build a centerNet network framework, where the centerNet network framework includes a backbone network, a feature extraction layer, and a detection head. Among them, the detection head includes three branches: a heat map, an offset, and a size, which are respectively used to predict the target center point, the center point offset, and the target size; S320. Introduce a multi-scale processing mechanism into the centerNet network framework to detect targets at different scales, improve the detection ability for small-scale and large-scale targets, and thus construct an initial recognition model.

[0056] In this embodiment, the CenterNet network framework can select different backbone networks according to specific application scenarios and computing resources, such as ResNet, DenseNet, Hourglass, etc. In the scenario of flame and smoke recognition, if the depth and complexity of features are emphasized and computing resources permit, deeper networks such as ResNet-50 or ResNet-101 can be selected because they have strong feature extraction capabilities and can process complex image information. If the need for multi-scale feature refinement of slender targets is considered, the Hourglass network is a good choice. Its unique downsampling and upsampling structures can capture feature information at different levels. For example, in a typical application of flame and smoke recognition in a sorting center, if the data volume is large and high detection accuracy is desired, ResNet-101 can be used as the backbone network because its depth can extract richer feature information and has better discrimination ability for the subtle differences between different types of flames (such as open flames and smoldering flames) and smokes (such as white smoke and black smoke).

[0057] Based on the backbone network, this embodiment extracts features from the input image through a series of operations such as convolution and pooling. Assuming the input image is [image description], after passing through the backbone network, a feature map F is obtained. This process gradually reduces the spatial resolution of the image while increasing the dimension of the features. For example, if the input is an RGB image with a size of 512*512*512 (width, height, channels), after passing through the backbone network and the feature extraction layer, a feature map of 64*64*256 may be obtained. Although the size of the feature map becomes smaller, the feature representation at each position becomes more abstract and advanced, containing rich semantic information and being able to better describe various objects and background information in the image. During the feature extraction process, convolution kernels and pooling kernels of different sizes, as well as different strides and padding methods, are used to achieve feature abstraction at different levels. For flame and smoke recognition, larger pooling kernels and strides may lose some detailed information but can capture overall features such as the range of large-area smoke or flames, while smaller convolution kernels and strides can retain more details and help detect the fine features of small flames or small amounts of smoke.

[0058] The detection head consists of three branches: heatmap, offset, and size. Among them, the heatmap branch is mainly used to predict the center point position of the target. By performing a series of convolutional operations on the feature map, a heatmap H is generated. The value at each position in the heatmap H represents the probability that this position is the center point of the target. For example, for an image containing flames and smoke, the peak position in the heatmap corresponds to the center point of the flame or smoke. During training, the Focal Loss function is used to handle the learning of the heatmap. Since the targets of flames and smoke in the image are usually sparse, Focal Loss can solve the problem of unbalanced positive and negative samples, enabling the network to focus more on the learning of difficult samples (i.e., the real center points of the targets), and improving the accuracy of heatmap prediction. During inference, by finding the peaks (i.e., local maxima) in the heatmap and their corresponding positions, the center position of the target can be initially located. For the application in the sorting center, the heatmap can clearly show the center positions of the flames and smoke, providing a basis for subsequent processing.

[0059] Since the downsampling process will cause the spatial resolution of the feature map to decrease, the position of the target center point will have a certain offset. The offset branch aims to predict this offset of the center point. For the feature map F, an offset map O is output through a separate convolutional layer, which contains the offset information (in the x and y directions) of each possible center point position. During training, the Smooth L1 Loss is used to regress the offset, ensuring that the predicted center position can be accurately mapped back to the original image space. In flame and smoke recognition, accurate offset prediction can accurately restore the center point found on the heatmap to the original image, avoiding the positioning error caused by downsampling, and ensuring the accurate positioning of the center points of flames and smoke, which is particularly important for the positioning of small targets (such as small flames).

[0060] The size branch is responsible for predicting the size of the target. By performing another convolutional operation on the feature map F, a size map S is obtained, which can predict the width and height of the target. During the training process, the Smooth L1 Loss is also used for the regression training of the size, enabling the model to accurately predict the size range of the flames and smoke. For flames and smoke of different sizes, the size branch can help determine their ranges, such as distinguishing between small flames and large-area flames, and between a small amount of smoke and dense smoke, which is helpful for subsequent fire assessment and processing decisions.

[0061] By introducing a multi-scale processing mechanism into the CenterNet network framework, better detection can be achieved for small flames, a small amount of smoke, or large-scale fires and thick smoke in the sorting center. For small targets, the use of small convolutional kernels, high-resolution inputs, and high-level feature maps can ensure that their fine features are captured; for large targets, large convolutional kernels, low-resolution inputs, and sliding window techniques, etc., guarantee the accurate detection of their overall features and positions, reducing the situation of missed detections. Combining different feature extraction methods and multi-scale processing, the CenterNet network can better adapt to the complex environment of the sorting center, including flames and smoke under different lighting conditions, different perspectives, and different distances. Through multi-resolution inputs and feature fusion, the model can maintain relatively stable detection performance in different scene changes. Experiments show that after introducing the multi-scale processing mechanism, on a test set containing various sizes of flames and smoke, the detection recall rate for small targets can be increased by about 20%-30%, and the detection recall rate for large targets can be increased by about 15%-25%. Further, the synergistic effect of each branch of the detection head and the multi-scale processing mechanism makes the positioning of the target center point, the correction of the offset, and the prediction of the size more accurate. In practical applications, not only can the positions of flames and smoke be accurately found, but also their ranges can be accurately described, providing more accurate information for subsequent fire assessment and handling. For example, for the positioning error of flames and smoke, after introducing multi-scale processing and accurate offset prediction, it can be reduced from the original average error of 10 pixels to about 3-5 pixels, improving the positioning accuracy. At the same time, due to the combination of multi-scale feature fusion and different resolution inputs, the overall detection efficiency is also improved, reducing the processing time while maintaining high accuracy, which is crucial for real-time flame and smoke monitoring.

[0062] Please refer to Figure 6 , the sixth embodiment of the flame and smoke recognition in the sorting center in the embodiment of the present invention includes:

[0063] S321. Through a 3x3 convolutional kernel with a smaller size in the feature extraction layer of the CenterNet network framework, perform convolutional processing on a higher-level feature map to capture the fine features of small-scale targets;

[0064] S322. Use 5x5 and / or 7x7 convolutional kernels with a larger size to perform convolutional processing on a lower-level feature map to obtain the global features of large-scale targets;

[0065] S323. Through a top-down path and lateral connections, fuse features at different levels and different scales to implement the introduction of a multi-scale processing mechanism in the CenterNet network framework.

[0066] Specifically, in the feature extraction layer, convolutional kernels of different sizes are configured for feature maps at different levels. For example, on the lower-level feature maps closer to the input layer, large convolutional kernels of 5x5 and / or 7x7 are used. This is because the receptive field of the lower-level feature maps is large, and large convolutional kernels can obtain more global feature information. For large-scale flame or smoke targets, they can better capture their overall contours and approximate positions. On the higher-level feature maps, small convolutional kernels of 3x3 are used. The small convolutional kernels have a small receptive field and are more suitable for capturing the fine details of small-scale targets, such as the edges of small flames and the local textures of smoke. A top-down path and a lateral connection structure are constructed. The top-down path can transfer the semantic information of high-level features to the low-level, and the lateral connections are used to fuse features at different levels but the same spatial position. Taking a feature extraction layer with 5 layers as an example, after the feature map of the 5th layer is adjusted in channel number through a 1x1 convolution, it is added to the feature map of the 4th layer (lateral connection), and then upsampled and fused with the feature map of the 3rd layer (top-down path), and so on. In this way, the advantages of feature maps at different levels can be integrated, and the detection ability for targets of different scales can be enhanced.

[0067] Please refer to Figure 7 , the seventh embodiment of flame and smoke recognition in the sorting center in the embodiments of the present invention includes:

[0068] S321’: Construct a multi-scale feature representation by using multiple downsampling and upsampling operations through the Hourglass network structure;

[0069] S322’: Set skip connections in the Hourglass network structure to fuse the feature maps at different levels in the downsampling process with the feature maps at the corresponding positions in the upsampling, so as to introduce a multi-scale processing mechanism into the centerNet network framework.

[0070] Specifically, in the feature extraction layer part, the Hourglass network structure is introduced. Through multiple downsampling and upsampling operations, the Hourglass network can gradually refine the feature map. In the Hourglass network structure, the downsampling process is achieved through convolutional layers and pooling layers. Usually, convolutional operations with a stride greater than 1 or max-pooling operations are used to gradually reduce the size of the feature map. For example, when inputting an image with a size of 256×256, after the first downsampling module, the size of the feature map becomes 128×128, and after subsequent downsampling modules, the size may successively become 64×64, 32×32, etc. During this process, the receptive field continuously increases, and the network can obtain information in a larger area of the image, which is beneficial to capturing the overall features of large-scale targets, such as the shape of a large-area flame and the approximate range of thick smoke. The upsampling operation is the opposite of downsampling. By using methods such as transposed convolution (deconvolution) or interpolation, the reduced feature map is restored to a larger size. For example, gradually upsampling from a 32×32 feature map to 64×64, 128×128, etc. During the upsampling process, although the spatial resolution increases, due to the loss of some detailed information during downsampling, the upsampled feature map can supplement some details, especially suitable for capturing the details of slender targets such as flames and smoke, such as the edges of small flames and the textures of smoke. Skip connections are set in the Hourglass network to fuse the feature maps at different levels during the downsampling process with the feature maps at the corresponding positions during upsampling. For example, during the process of downsampling from 256×256 to 128×128, the information of the 128×128 feature map will be saved; when upsampling from 64×64 back to 128×128, the previously saved 128×128 feature map is added or concatenated with the currently upsampled 128×128 feature map. The advantage of doing this is that it can fuse the feature information at different scales, enabling the network to have both the global information of large-scale features and the detailed information of small-scale features. For the detection of flames and smoke, this fused feature can more accurately describe the target, improving the accuracy and reliability of the detection.

[0071] Please refer to Figure 8 , the eighth embodiment of the flame and smoke recognition in the sorting center according to the embodiment of the present invention includes:

[0072] S410. Use the random initialization method to assign initial values to all parameters in the initial recognition model;

[0073] S420. Randomly select picture samples from the training set and input them into the initial recognition model for training. Calculate the prediction result through forward propagation, and then calculate the loss value according to the Focal Loss function;

[0074] S430. Use the backpropagation algorithm to update the parameters of the network to gradually reduce the loss value;

[0075] S440. Repeat the above training steps until a predetermined number of training rounds is reached to obtain a trained initial recognition model;

[0076] S450. Use the test set to evaluate the trained initial recognition model, and adjust the hyperparameters according to the evaluation results to obtain an optimized flame and smoke recognition model.

[0077] Specifically, before model training, use the random initialization method to assign initial values to all parameters in the initial recognition model. For example, set the initial number of iterations (such as 5000 times), the initial learning rate (such as 10^-5), and batch_size (such as 300), and set the dataset path parameters and class parameters, and use parallel GPUs for training; randomly select a batch of images and corresponding bounding box annotation data from the training set and input them into the model for training; after feature extraction and processing of the prediction branch, obtain the predicted class probabilities and bounding box positions; calculate the classification loss and regression loss according to the prediction results and the true annotations. The classification loss uses focal loss, and the regression loss uses smooth L1 loss. Among them, focal loss is used as the classification loss function, and the difference between the class probability predicted by the initial recognition model and the true class is measured through the classification loss function. The calculation formula of Focal loss is as follows: L_cls = -α(1 - p_t)^γlog(p_t), where p_t is the class probability predicted by the initial recognition model, and α and γ are hyperparameters used to adjust the weights of different difficult and easy samples. When p_t is close to 1, (1 - p_t)^γ will become smaller, making the loss value of this sample smaller; conversely, when p_t is close to 0, (1 - p_t)^γ will become larger, making the loss value of this sample larger; smooth L1 loss is used as the regression loss function, and the difference between the bounding box position predicted by the initial recognition model and the true bounding box position is measured through the regression loss function. The calculation formula of Smooth L1 loss is as follows: L_reg = 0.5*(IoU - 1)^2*1 if |IoU - 1| < 1, otherwise L_reg = |IoU - 1| - 0.5, where IoU represents the IOU value between the predicted bounding box and the true bounding box. When the IoU value is close to 1, it means that the predicted box is very close to the true box, and the loss value is smaller at this time; when the IoU value deviates from 1, the loss value will gradually increase. Focal Loss can dynamically adjust the weights of different class samples by improving the standard cross-entropy loss function.

[0078] Next, according to the gradient of the loss function, backpropagate to all the parameters of the model, update the parameter values, and repeat the training until the preset number of training rounds is reached; use the test set to evaluate the performance of the trained model, calculate metrics such as accuracy, recall, and mAP to evaluate the performance of the model; finally, test the verification results and adjust the hyperparameters, such as the learning rate, batch size, etc., to optimize the performance of the model and obtain an optimized flame and smoke recognition model.

[0079] The present invention inputs the image of the area of the sorting center to be recognized obtained in real time into the flame and smoke recognition model, outputs the recognition result, and determines whether to issue an alarm according to the recognition result. The present invention constructs a flame and smoke recognition model based on the centerNet network. Through the flame and smoke recognition model, automatic and accurate recognition of flames and smoke can be realized, which not only reduces the dependence on manual monitoring, but also reduces the labor cost and the intensity of monitoring work; at the same time, it can also reduce false alarms and missed alarms and improve the reliability of the monitoring system.

[0080] The method for recognizing flames and smoke in the sorting center in the embodiment of the present invention has been described above. Next, the device for recognizing flames and smoke in the sorting center in the embodiment of the present invention will be described. Please refer to Figure 9 , an embodiment of the device for recognizing flames and smoke in the sorting center in the embodiment of the present invention includes:

[0081] The annotation and classification module 10 is used to collect image samples from scenes with flames and smoke, perform annotation and classification processing on the collected image samples, and obtain an image sample data set;

[0082] The data division module 20 is used to divide the image sample data set into a training set and a test set according to a predetermined ratio. Both the training set and the test set include flame image samples, smoke image samples, and normal image samples;

[0083] The model construction module 30 is used to construct an initial recognition model based on the centerNet network;

[0084] The training module 40 is used to train the initial recognition model through the training set, adjust the parameters of the initial recognition model, and obtain a flame and smoke recognition model after testing by the test set;

[0085] The recognition module 50 is used to input the image of the area of the sorting center to be recognized obtained in real time into the flame and smoke recognition model, output the recognition result, and determine whether to issue an alarm according to the recognition result.

[0086] The present invention constructs a flame and smoke recognition model based on the CenterNet network. Through the flame and smoke recognition model, automatic and accurate recognition of flames and smoke can be achieved, which not only reduces the dependence on manual monitoring, but also reduces labor costs and the intensity of monitoring work. At the same time, it can also reduce false alarms and missed alarms, and improve the reliability of the monitoring system.

[0087] Please refer to Figure 10 , another embodiment of the flame and smoke recognition device in the distribution center of the embodiment of the present invention includes:

[0088] The annotation and classification module 10 is used to collect image samples from scenes with flames and smoke, perform annotation and classification processing on the collected image samples, and obtain an image sample data set;

[0089] The data division module 20 is used to divide the image sample data set into a training set and a test set according to a predetermined ratio. Both the training set and the test set include flame image samples, smoke image samples, and normal image samples;

[0090] The model construction module 30 is used to construct an initial recognition model based on the CenterNet network;

[0091] The training module 40 is used to train the initial recognition model through the training set, adjust the parameters of the initial recognition model, and obtain a flame and smoke recognition model after testing by the test set;

[0092] The recognition module 50 is used to input the image of the area of the distribution center to be recognized obtained in real time into the flame and smoke recognition model, output the recognition result, and determine whether to issue an alarm according to the recognition result.

[0093] In this embodiment, the annotation and classification module 10 includes:

[0094] The amplification unit 11 is used to perform data amplification processing on the image samples to obtain expanded picture samples;

[0095] The classification unit 12 is used to perform annotation and classification processing on the expanded picture samples to obtain an image sample data set.

[0096] In this embodiment, the amplification unit 11 includes:

[0097] The deformation subunit 111 is used to perform mirroring, rotation, scaling, cropping, translation, or Gaussian noise processing on the image samples to obtain deformed picture samples;

[0098] The transformation subunit 112 is used to perform Cutout processing and / or Mix-up processing on the image samples and the deformed picture samples to obtain expanded picture samples.

[0099] In this embodiment, the classification unit 12 includes:

[0100] An annotation subunit 121, configured to perform feature annotation on the flames and smoke existing in the augmented picture samples through the Labelme tool, so as to obtain image samples annotated with feature information;

[0101] A classification subunit 122, configured to classify the image samples according to the feature information to obtain flame image samples, smoke image samples, and normal image samples, and the flame image samples, smoke image samples, and normal image samples constitute the image sample dataset.

[0102] In this embodiment, the model construction module 30 includes:

[0103] A building unit 31, configured to build a centerNet network framework, where the centerNet network framework includes a backbone network, a feature extraction layer, and a detection head. Among them, the detection head includes three branches of a heat map, an offset, and a size, which are respectively used to predict the target center point, the center point offset, and the target size;

[0104] An introduction unit 32, configured to introduce a multi-scale processing mechanism into the centerNet network framework to detect targets at different scales, improve the detection ability for small-scale and large-scale targets, and thus build an initial recognition model.

[0105] In this embodiment, the introduction unit 32 includes:

[0106] A small convolution kernel processing subunit 321, configured to perform convolution processing on a higher-level feature map by using a smaller 3x3 convolution kernel in the feature extraction layer of the CenterNet network framework to capture the fine features of small-scale targets;

[0107] A large convolution kernel processing subunit 322, configured to perform convolution processing on a lower-level feature map by using larger 5x5 and / or 7x7 convolution kernels to obtain the global features of large-scale targets;

[0108] A fusion subunit 323, configured to fuse features of different levels and different scales through a top-down path and lateral connections, so as to introduce a multi-scale processing mechanism into the centerNet network framework.

[0109] In this embodiment, the introduction unit 32 may alternatively include:

[0110] A sampling subunit 321', configured to construct a multi-scale feature representation by adopting multiple downsampling and upsampling operations through an Hourglass network structure;

[0111] The skip connection sub-unit 322' is used to set skip connections in the Hourglass network, fuse feature maps at different levels during the downsampling process with the feature maps at the corresponding positions during the upsampling, so as to introduce a multi-scale processing mechanism into the centerNet network framework.

[0112] In this embodiment, the training module 40 includes:

[0113] The initialization unit 41 is used to assign initial values to all parameters in the initial recognition model using the random initialization method;

[0114] The calculation unit 42 is used to randomly select picture samples from the training set and input them into the initial recognition model for training, calculate the prediction result through forward propagation, and then calculate the loss value according to the Focal Loss function;

[0115] The update unit 43 is used to update the parameters of the network using the backpropagation algorithm, so that the loss value gradually decreases;

[0116] The repetition unit 44 is used to repeat the above training steps until a predetermined number of training rounds is reached to obtain a trained initial recognition model;

[0117] The evaluation unit 45 is used to evaluate the trained initial recognition model using the test set, and adjust the hyperparameters according to the evaluation results to obtain an optimized flame and smoke recognition model.

[0118] In this embodiment, the loss calculation unit 43 includes: a classification loss calculation sub-unit 431, which uses focal loss as the classification loss function, measures the difference between the class probability predicted by the initial recognition model and the true class through the classification loss function. The calculation formula of Focal loss is as follows: L_cls = -α(1 - p_t)^γlog(p_t), where p_t is the class probability predicted by the initial recognition model, and α and γ are hyperparameters used to adjust the weights of different difficult and easy samples. When p_t is close to 1, (1 - p_t)^γ will become smaller, making the loss value of this sample smaller; conversely, when p_t is close to 0, (1 - p_t)^γ will become larger, making the loss value of this sample larger;

[0119] A regression loss calculation sub-unit 432 is configured to use the smooth L1 loss as a regression loss function to measure the difference between the position of the bounding box predicted by the initial recognition model and the position of the true bounding box through the regression loss function. The calculation formula of the SmoothL1 loss is as follows: L_reg = 0.5 * (IoU - 1)^2 * 1 if |IoU - 1| < 1, otherwise L_reg = |IoU - 1| - 0.5, where IoU represents the IOU value of the predicted bounding box and the true bounding box. When the IoU value is close to 1, it means that the predicted box is very close to the true box, and the loss value is small at this time; when the IoU value deviates from 1, the loss value will gradually increase.

[0120] Above Figure 9 And Figure 10 The sorting center flame and smoke recognition device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. Next, the sorting center flame and smoke recognition device in the embodiment of the present invention will be described in detail from the perspective of hardware processing.

[0121] Figure 11 FIG. is a schematic structural diagram of a sorting center flame and smoke recognition device provided by an embodiment of the present invention. The sorting center flame and smoke recognition device 100 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 11 (for example, one or more processors) and a memory 12, and one or more storage media 13 for storing application programs 133 or data 132 (for example, one or more mass storage devices). Among them, the memory 12 and the storage media 13 may be transient storage or persistent storage. The program stored in the storage media 13 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the sorting center flame and smoke recognition device 100. Further, the processor 11 may be configured to communicate with the storage media 13 and execute a series of instruction operations in the storage media 13 on the sorting center flame and smoke recognition device 100.

[0122] The sorting center flame and smoke recognition device 100 may further include one or more power supplies 14, one or more wired or wireless network interfaces 15, one or more input / output interfaces 16, and / or one or more operating systems 131, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that Figure 11 The device structure shown does not limit the sorting center flame and smoke recognition device 100, and may include more or fewer components than shown, or combine some components, or have different component arrangements.

[0123] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the method for flame and smoke recognition in a sorting center.

[0124] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, or unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0125] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0126] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.

Claims

1. A flame and smoke identification method for a distribution center, characterized in that: The distribution center flame and smoke identification method comprises: Collect image samples from scenes with flames and smoke, annotate and classify the collected image samples, and obtain an image sample data set; Dividing the image sample data set into a training set and a test set according to a predetermined ratio, wherein the training set and the test set both include flame image samples, smoke image samples and normal image samples; Build an initial recognition model based on the centerNet network; The initial recognition model is trained by the training set, the parameters of the initial recognition model are adjusted, and after being tested by the test set, a flame and smoke recognition model is obtained; The image of the distribution center area to be identified acquired in real time is input into the flame and smoke identification model, and the identification result is output, and it is determined whether to issue an alarm according to the identification result.

2. The flame and smoke identification method of the distribution center according to claim 1 is characterized in that: The collected image samples are labeled and classified to obtain an image sample data set, including the following steps: Performing data augmentation processing on the image sample to obtain an expanded image sample; The expanded image samples are labeled and classified to obtain an image sample data set.

3. The flame and smoke identification method of the distribution center according to claim 2, characterized in that: Performing data augmentation processing on the image sample to obtain an expanded image sample includes the following steps: Mirroring, rotating, scaling, cropping, translating or Gaussian noise processing are performed on the image sample to obtain a deformed image sample; Cutout processing and / or Mix-up processing are performed on the image samples and the deformed image samples to obtain expanded image samples.

4. The flame and smoke identification method of the distribution center according to claim 2 is characterized in that: The expanded image samples are labeled and classified to obtain an image sample data set, including the steps of: Use the Labelme tool to label the flames and smoke in the expanded image samples to obtain image samples labeled with feature information; The image samples are classified according to the feature information to obtain flame image samples, smoke image samples and normal image samples, and the flame image samples, smoke image samples and normal image samples constitute the image sample data set.

5. The flame and smoke identification method of the distribution center according to claim 1, characterized in that: Building an initial recognition model based on the centerNet network includes the following steps: Build a centerNet network framework, which includes a backbone network, a feature extraction layer and a detection head. The detection head includes three branches: heat map, offset and size, which are used to predict the target center point, center point offset and target size respectively. A multi-scale processing mechanism is introduced into the centerNet network framework to detect targets at different scales, improve the detection capabilities of small-scale and large-scale targets, and thus build an initial recognition model.

6. The flame and smoke identification method of the distribution center according to claim 1, characterized in that: A multi-scale processing mechanism is introduced into the centerNet network framework, including the following steps: In the feature extraction layer of the CenterNet network framework, a smaller 3x3 convolution kernel is used to convolve the higher-level feature maps to capture the fine features of small-scale targets. Use larger 5x5 and / or 7x7 convolution kernels to convolve lower-level feature maps to obtain global features of large-scale objects; Through top-down paths and lateral connections, features at different levels and scales are fused to introduce a multi-scale processing mechanism into the centerNet network framework.

7. The flame and smoke identification method of the distribution center according to claim 1, characterized in that: A multi-scale processing mechanism is introduced into the centerNet network framework, including the steps The Hourglass network structure uses multiple downsampling and upsampling operations to construct multi-scale feature representation; A skip connection is set in the Hourglass network structure to fuse the feature maps of different levels in the downsampling process with the feature maps of the corresponding positions in the upsampling process, thereby introducing a multi-scale processing mechanism in the centerNet network framework.

8. A flame and smoke identification device for a distribution center, characterized in that: include: The labeling and classification module is used to collect image samples from scenes with flames and smoke, label and classify the collected image samples, and obtain an image sample data set; A data partitioning module, used for partitioning the image sample data set into a training set and a test set according to a predetermined ratio, wherein the training set and the test set both include flame image samples, smoke image samples and normal image samples; Model building module, used to build the initial recognition model based on the centerNet network; A training module, used to train the initial recognition model through the training set, adjust the parameters of the initial recognition model, and obtain a flame and smoke recognition model after testing with the test set; The recognition module is used to input the real-time acquired image of the distribution center area to be recognized into the flame and smoke recognition model, output the recognition result, and determine whether to issue an alarm according to the recognition result.

9. A flame and smoke identification device for a distribution center, characterized in that: comprising a memory and at least one processor, wherein the memory has computer-readable instructions stored therein; The at least one processor calls the computer-readable instructions in the memory to execute the various steps of the distribution center flame and smoke identification method as claimed in any one of claims 1-7.

10. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the steps of the distribution center flame and smoke identification method as claimed in any one of claims 1 to 7 are implemented.