Risk factor identification method for disaster patrol image data
By using the YOLOv8 model to build a network and design a loss function, the problem of insufficient identification of risk factors in disaster inspection image data in the existing technology is solved, and high-precision and real-time identification of risk factors is achieved.
Patent Information
- Application Number
- CN202510209942.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-27
AI Technical Summary
The existing YOLO model lacks the accuracy of risk factor identification in disaster inspection image data, especially in complex backgrounds and low resolutions, which are difficult to meet high requirements.
YOLOv8 is used as the basic framework for risk factor identification tasks, a network model is built, and a loss function is designed to identify risk factors on real-time image data through training models.
The model's identification accuracy of risk factors in disaster inspection images is improved, computing efficiency and memory usage are optimized, real-time identification of risk factors is realized, and it can respond quickly when disaster situations occur.
Smart Images

Figure CN120219931A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of disaster risk assessment, and particularly to a method for identifying risk factors in disaster inspection image data. Background Art
[0002] The automatic analysis of disaster inspection image data and the identification of risk factors are one of the key technologies in disaster management and emergency response systems. Researchers have proposed various methods for the problem of risk factor identification.
[0003] In traditional algorithms, disaster inspection relies on manual analysis of image data to identify possible risk factors, such as damaged areas in the disaster area, road obstacles, and crowded situations of people, etc. This manual method is not only inefficient but also limited by the experience and accuracy of analysts, and it is impossible to quickly identify risks in real-time disasters, with relatively large limitations. In recent years, the successful application of deep learning algorithms such as convolutional neural networks (CNNs) in image processing has made it gradually possible to automate disaster risk identification. Risk identification methods based on deep learning mostly rely on classic YOLO series object detection models. Using unmanned aerial vehicles (UAVs) equipped with cameras for rapid post-disaster inspections and combining image processing technologies to automatically identify risk factors have become research hotspots. These models achieve a preliminary analysis of disaster inspection images by classifying and locating different objects in the images, or use semantic segmentation technology to identify disaster areas, etc.
[0004] However, due to the often complex background, low resolution, and relatively complex disaster scenarios of disaster inspection image data, the detection accuracy of existing YOLO models often fails to meet high requirements and does not fully combine the scene characteristics of disaster inspection images, resulting in poor identification effects of risk factors in some special environments. Summary of the Invention
[0005] The main purpose of the present invention is to provide a method for identifying risk factors in disaster inspection image data, which can improve the identification accuracy of the model for risk factors in disaster inspection images.
[0006] To achieve the above object, the first aspect of the present application provides a method for identifying risk factors in disaster inspection image data, the method comprising:
[0007] Obtain image data of different risk factors and perform data preprocessing to obtain original image data;
[0008] Label the risk factors in each image of the original image data to obtain sample image data;
[0009] Use YOLOv8 as the basic framework for the risk factor identification task, build a network model, and design the loss function of the network model;
[0010] Train the network model using the sample image data to obtain a trained risk factor identification model;
[0011] Identify risk factors for real-time image data based on the trained risk factor identification model.
[0012] Optionally, the obtaining of the image data of different risk factors includes:
[0013] Collect the image data of different risk factors by means of drone shooting, and the image data covers different time periods and weather conditions; the categories of the risk factors include, but are not limited to, any one or several of the following: geological changes, waterlogging, traffic damage and blockage.
[0014] Optionally, the data preprocessing includes:
[0015] Convert the category identifier of the image data into a format recognizable by the network model, and normalize the size of the image data.
[0016] Optionally, the annotation of the risk factors in each image of the original image data includes:
[0017] Perform target box annotation on each image in the original image data to generate a txt file corresponding to the image, and the txt file stores the category of the target risk area in the image and the basic image information, and the basic image information includes the center coordinates of the target risk area, the width and height of the target risk area.
[0018] Optionally, the method further includes:
[0019] If the target risk area is occluded, if the occluded area of the target risk area is not greater than half of the overall area of the target risk area, perform annotation; if the occluded area of the target risk area is greater than half of the overall area of the target risk area, abandon annotation.
[0020] Optionally, some convolutional modules are introduced into the network model to synchronously reduce the computational complexity and memory access cost;
[0021] The number of floating-point operations of the partial convolutional module is shown in the following formula:
[0022]
[0023] where h and w respectively represent the height and width of the feature map, k represents the size of the convolutional kernel, and c p represents the number of channels participating in the convolution;
[0024] The memory access pattern of the partial convolution module is shown in the following formula:
[0025]
[0026] Optionally, the loss function of the network model includes a detection box loss and a classification loss;
[0027] The detection box loss adopts a complete cross-union loss function;
[0028] The classification loss adopts a VFL loss function
[0029] The second aspect of the present application provides a risk factor identification device for disaster inspection image data, including:
[0030] An acquisition module, configured to acquire image data of different risk factors;
[0031] A preprocessing module, configured to perform data preprocessing on the image data to obtain original image data;
[0032] A labeling module, configured to label the risk factors in each image of the original image data to obtain sample image data;
[0033] A model building module, configured to use YOLOv8 as the basic framework for the risk factor identification task, build a network model, and design the loss function of the network model;
[0034] A model training module, configured to train the network model with the sample image data to obtain a trained risk factor identification model;
[0035] An identification module, configured to identify risk factors in real-time image data based on the trained risk factor identification model.
[0036] The third aspect of the present application provides an electronic device, including a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor is caused to execute the steps of the first aspect and any possible implementation manner thereof.
[0037] The fourth aspect of the present application provides a computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, the processor is caused to execute the respective steps in the method described in the first aspect.
[0038] The present application provides a method for identifying risk factors in disaster inspection image data. By obtaining image data of different risk factors and performing data preprocessing, the original image data is obtained; the risk factors in each image of the original image data are labeled to obtain sample image data; YOLOv8 is used as the basic framework for the risk factor identification task to build a network model and design the loss function of the network model; the sample image data is used to train the network model to obtain a trained risk factor identification model; the trained risk factor identification model is used to identify risk factors in real-time image data; this method can improve the identification accuracy of the model for risk factors in disaster inspection images. At the same time, by optimizing the computational efficiency and memory occupancy of the model, the real-time performance of risk factor identification is improved, enabling the system to respond quickly when a disaster occurs. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0040] Among them:
[0041] Figure 1 is a schematic flowchart of a method for identifying risk factors in disaster inspection image data provided by an embodiment of the present application;
[0042] Figure 2 is a schematic diagram of disaster inspection image data provided by an embodiment of the present application;
[0043] Figure 3 is a schematic diagram of the structure of YOLOv8 provided by an embodiment of the present application;
[0044] Figure 4 is a schematic diagram of a partial convolution module provided by an embodiment of the present application;
[0045] Figure 5 is a schematic diagram of the qualitative results of the model performance provided by an embodiment of the present application;
[0046] Figure 6 is a schematic diagram of the structure of a device for identifying risk factors in disaster inspection image data provided by an embodiment of the present application;
[0047] Figure 7 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without making creative efforts belong to the scope of protection of this application.
[0049] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0050] Referring to "embodiments" in this context means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0051] The YOLO (You Only Look Once) algorithm involved in the embodiments of this application was proposed by scholars such as Joseph Redmon in 2016 and is a revolutionary real-time object detection algorithm. It uses CNN as the basic framework and includes multiple convolutional layers and fully connected layers. In the YOLO v1 version, the network structure was inspired by GoogLeNet and includes 24 convolutional layers and 2 fully connected layers. YOLO divides the input image into an S×S grid, and each grid cell is responsible for detecting the objects whose center points fall within that cell. This strategy simplifies the object localization process and allows the model to quickly localize and classify multiple objects.
[0052] Specifically, the YOLO algorithm divides the input image into multiple grid cells, and each grid cell predicts a certain number of bounding boxes and class probabilities. These bounding boxes contain information such as the center coordinates, width, height, and confidence of the target. The confidence indicates whether the target object is contained in the bounding box, while the class probability indicates which class the target belongs to. The final detection result is obtained by screening and filtering based on the confidence and class probability among all grid cells and bounding boxes. With the continuous development of technology, the YOLO algorithm has also undergone iterations and optimizations in multiple versions. These iterations and optimizations of the versions have enabled the YOLO algorithm to maintain a leading position in the field of object detection and provide strong technical support for related applications. The YOLO algorithm has been widely applied in fields such as object detection, intelligent monitoring, and autonomous driving.
[0053] In the embodiments of this application, the partial convolution mainly utilizes the redundant information in the feature map, only performs convolution operations on some input channels, and keeps other channels unchanged. This can reduce computational redundancy and the number of memory accesses, thereby improving the computational speed. Its core lies in its unique convolution operation, which combines standard convolution and a masking mechanism. Based on the input feature map and the corresponding masking map, this module only performs convolution calculations on the unmasked (or valid) pixels. This mechanism ensures that the convolution operation focuses on the valid region and reduces the impact of the invalid region on the output.
[0054] The partial convolution operation first determines the validity of each pixel according to the masking map. For pixels with a masking value of 1 (i.e., valid pixels), they participate in the convolution calculation; for pixels with a masking value of 0 (i.e., invalid pixels), they are ignored during convolution. To maintain the normalization of the convolution result, the partial convolution introduces a scaling factor, which is equal to the ratio of the number of valid pixels to the total number of pixels covered by the convolution kernel. This design ensures that even in the presence of a mask, the output of the convolution operation can maintain a stable numerical range. In addition, the partial convolution module usually also contains an activation function (such as ReLU) and a batch normalization layer (Batch Normalization) to enhance the model's non-linear expression ability and stability. The activation function introduces non-linearity, enabling the model to learn more complex feature representations; while the batch normalization layer helps to accelerate the training process and improve the generalization ability of the model. Therefore, the partial convolution module is competitive with existing operators in terms of speed and accuracy, providing strong support for the practical application of the partial convolution module. This is particularly useful when dealing with images with irregular shapes or missing data, such as tasks like image inpainting, image denoising, and image segmentation. It can also be used for edge detection. By designing a specific convolution kernel, the edge information in the image can be highlighted, thereby achieving precise edge detection.
[0055] The embodiments of this application will be described below with reference to the accompanying drawings in the embodiments of this application.
[0056] Please refer to Figure 1 , which is a schematic flowchart of a method for identifying risk factors of disaster inspection image data provided by an embodiment of this application. As Figure 1 shown, this method includes:
[0057] 101. Obtain image data of different risk factors, and perform data preprocessing to obtain original impact data.
[0058] In the embodiment of this application, the execution subject of the method can be a device for identifying risk factors of disaster inspection image data, which can be implemented by a terminal device in actual application, such as a server or a computer.
[0059] In an optional implementation manner, the obtaining of the image data of different risk factors includes:
[0060] Collect the image data of the above different risk factors through the method of drone shooting. The above image data covers different time periods and weather conditions; the categories of the above risk factors include, but are not limited to, any one or more of the following: geological changes, water accumulation, traffic damage and blockage.
[0061] In the disaster inspection work, data collection is the primary link. In the embodiment of this application, the image data covering typical risk factors such as geological changes, water accumulation, traffic damage and blockage can be widely collected mainly through the method of drone shooting. At the same time, ensure that the data covers different time periods and weather conditions to comprehensively reflect the real situation of the disaster site.
[0062] Figure 2 This is a schematic diagram of the disaster inspection image data provided by an embodiment of this application.
[0063] Specifically, the drone can be deployed above the disaster site and automatically cruise according to the preset flight path and altitude parameters. The high-definition camera carried by the drone takes image data of key risk factors such as geological changes, water accumulation, traffic damage and blockage at a predetermined time interval and under weather conditions. Transmit the taken image data to the ground. Some image data, such as Figure 2 shown, can be color images in actual application.
[0064] The collected original image data has various formats and different resolutions, so preprocessing is required.
[0065] In an optional implementation manner, the above data preprocessing includes:
[0066] Convert the category identifier of the above image data into a format recognizable by the above network model, and perform normalization processing on the size of the above image data.
[0067] The data preprocessing steps in the embodiments of this application may include converting the image data into a unified format and adjusting it to the same resolution. In addition, image enhancement processing can be performed, such as adjusting brightness, contrast, rotation, and scaling, to increase the diversity and robustness of the data and improve the model's ability to recognize disaster features.
[0068] In the embodiments of this application, a data preprocessing module is set up to perform format conversion and adjustment on the labeled objects in the dataset. Specifically, for each target object, detailed information extraction and formatting processing are carried out, including converting its class identifier (cls_id) into a format recognizable by the model, and normalizing key information such as the center point position (x, y) and the width w and height h of the bounding box, so as to ensure that this data can be seamlessly connected to the model and provide direct and efficient support for the subsequent processing of the model. On this basis, the dataset is further divided into a training set and a test set to ensure that the model can obtain fair and effective data support during both the training and evaluation phases.
[0069] Optionally, in order to further improve the generalization ability of the model, a variety of data augmentation techniques are introduced in the data loading link, such as classic methods like image scaling, random cropping, horizontal or vertical flipping, and the Mosaic data augmentation technique is introduced, with the expectation of enriching the data diversity and providing more comprehensive and in-depth support for the model's learning and reasoning. Especially Mosaic augmentation, which involves randomly selecting 4 images and generating a composite image containing various elements for training through cropping, splicing, etc. This method greatly enriches the diversity of training data and significantly improves the model's ability to handle complex scenarios.
[0070] 102. Label the risk factors in each image of the above original image data to obtain sample image data.
[0071] Specifically, in this labeling task, it is necessary to accurately label the risk factors in each image, including the positions and categories of important targets such as geological changes, waterlogging, and traffic damage and blockages. After the labeling is completed, the labeled data also needs to be sorted out for subsequent model training. The labeled data is converted into a format suitable for the training of this model, such as converting it into XML, CSV, or TXT files when using the YOLOv8 model, thus ensuring the compatibility and readability of the data.
[0072] In an alternative embodiment, the labeling of the risk factors in each image of the above original image data includes:
[0073] Perform target box annotation on each image in the above original image data, generate a txt file corresponding to the above image, and the above txt file stores the category of the target risk area in the above image and the basic image information, and the above basic image information includes the center coordinates of the above target risk area, the width and height of the above target risk area.
[0074] Specifically, in the embodiment of the present application, the open-source software LabelImg annotation tool can be selected to annotate the target boxes of the data set, and generate a txt annotation file with the same name as the image. Among them, the generated txt file can store the detection category (category of the target risk area) of the annotated image and various information, where cls_id can be used to represent the category number, x and y represent the center coordinates of the target, and w and h represent the width and height of the target.
[0075] Further optionally, the above method further includes:
[0076] If the above target risk area is occluded, if the occluded area of the above target risk area is not greater than half of the overall area of the above target risk area, annotation is performed; if the occluded area of the above target risk area is greater than half of the overall area of the above target risk area, annotation is abandoned.
[0077] Specifically, since the data set comes from multiple scenarios and time periods, sometimes the target will be occluded by factors such as trees, houses, or lighting angles. In this case, if the occluded area of the target is less than half of the overall target area, annotation will still be performed; but if the occluded area is greater than half of the overall target area, annotation is abandoned. In practical applications, the judgment logic can be modified and adjusted as needed. For example, annotation is only performed when the occluded area of the target is greater than 2 / 3 of the overall target area. The embodiment of the present application does not limit this.
[0078] 103. Use YOLOv8 as the basic framework for the risk factor identification task, build a network model, and design the loss function of the above network model.
[0079] In the embodiment of the present application, YOLOv8 is selected as the basic framework for the risk factor identification task, based on the following key points: First, YOLOv8 is the latest iteration based on the widely recognized YOLO series of models, and this series of models is famous for its excellent real-time performance and high precision. YOLOv8 inherits the advantages of this series and achieves higher efficiency and accuracy through further optimization. For images such as disaster inspections, YOLOv8 significantly improves the model's recognition ability through improved feature extraction and fusion strategies. In addition, the architecture design of YOLOv8 takes into account flexibility and scalability, making it possible to customize and adjust for specific applications.
[0080] Figure 3 A schematic diagram of the structure of YOLOv8 provided by an embodiment of the present application.
[0081] As Figure 3 shown, the YOLOv8 network structure mainly consists of three parts: Backbone (main network), Neck (neck network), and Head (head network). After the image is input into the Backbone (main network), it passes through multiple convolutional (Conv) layers and C2f modules in sequence to achieve the initial extraction and abstraction of the input image features. The SPPF module is used at the end of the main network to further fuse global features and enhance the network's ability to capture risk features. The Neck (neck network) includes convolutional (Conv) layers, C2f modules, upsampling (Upsample), and concatenation (Concat) operations. Through the upsampling operation, the deep low-resolution feature map is restored to a higher resolution and concatenated with the shallow feature map to achieve the fusion of features at different levels and improve the detection accuracy. The Head (head network) outputs the final detection task. The data after feature fusion in the Neck (neck network) is input into the Detect (detection) module, and the detection module outputs the detection results of relevant risk factors, including the category and location of the target (such as the bounding box). Figure 3 The example in the lower right corner of [Figure] shows the detected "pothole" and its confidence score.
[0082] To reduce the computational complexity and improve the computational efficiency, an embodiment of the present application also introduces a partial convolution module in YOLOv8 to synchronously reduce the computational complexity and memory access cost.
[0083] In an embodiment of the present application, in view of the characteristics of these images, the YOLOv8 model is specifically optimized. Partial convolution is introduced, and the basic structure of the original model is retained, including the input layer, feature extraction layer, and detection layer. Among them, the input layer is responsible for receiving image data, the feature extraction layer extracts image features through structures such as convolutional layers, and the detection layer is responsible for target detection and classification based on the extracted features. Through this series of improvements and optimizations, the detection performance and accuracy of the model on the disaster inspection image data can be improved, and the computational complexity can be reduced.
[0084] Figure 4 A schematic diagram of a partial convolution module provided by an embodiment of the present application.
[0085] Specifically, partial convolution selectively performs standard convolution operations on a subset of the input channels to extract spatial features, while the remaining channels remain unchanged. To ensure smooth and orderly memory access, the first or last group of consecutive c in the feature map is selected pOne channel serves as the reference sample for calculation. This design not only effectively reduces the computational complexity but also optimizes the memory access efficiency, thus significantly improving the overall operation performance. The number of floating-point operations of partial convolution is shown in Equation (1):
[0086] Flops = h × w × k 2 × c p 2 (1)
[0087] Wherein, the height and width of the feature map are represented by h and w respectively, the size of the convolution kernel is represented by k, and the number of channels participating in the convolution is represented by c p to represent. In actual operation, usually set r = c p / c = 1 / 4, which results in the number of floating-point operations (FLOPs) of partial convolution being only 1 / 16 of that of traditional convolution.
[0088] The design of partial convolution also optimizes the memory access pattern. By selecting continuous channels for calculation, partial convolution can make full use of the cache mechanism in modern hardware architectures, reducing the randomness and latency of memory access, thereby further improving the calculation efficiency. The memory access pattern of partial convolution can be expressed as Equation (2):
[0089]
[0090] The memory access amount required for partial convolution is only 1 / 4 of that of traditional convolution, because the remaining c - c p channels do not participate in the calculation, so there is no need to access their memory.
[0091] In an alternative embodiment, the loss function of the above network model includes detection box loss and classification loss;
[0092] The above detection box loss adopts the complete cross-union loss function;
[0093] The above classification loss adopts the VFL loss function.
[0094] In the embodiments of the present application, when designing the loss function, the YOLO series composite loss function suitable for object detection is selected. This function covers the detection box loss and classification loss, and can comprehensively evaluate the accuracy of the model in target position, existence probability, and category judgment. For the improved model characteristics, the loss function is optimized. By adjusting the weights of each loss, the model pays more attention to the accurate recognition of the disaster target position and category during the training process, thereby improving the training effect and recognition accuracy of the model. This optimization strategy helps the model achieve more efficient and accurate object detection in disaster inspection image data.
[0095] Specifically, the loss function designed by the optimized and improved YOLOv8 algorithm model in the embodiments of this application includes two branches: detection box loss and classification loss.
[0096] To optimize bounding box regression, the improved YOLOv8 algorithm model adopts the Complete-IoU (CIoU) loss function, which is particularly prominent in adjusting the similarity between the predicted box and the actual box. The CIoU loss function provides an efficient method for comprehensively evaluating box similarity by considering the overlap degree (Intersection over Union, IoU) between bounding boxes, the distance between the center points, and the aspect ratio difference.
[0097] The CIoU loss function mentioned in the embodiments of this application is an improved loss function for bounding box regression in object detection. It is extended based on the traditional IoU (Intersection over Union) loss function to address some limitations of the IoU loss function. When the predicted box and the ground truth box do not overlap at all, the gradient of the IoU loss function is zero, which makes it difficult for the model to adjust the predicted box to be close to the ground truth box during the early training stage. In addition, IoU only measures the area of the overlapping region and does not reflect the proximity of the two boxes. Even if the distances between two predicted boxes and the ground truth box are very different, as long as they do not overlap, the IoU value is zero.
[0098] The definition of the CIoU loss function is as follows:
[0099]
[0100] Among them, IoU represents the intersection over union, which measures the overlap degree between the predicted box and the ground truth box; ρ(b, b gt ) is the Euclidean distance between the predicted bounding box b and the ground truth bounding box b gt , representing the distance between their center points;
[0101] C is the diagonal length of the smallest closed rectangle (i.e., the smallest bounding rectangle) that contains the predicted bounding box and the ground truth bounding box;
[0102] α is the weight coefficient used to balance the influence of the aspect ratio term v;
[0103] v is the aspect ratio consistency, which measures the consistency of the aspect ratios of the predicted bounding box and the ground truth bounding box.
[0104] For the classification loss, the VFL loss function is adopted. This loss function is developed on the basis of FocalLoss, aiming to solve the problem of imbalance between positive and negative samples in object detection problems, and further enhance the attention to the IoU index to improve the performance of dense object detectors. FocalLoss solves the problem that easy-to-classify samples contribute too much to the loss function by introducing the γ parameter, making the model pay more attention to difficult samples. The operation principle of the VFL loss function is as shown in Equation (4):
[0105]
[0106] where p is the probability of the target class. When it is a positive sample, q is the IoU value between the predicted box and the ground truth box, and when it is a negative sample, q = 0. It can be seen from Equation (4) that when the sample is a positive sample, the ordinary loss function is used instead of FocalLoss, but there is an additional adaptive IoU weighting to highlight the main sample (within the positive samples, the main sample has the largest IOU value); when the sample is a negative sample, FocalLoss is used. It can be seen that VFL asymmetrically weights positive and negative samples to highlight the main sample.
[0107] 104. Use the above sample image data to train the above network model to obtain a trained risk factor identification model.
[0108] In one implementation, during the model training and validation process, the following environmental parameter configurations are adopted:
[0109] Adopt the pytorch deep learning framework, adopt the end-to-end training method, and optimize with the designed loss function as the target. During the training process, it is executed on a GeForce RTX 3090 (24GB), the batch-size is set to 16, a total of 150 epochs are carried out, the learning rate is 0.001, and the model parameters are gradually adjusted to enable the model to better identify the disaster types. Specifically, the network training parameters and methods can be adjusted according to needs, and the embodiments of the present application do not limit this.
[0110] 105. Based on the above trained risk factor identification model, identify risk factors for real-time image data.
[0111] After the model training is completed, verification and application can be carried out. In the embodiments of the present application, the images output by the model show that this method can automatically identify each risk factor and mark the approximate location.
[0112] In one implementation, quantitative evaluation of the model performance can be carried out. Specifically, the trained model is evaluated on the test dataset, and the following indicators can be mainly used:
[0113] Param: The number of parameters of the improved network model is 3.37M, indicating that the model can improve the computing efficiency and achieve real-time performance.
[0114] Flops: The Flops of the improved network model is 56G, indicating that the network model requires less computation during the inference process, thus reducing the computation time.
[0115] Precision: The best Precision value is 87.3%, indicating that the improved network can better identify the disaster risk factors and determine which type of risk factors.
[0116] Recall: The best Recall value is 85.2%. This indicates that the proportion of the factors correctly identified as disaster risk factors by the improved model is high, and there is no misjudgment.
[0117] mAP (mean Average Precision): The best mAP value is 84.2%, indicating that the improved network model has good recognition performance and can automatically identify the disaster risk.
[0118] Figure 5 This is a schematic diagram of the qualitative results of the model performance provided by the embodiment of the present application. Figure 5 The shown images display the qualitative experimental results of the model on the test dataset. Specifically:
[0119] The first column shows the images of the disaster risks, including the impacts of landslides, traffic jams, and large-area waterlogging on power facilities.
[0120] The second column shows the results after processing by the method proposed in the embodiment of the present application. These images show that the method can automatically identify each risk factor and mark its approximate location. Examples show the detected risk factors such as "pothole", "waterlogging", "traffic damage", etc. and their confidence scores.
[0121] In the field of disaster risk factor identification technology, traditional methods mainly rely on convolutional neural networks and other versions of the YOLO series. Although these methods have achieved preliminary identification of disaster images to a certain extent, in the face of complex and changing disaster scenarios, their detection accuracy and efficiency often fail to meet the high standards required for practical applications. In contrast, the YOLOv8 network model is introduced in the embodiments of this application, which can more deeply mine the key information in the images, thus achieving more accurate positioning and identification of disaster risk factors. YOLOv8 also focuses on the lightweight design of the model. By reducing network parameters and computational complexity, it has achieved a significant improvement in the running efficiency of the model while maintaining high-precision identification capabilities. For images such as disaster inspections, YOLOv8 has significantly improved the model's identification ability through improved feature extraction and fusion strategies.
[0122] In addition, most of the network models used in traditional methods for identifying risk factors in disaster inspection images adopt full convolution, resulting in large network overhead and high computational complexity, and unable to achieve real-time identification effects. In the embodiments of this application, by introducing partial convolution, unnecessary calculations are reduced. At the same time, combined with the detection ability of YOLOv8, the computational efficiency of the model has been significantly improved. Standard convolution operations are selectively performed on a subset of the input channels to extract spatial features, while the remaining channels remain unchanged. This not only effectively reduces computational complexity but also optimizes memory access efficiency, significantly improving the overall operation performance, thus achieving real-time effects. Partial convolution increases the flexibility of the model by selectively performing convolution operations on the input channels. This flexibility enables the model to better adapt to different disaster scenarios and image features, thereby improving the robustness and generalization ability of the model.
[0123] Based on the description of the foregoing method embodiments, the embodiments of this application also provide a device for identifying risk factors in disaster inspection image data.
[0124] Figure 6 It is a structural schematic diagram of a device for identifying risk factors in disaster inspection image data provided by the embodiments of this application. As Figure 6 shown, the device 600 for identifying risk factors in disaster inspection image data includes:
[0125] An acquisition module 610, configured to acquire image data of different risk factors;
[0126] A preprocessing module 620, configured to perform data preprocessing on the above-mentioned image data to obtain original image data;
[0127] A labeling module 630, configured to label the risk factors in each image of the above-mentioned original image data to obtain sample image data;
[0128] The model construction module 640 is used to build a network model with YOLOv8 as the basic framework for the risk factor identification task, and design the loss function of the above network model;
[0129] The model training module 650 is used to train the above network model with the above sample image data to obtain a trained risk factor identification model;
[0130] The identification module 660 is used to identify risk factors for real-time image data based on the above trained risk factor identification model.
[0131] It can be understood that the relevant content of each module involved in Figure 6 has been described in detail in the foregoing method embodiments. Specifically, reference can be made to the content in the method embodiments; that is Figure 6 The provided risk factor identification device 600 for disaster inspection image data can execute any step in the Figure 1 shown embodiments, which will not be elaborated here.
[0132] In an embodiment of the present application, an electronic device is also proposed. Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an electronic device provided in an embodiment of the present application. As Figure 7 shown, the electronic device 700 includes a processor 701 and a memory 702. The memory 702 stores a computer program. When the computer program is executed by the processor 701, it will execute any step in the Figure 1 shown method embodiments. The electronic device 700 may also include input / output devices, etc. In a specific implementation manner, the electronic device may be a terminal device, etc.
[0133] In an embodiment, a computer-readable storage medium is also proposed. The computer-readable storage medium stores a computer program. When the computer program is executed by the processor 701, it causes the processor 701 to execute any step in the above method embodiments.
[0134] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0135] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0136] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for identifying risk factors of disaster inspection image data, characterized in that: The method comprises: Obtain imaging data of different risk factors and perform data preprocessing to obtain original impact data; Annotating the risk factors in each image in the original image data to obtain sample image data; Use YOLOv8 as the basic framework for risk factor identification tasks, build a network model, and design the loss function of the network model; Using the sample image data to train the network model to obtain a trained risk factor identification model; Risk factors are identified on real-time image data based on the trained risk factor identification model.
2. The risk factor identification method for disaster inspection image data according to claim 1 is characterized in that: The obtaining of imaging data of different risk factors includes: Image data of the different risk factors are collected by drone photography, and the image data covers different time periods and weather conditions; the categories of the risk factors include but are not limited to any one or more of the following: geological changes, water accumulation, and traffic damage and blockage.
3. The risk factor identification method for disaster inspection image data according to claim 2 is characterized in that: The data preprocessing includes: The category identifier of the image data is converted into a format recognizable by the network model, and the size of the image data is normalized.
4. The risk factor identification method for disaster inspection image data according to claim 1 is characterized in that: The step of labeling the risk factors in each image in the original image data includes: A target box is annotated for each image in the original image data to generate a txt file corresponding to the image. The txt file stores the category of the target risk area in the image and basic image information. The basic image information includes the center coordinates of the target risk area and the width and height of the target risk area.
5. The method for identifying risk factors of disaster inspection image data according to claim 4 is characterized in that: The method further comprises: If the target risk area is blocked, if the area of the target risk area blocked is not larger than half of the overall area of the target risk area, it will be marked; if the area of the target risk area blocked is larger than half of the overall area of the target risk area, the marking will be abandoned.
6. The risk factor identification method for disaster inspection image data according to claim 1 is characterized in that: A partial convolution module is introduced into the network model to simultaneously reduce computational complexity and memory access cost; The number of floating-point operations of the partial convolution module is shown in the following formula: Among them, h and w represent the height and width of the feature map, k represents the size of the convolution kernel, and c p Indicates the number of channels involved in convolution; The memory access pattern of the partial convolution module is shown in the following formula:
7. The risk factor identification method for disaster inspection image data according to claim 1 is characterized in that: The loss function of the network model includes detection box loss and classification loss; The detection box loss adopts a complete cross-joint loss function; The classification loss adopts the VFL loss function.
8. A risk factor identification device for disaster inspection image data, characterized in that: include: An acquisition module, used to acquire imaging data of different risk factors; A preprocessing module, used for performing data preprocessing on the image data to obtain original impact data; A labeling module, used for labeling the risk factors in each image in the original image data to obtain sample image data; A model building module is used to build a network model using YOLOv8 as the basic framework for risk factor identification tasks and to design a loss function for the network model; A model training module, used to train the network model using the sample image data to obtain a trained risk factor identification model; The identification module is used to identify risk factors for real-time image data based on the trained risk factor identification model.
9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.