Image-based mining area dangerous behavior visual identification method and system

By applying the target fast area recursive convolutional neural network model and target behavior recognition model in the mining area, the dangerous behavior of mining area operators is solved, and more efficient risk behavior monitoring is achieved.

CN120183045APending Publication Date: 2025-06-20KAILUAN GRP MINING ENG CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510359376.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing image-based hazardous behavior recognition methods have low recognition accuracy under the influence of factors such as complex background, variable lighting and behavioral diversity in the mining area, making it difficult to meet the actual safety monitoring needs of the mining area.

Method used

The target fast area recursive convolutional neural network model is used to process the mining area images, identify the operators, and determine their dangerous behavior through the target behavior recognition model. Finally, the behavior is marked based on the risk level and probability of occurrence.

Benefits of technology

It improves the accuracy of visual identification of dangerous behaviors in mining areas, can detect dangerous behaviors in a timely and accurate manner, and effectively avoid accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183045A_ABST
    Figure CN120183045A_ABST
Patent Text Reader

Abstract

The invention provides an image-based mining area dangerous behavior visual recognition method and system, and belongs to the technical field of behavior recognition, and the method comprises the steps: processing a mining area image based on a target rapid region recursion convolutional neural network model, and obtaining a first target which is an operator; identifying the first target based on a target behavior identification model, and determining a dangerous behavior of the first target; and marking the dangerous behavior of the first target based on the dangerous level and the occurrence probability corresponding to the dangerous behavior. According to the image-based mining area dangerous behavior visual identification method and system provided by the invention, the accuracy of mining area dangerous behavior visual identification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of behavior recognition, and more specifically, relates to a method and system for visual recognition of dangerous behaviors in mining areas based on images. Background Art

[0002] In the mining operation environment, due to its complex and potentially dangerous characteristics, the safety of workers faces many threats. With the development of image processing and artificial intelligence technologies, using image recognition technology to identify and warn of dangerous behaviors in mining areas has become an effective way to improve the safety management level of mining areas. However, existing image-based dangerous behavior recognition methods have the problem of low recognition accuracy under the influence of factors such as complex backgrounds, variable lighting, and behavior diversity in mining areas, and it is difficult to meet the actual safety monitoring requirements of mining areas. Summary of the Invention

[0003] The purpose of the present disclosure is to provide a method and system for visual recognition of dangerous behaviors in mining areas based on images to improve the accuracy of visual recognition of dangerous behaviors in mining areas.

[0004] In the first aspect of the embodiments of the present disclosure, a method for visual recognition of dangerous behaviors in mining areas based on images is provided, including: Processing a mining area image based on a region-based fast convolutional neural network model for object detection to obtain a first object, where the first object is a worker; Identifying the first object based on an object behavior recognition model to determine the dangerous behavior of the first object; Labeling the dangerous behavior of the first object based on the risk level and occurrence probability corresponding to the dangerous behavior.

[0005] In the second aspect of the embodiments of the present disclosure, a system for visual recognition of dangerous behaviors in mining areas based on images is provided, including: An object localization module for processing a mining area image based on a region-based fast convolutional neural network model for object detection to obtain a first object, where the first object is a worker; A behavior recognition module for identifying the first object based on an object behavior recognition model to determine the dangerous behavior of the first object; A behavior labeling module for labeling the dangerous behavior of the first object based on the risk level and occurrence probability corresponding to the dangerous behavior.

[0006] In the third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned method for visual recognition of dangerous behaviors in mining areas based on images are implemented.

[0007] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned image-based visual recognition method for dangerous behaviors in a mining area are implemented.

[0008] The beneficial effects of the image-based visual recognition method and system for dangerous behaviors in a mining area provided by the embodiments of the present disclosure are as follows: The present disclosure processes mining area images using a region-based fast convolutional neural network model for targets, which can efficiently identify the first target, improving the accuracy and speed of target detection. Subsequently, the present disclosure identifies the dangerous behaviors of the first target through a target behavior recognition model, which can timely and accurately detect the dangerous behaviors of the first target and effectively avoid accidents. At the same time, the present disclosure labels the dangerous behaviors, enabling targeted identification of dangerous behaviors. Therefore, the present disclosure can improve the accuracy of visual recognition of dangerous behaviors in a mining area. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0010] Figure 1 It is a schematic flowchart of an image-based visual recognition method for dangerous behaviors in a mining area provided by an embodiment of the present disclosure; Figure 2 It is a structural block diagram of an image-based visual recognition system for dangerous behaviors in a mining area provided by an embodiment of the present disclosure; Figure 3 It is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0012] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments with reference to the drawings.

[0013] Please refer toFigure 1 , Figure 1 The flowchart of the image-based visual recognition method for dangerous behaviors in mining areas provided by an embodiment of the present disclosure. The method includes: S101: Process the mining area image based on the Faster Region-based Convolutional Neural Networks (Faster R-CNN) model to obtain a first target, where the first target is a worker.

[0014] In this embodiment, the Faster Region-based Convolutional Neural Networks (Faster R-CNN) model is an object detection model. The object detection model is used to find target objects in images or videos and determine the positions and categories of the target objects. The object detection model used in this embodiment is trained to be able to detect workers in mining area images. The mining area image is an image taken in a mining area environment, which may include the mining operation site, various operations performed by workers, and the areas where workers are located, etc., and can be used for subsequent identification of dangerous behaviors. The first target is the object to be detected by the object detection model from the mining area image. The target Faster R-CNN model is a trained Faster R-CNN model used to detect workers in mining area images.

[0015] Processing the mining area image based on the target Faster R-CNN model to obtain the first target includes: Input the mining area image into the backbone network to obtain a feature map; Process the feature map based on the Regional Proposal Network (RPN) to generate multiple candidate regions; each candidate region has a different size; Map the multiple candidate regions to the feature map to obtain multiple candidate region feature maps; Based on the Region of Interest (RoI) pooling layer, adjust each candidate region feature map to a feature vector of a fixed size; Based on the fully connected layer, perform deep feature learning on the feature vector, and use a classifier to determine whether there is a worker in the candidate region, and use a regressor to adjust the position of the bounding box of the worker to obtain the first target, that is, the worker, and the position of the worker.

[0016] The backbone network can be a lightweight network or a heavyweight network. The lightweight networks include Mobile Convolutional Neural Network (MobileNet), EfficientNet (EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks), etc. The heavyweight networks include Visual Geometry Group Network (VGGNet), Residual Network (ResNet), etc.

[0017] In this embodiment, determining the backbone network based on the detection accuracy includes: In response to the detection accuracy being greater than or equal to the first accuracy threshold, using the heavyweight network as the backbone network; In response to the detection accuracy being less than the first accuracy threshold, using the lightweight network as the backbone network; Among them, the computing resources required by the heavyweight network are greater than those required by the lightweight network.

[0018] Alternatively, determining the backbone network based on the detection efficiency includes: In response to the detection efficiency being greater than or equal to the first efficiency threshold, using the lightweight network as the backbone network; In response to the detection efficiency being less than the first efficiency threshold, using the heavyweight network as the backbone network; Among them, the operation time of the heavyweight network is greater than that of the lightweight network.

[0019] Both the first accuracy threshold and the first efficiency threshold are pre-set reference standard values, which can be set empirically according to the requirements of the actual scenario.

[0020] The above process of selecting the backbone network is based on actual needs and selection during the training process.

[0021] In this embodiment, processing the feature map based on the target RPN to generate multiple candidate regions includes: Dividing the feature map based on a sliding window to generate multiple anchor boxes, and the multiple anchor boxes are anchor boxes of different scales and ratios; Classifying each anchor box based on a convolutional neural network, determining whether the first target is included in each anchor box, and performing a regression operation to adjust the position and size of each anchor box; If the first target is included in the anchor box, using the anchor box and its location as a candidate region to obtain multiple candidate regions; Among them, the size of the sliding window can be set based on experience.

[0022] S102: Identify the first target based on the target behavior recognition model to determine the dangerous behavior of the first target.

[0023] In this embodiment, the target behavior recognition model is used to identify the model of the behavior exhibited by the first target, and can determine whether the behavior of the first target is a dangerous behavior. Dangerous behaviors may include operators not wearing safety helmets, operating equipment in violation of regulations, staying in dangerous areas, and not wearing safety ropes when working at height.

[0024] This embodiment uses a trained target behavior recognition model, which is obtained through a large amount of data training. Then, the first target (mine operator) is used as input to the target behavior recognition model, which analyzes the behavior of the mine operator, extracts the behavior features, and matches them with the behavior patterns learned in training. Finally, based on the matching results, the model determines whether the behavior of the mine operator is a dangerous behavior and outputs the corresponding judgment result.

[0025] The target behavior recognition model consists of two parts: the directional gradient histogram algorithm and the support vector machine; The first target is identified based on the target behavior identification model, and the dangerous behavior of the first target is determined, including: Determining local features of the first target based on a histogram of oriented gradients algorithm; The local features of the first target are classified based on a support vector machine to determine the dangerous behavior of the first target.

[0026] In this embodiment, the Histogram of Oriented Gradient (HOG) algorithm can construct features by calculating and counting the gradient direction histogram of the local area of ​​the image. For the first target, the HOG algorithm is used to extract the local contour and shape features of the first target.

[0027] Specifically, for the first target (operator) detected in the mining area image, the HOG algorithm is used to divide the area where it is located into multiple cells, calculate the gradient direction and amplitude of the pixels in each cell, and then count the gradient direction histogram of each cell. Adjacent cells are combined into blocks, and the histogram in the block is normalized to obtain the final HOG feature description, which describes the local shape and contour information of the operator, that is, the local features of the first target.

[0028] The parameters of the HOG algorithm include the pre-set cell size and the number of bins in the gradient direction, so as to calculate the HOG features of the area where the operator is located.

[0029] Assume that the resolution of the image is W×H (where W represents the width of the image and H represents the height of the image). To ensure that the cells have an appropriate coverage range in the image and considering the computational efficiency, the cell size can be determined according to a certain ratio.

[0030] Let the proportionality coefficient be r (0 < r < 1), then the calculation formulas for the cell width cell_width1 and the cell height cell_height1 are:

[0031]

[0032] Among them, represents the floor operation.

[0033] Assume that the approximate size range of the first target in the mining area image is known. The cell size can be determined according to the average width target_avg_width and the average height target_avg_height of the first target. Let a scaling factor s (s > 0) be used to adjust the ratio of the cell relative to the target size, then:

[0034]

[0035] Therefore, the calculation formula for the cell size is:

[0036]

[0037] Among them, cell_width is the width of the cell, cell_height is the height of the cell, k is the balance coefficient (0 < k < 1), which can flexibly control the influence degree of the image resolution and the target size on the cell size, cell_width1 is the first width of the cell, cell_height1 is the first height of the cell, cell_width2 is the second width of the cell, and cell_height2 is the second height of the cell.

[0038] The calculation formula for the number of bins N in the gradient direction is:

[0039]

[0040]

[0041]

[0042] ,

[0043] Among them, is the number of bins, is the basic number of bins, is the adjustment factor ( ), is the roughness of the first target edge, is the standard deviation of the distance, is the average value of the distance, is the Euclidean distance between adjacent points, is the point coordinates, , is the discretized point of the first target edge, is the total quantity. The higher the roughness, the higher the edge complexity.

[0044] Support Vector Machine (SVM) is used to classify the local features of the first target extracted by the HOG algorithm to determine whether the behavior of the operator is a dangerous behavior. SVM is a trained model. Specifically, the local features of the first target are used as the input of SVM, and SVM classifies the above features according to the classification boundary learned in the training stage. If SVM determines that the feature belongs to the feature category of dangerous behavior, it is determined that the behavior of the first target is a dangerous behavior; otherwise, it is a normal behavior.

[0045] Among them, in this embodiment, the penalty parameter of SVM can be adjusted to control the penalty degree for misclassified samples, that is, the larger the penalty parameter, the stricter the penalty for misclassification.

[0046] In this embodiment, an initial penalty parameter is determined; In response to the classification accuracy being greater than or equal to the first accuracy threshold, the initial penalty parameter is increased by the first step size to obtain the target penalty parameter. The first accuracy threshold is a preset reference value, which can be empirically set according to the requirements of the actual scenario. The first step size can be empirically set.

[0047] S103: Label the dangerous behavior of the first target based on the danger level and occurrence probability corresponding to the dangerous behavior.

[0048] In this embodiment, the risk levels corresponding to the dangerous behaviors can be empirically divided according to the historical number of casualties, economic losses, and time losses. The risk levels can be the first level, the second level, and the third level, that is, the first level is the low-risk level, the second level is the medium-risk level, and the third level is the high-risk level. For example, the behaviors of the low-risk level can only cause minor injuries or losses, such as not wearing gloves correctly; the behaviors of the medium-risk level can cause a certain degree of injury or equipment damage, such as operating general equipment in violation of regulations; the behaviors of the high-risk level can trigger serious accidents and even endanger lives, such as smoking in an area with excessive gas concentration, not wearing a safety belt during high-altitude operations, etc.

[0049] The occurrence probability corresponding to the dangerous behavior is the frequency of a certain dangerous behavior occurring within a period of time, which can be determined by statistical calculation of historical data. For example, according to the records of the past year, the behavior of miners not wearing safety helmets in a certain mining area occurs 6 times on average per month. The occurrence probability of this behavior is 6 / 30 = 0.2, that is, 20%.

[0050] Labeling the dangerous behaviors of the first target based on the risk levels and occurrence probabilities of the dangerous behaviors, including: Determining the first similarity based on the risk levels and occurrence probabilities of the dangerous behaviors. The first similarity is the risk similarity between the dangerous behaviors of the first target and multiple standard dangerous behaviors; Labeling the dangerous behaviors of the first target based on the first similarity.

[0051] The calculation formula for the risk similarity is:

[0052] Among them, is the risk similarity between the dangerous behavior of the first target and each standard dangerous behavior; is the number of factor dimensions considered. In this embodiment, two factor dimensions are considered, that is, m = 2; is the weight of the th factor dimension, and ; is the similarity score under the th factor dimension.

[0053] is the similarity score under the first factor dimension. The first factor dimension is the risk level. Therefore, the calculation formula for the similarity score under the risk level is:

[0054] Among them, is the quantified value of the risk level of the dangerous behavior of the first target, The quantified value of the risk level for each standard dangerous behavior, i.e., the first level corresponds to 1, the second level corresponds to 2, and the third level corresponds to 3; is the maximum value of the quantified risk level value, is the minimum value of the quantified risk level value.

[0055] is the similarity score under the second factor dimension. The second factor dimension is the occurrence probability. Therefore, the calculation formula for the similarity score under the occurrence probability is:

[0056] where, is the occurrence probability of the first target dangerous behavior, and the value range is [0, 1]; is the occurrence probability of each standard dangerous behavior, and the value range is [0, 1].

[0057] Label the dangerous behavior of the first target based on the danger similarity, including: In response to the danger similarity being greater than or equal to the first similarity threshold, determine the target labeling strategy corresponding to the dangerous behavior of the first target; Label the dangerous behavior of the first target based on the target labeling strategy.

[0058] In this embodiment, the first similarity threshold is a pre-set reference standard value, which can be set according to experience. When the danger similarity exceeds the first similarity threshold, it indicates that the dangerous behavior of the first target is highly similar to the corresponding standard dangerous behavior, and the standard strategy corresponding to this standard dangerous behavior can be used as the target labeling strategy.

[0059] The target labeling strategy can be a color standard strategy. If focusing on risk level reminder, the dangerous behavior of the first target can be labeled with a prominent color to facilitate relevant personnel's attention, thereby giving a risk reminder; detailed risk reminders can also be added, such as "This behavior is highly similar to the standard behavior that triggers a major safety accident. Please pay attention immediately!" It can be concluded from the above that the present disclosure processes the mining area image using the object detection model, which can efficiently identify the first target, improving the accuracy and speed of object detection. Then, the present disclosure identifies the dangerous behavior of the first target through the target behavior recognition model, which can timely and accurately discover the dangerous behavior of the first target and effectively avoid the occurrence of accidents. At the same time, the present disclosure labels the dangerous behavior, which can specifically identify the dangerous behavior. Therefore, the present disclosure can improve the accuracy of visual recognition of dangerous behaviors in the mining area.

[0060] In an embodiment of the present disclosure, the method for visual recognition of dangerous behaviors in the mining area based on images further includes: Determine the parameters of the target region proposal network; Determine the parameters of the target fast region convolutional neural network; Based on the target region proposal network and the target fast region convolutional neural network, obtain the target fast region recurrent convolutional neural network model; Among them, the target region proposal network is a trained region proposal network, and the target fast region convolutional neural network is a trained fast region convolutional neural network.

[0061] In this embodiment, train the region proposal network to obtain the target region proposal network, and fix the parameters of the target region proposal network; Train the fast region convolutional neural network to obtain the target fast region convolutional neural network, and fix the parameters of the target fast region convolutional neural network; Adjust the parameters of the target region proposal network and the parameters of the target fast region convolutional neural network until, after the first number of training rounds, when the detection accuracy of the target fast region recurrent convolutional neural network model meets the first condition, stop the joint training to obtain the target fast region recurrent convolutional neural network model; The first condition is that the detection accuracy of the target fast region recurrent convolutional neural network model is greater than or equal to the first accuracy rate, and the second number is continuously satisfied with the detection accuracy rate greater than or equal to the first accuracy rate.

[0062] The first image set is a collection of a large number of historical mining area images, including the dangerous behavior positions and behavior types of operators. The joint training is that the target RPN and the target fast region convolutional neural network (Fast Region-based Convolutional Neural Networks‌‌, Fast R-CNN) adopt an alternating training method, that is, first train the target RPN for the first time interval, then fix the parameters of the target RPN, and then train the target Fast R-CNN; then fine-tune the target RPN, and then fine-tune the Fast R-CNN. After the first number of rounds, when the detection accuracy of the target Faster R-CNN model meets the first condition, stop the joint training to obtain the target Faster R-CNN model.

[0063] The first condition is that the detection accuracy of the target Faster R-CNN model is greater than or equal to the first accuracy rate and the second number is continuously satisfied with the detection accuracy rate greater than or equal to the first accuracy rate.

[0064] The first time interval is an initial time set, which can be set according to experience; the first quantity depends on whether the detection accuracy of the target Faster R-CNN model meets the first condition, and an initial first quantity can be determined. In response to the detection accuracy of the target Faster R-CNN model meeting the first condition when it is less than the initial first quantity, the initial first quantity can be reduced to obtain the target first quantity, thereby improving the efficiency of model training; in response to the detection accuracy of the target Faster R-CNN model still not meeting the first condition when it is equal to the initial first quantity, the initial first quantity can be increased to obtain the target first quantity, thereby meeting the requirements of detection accuracy.

[0065] After the joint training is completed, the target RPN and the target Fast R-CNN are connected to form a complete target Faster R-CNN model. At the same time, the target Faster R-CNN model inherits the parameters of the target RPN and the target Fast R-CNN.

[0066] It can be concluded from the above that this embodiment can not only significantly improve the accuracy and efficiency of target detection, but also effectively reduce the false alarms and missed alarms of subsequent dangerous behaviors, thereby improving the overall safety management level of the mining area.

[0067] In an embodiment of the present disclosure, the training process of the target region proposal network includes: Input the first image set into the backbone network to obtain a feature map; Input the feature map into the region proposal network to obtain initial candidate regions, where the initial candidate regions include anchor box categories and anchor box positions; Determine a first loss function based on the anchor box categories, determine a second loss function based on the anchor box positions, and perform weighted calculation on the first loss function and the second loss function to obtain a first comprehensive loss function; Update the parameters of the region proposal network based on the first loss value of the first comprehensive loss function until the first loss value is less than the first threshold or the number of iterations reaches the second threshold to obtain the target region proposal network.

[0068] In this embodiment, the backbone network is a pre-trained deep convolutional neural network, such as ResNet, VGG, etc. The backbone network is used to extract features from the input first image set to obtain a feature map. The convolutional layer of the backbone network will perform feature extraction operations on the first image set, that is, through the convolutional operation of the convolutional kernel, the pixel information in the image is converted into features at different levels, and finally a feature map is obtained. The above feature map contains the feature information of various targets in the image, providing a necessary data basis for the subsequent RPN to generate candidate regions.

[0069] The Region Proposal Network (RPN) receives the feature maps output by the backbone network. In the RPN, anchor boxes of different sizes and ratios are preset. The above-mentioned anchor boxes are windows on the feature maps, used to search for targets at different positions and scales. The RPN processes the feature maps through convolutional operations. At each position of the feature map, based on the anchor boxes, it predicts whether there is a first target and the position of the first target. For each anchor box, two key pieces of information are output: one is the anchor box category, that is, to judge whether the first target is contained in the anchor box, which is a binary classification problem, and the result is either containing the first target or not containing the first target; the other is the anchor box position, which is represented by the offset relative to the original anchor box position. Through these offsets, the position and size of the anchor box can be adjusted to more accurately frame the first target. For example, if an anchor box has a high probability of containing a target after being processed by the RPN and corresponding position offsets are obtained, then this anchor box becomes an initial candidate region and may contain the first target in the image.

[0070] The calculation formula of the first loss function is:

[0071] It can be concluded from the above that is the first loss function, is the number of anchor boxes, is the probability of the anchor box category prediction, is the th anchor box, is the true category label ( ), 0 indicates non-first target, and 1 indicates first target.

[0072] The calculation formula of the second loss function is:

[0073]

[0074]

[0075]

[0076]

[0077]

[0078] Among them, are the predicted position coordinates of the anchor box, are the true position coordinates of the anchor box, , , , are all the offsets between the predicted position coordinates and the true position coordinates of the anchor box.

[0079] The calculation formula of the first comprehensive loss function is as follows:

[0080] Among them, is the weight coefficient, , which is used to balance the importance of the class loss and the location loss.

[0081] In this embodiment, the loss value calculated using the first comprehensive loss function is used to update the parameters of the region proposal network through the backpropagation algorithm. Backpropagation will reverse the loss value along the connections of the network, calculate the gradient of each parameter with respect to the loss value, and then adjust the values of the parameters according to the gradient, so that the prediction result of the network is closer to the true label.

[0082] This process of updating the parameters is continuously repeated, that is, multiple iterative trainings are performed. In each iteration, a new loss value is calculated, and the parameters are continuously updated according to the loss value. When the first loss value is less than a pre-set first threshold (indicating that the prediction of the model is accurate enough), or the number of iterations reaches a second threshold (the pre-set maximum number of training times to prevent overtraining or non-convergence of training), the training is stopped. At this time, the obtained region proposal network is the target region proposal network, which can relatively accurately generate candidate regions containing the target.

[0083] It can be concluded from the above that the RPN generates initial candidate regions, effectively narrowing the search range. In this embodiment, through the first comprehensive loss function, both the anchor box category and the location information are considered, ensuring the comprehensiveness and pertinence of the training. The iterative optimization process continuously fine-tunes the network parameters until the established conditions are met, thereby obtaining a target RPN with excellent performance. This embodiment not only improves the detection accuracy but also speeds up the detection speed.

[0084] In an embodiment of the present disclosure, the training process of the target fast region convolutional neural network includes: Determine a feature vector of a fixed size; Input the feature vector into the fully connected layer of the fast region convolutional neural network to obtain the category and location of each candidate region; Based on the category of the candidate region, determine the third loss function, and based on the location of the candidate region, determine the fourth loss function. Perform weighted calculation on the third loss function and the fourth loss function to obtain the second comprehensive loss function; Update the parameters of the fast region convolutional neural network based on the second loss value of the second comprehensive loss function until the second loss value is less than the third threshold or the number of iterations reaches the fourth threshold, to obtain the target fast region convolutional neural network.

[0085] In this embodiment, determining the feature vector of a fixed size includes: Determine target candidate regions based on the target region proposal network. The target candidate regions include multiple candidate regions, and the size of each candidate region is different; Map the target candidate regions to the feature map to obtain multiple candidate region feature maps; Based on the region of interest pooling layer, transform each candidate region feature map to generate a feature vector with a fixed size.

[0086] During the previous training process, the target RPN has been able to generate initial candidate regions that may contain the target based on the input feature map. In this embodiment, the target candidate regions are screened from the above-mentioned initial candidate regions. The initial candidate regions generated by the target RPN are based on anchor boxes, and the anchor boxes have different sizes and ratios, so the sizes of the obtained candidate regions are also different. The above-mentioned candidate regions represent the regions in the image where dangerous behavior targets may exist. For example, screening is performed based on the probability threshold that the candidate region contains the target, and only the initial candidate regions with a probability greater than a certain value will be selected as the target candidate regions.

[0087] The feature map output by the backbone network contains rich feature information of the image. Mapping the target candidate regions to this feature map is to obtain the features corresponding to each candidate region. Since the feature map is the result of feature extraction from the original image, the mapped region of the candidate region on the feature map contains the feature representation of the target within the candidate region. Through this mapping operation, multiple candidate region feature maps are obtained. Each candidate region feature map corresponds to a target candidate region, and these feature maps retain the feature information of the target candidate region in the original image.

[0088] The RoI pooling layer is used to transform candidate region feature maps of different sizes into feature vectors with a fixed size. Because the subsequent fully connected layer requires the input features to have a fixed dimension, and the sizes of the candidate region feature maps are different, the RoI pooling layer is needed for transformation. The RoI pooling layer divides the candidate region feature map into a fixed number of sub-regions, and then performs a pooling operation (such as max pooling) on each sub-region, thereby obtaining a feature vector with a fixed size. In this way, each candidate region is transformed into a feature representation with a fixed dimension, which is convenient for input into the fully connected layer for processing.

[0089] The fully connected layer of Fast R-CNN receives the fixed-size feature vector output by the RoI pooling layer. The fully connected layer maps the feature vector to the category and location information of the target through a series of weight matrix multiplications and activation function operations. For category prediction, the fully connected layer outputs the probability that each candidate region belongs to a different category. For example, in the detection of dangerous behaviors in mining areas, there may be categories such as not wearing a safety helmet and operating equipment in violation of regulations. These probabilities can be used to determine which category the target in the candidate region belongs to. For location prediction, the fully connected layer outputs the offset of the bounding box coordinates of the candidate region. These offsets can be used to adjust the position of the candidate region so that it can frame the target more accurately.

[0090] The calculation formula of the third loss function is:

[0091] in, is the third loss function, is the number of candidate regions, To predict the probability that the candidate region belongs to each category, is the true candidate region category label, For the candidate regions.

[0092] The calculation formula of the fourth loss function is:

[0093]

[0094]

[0095]

[0096]

[0097]

[0098] The calculation formula for the degree of deviation is:

[0099] The calculation formula of the adaptive adjustment factor is:

[0100] in, is the predicted position coordinate of the candidate area, is the actual position coordinate of the candidate area, is the predicted location area of ​​the candidate region, , , , They are all the offsets between the predicted position coordinates and the true position coordinates of the candidate regions, is the numerical value of the deviation degree, is the adaptive adjustment factor, is a hyperparameter used to control the variation range of the adjustment factor.

[0101] The fourth loss function comprehensively considers the position deviation, area, and the adaptive penalty of the deviation degree, and can flexibly handle the position prediction errors in different situations.

[0102] The calculation formula of the second comprehensive loss function is:

[0103] where, is the weight coefficient, , which is used to balance the importance of the class loss and the position loss.

[0104] After calculating the loss value of the second comprehensive loss function in this embodiment, the parameters of FastR-CNN are updated through the backpropagation algorithm. The backpropagation algorithm will propagate the loss value from the output layer back to the input layer along the connection path of the network, and calculate the gradient of each parameter with respect to the loss value. According to the magnitude and direction of the gradient, an optimizer is used to adjust the parameters of Fast R-CNN to gradually reduce the loss value. During the training process, this parameter update step is continuously repeated, that is, multiple iterations are performed. Each iteration will calculate a new loss value and continue to update the parameters according to the loss value. When the second loss value is less than a pre-set third threshold, it means that the prediction of Fast R-CNN is accurate enough and the training can be stopped; or when the number of iterations reaches a pre-set fourth threshold, even if the loss value is not less than the third threshold, the training is also stopped. The FastR-CNN obtained at this time is the target Fast R-CNN, which can accurately classify and perform position regression on the input candidate regions, and realize the detection of the dangerous behavior target (i.e., the first target) in the mining area.

[0105] It can be concluded from the above that the training process of the target Fast R-CNN effectively reduces the number of candidate regions through the precise target RPN, improving the detection efficiency. Mapping the candidate regions to the feature map and performing the conversion of a fixed size ensures the consistency of the input data, which is beneficial to improving the stability and accuracy of the model. At the same time, by comprehensively considering the loss functions of the class and the position for parameter update, the model can optimize the object detection and localization performance simultaneously.

[0106] Corresponding to the above-mentioned embodiment of the method for visual recognition of dangerous behaviors in mining areas based on images, Figure 2The block diagram of the image-based visual recognition system for dangerous behaviors in mining areas provided by an embodiment of the present disclosure. For ease of illustration, only parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 , the image-based visual recognition system 20 for dangerous behaviors in mining areas includes: a target positioning module 21, a behavior recognition module 22, and a behavior annotation module 23.

[0107] Among them, the target positioning module 21 is configured to process a mining area image based on a target fast region recursive convolutional neural network model to obtain a first target, and the first target is an operator; The behavior recognition module 22 is configured to recognize the first target based on a target behavior recognition model to determine the dangerous behavior of the first target; The behavior annotation module 23 is configured to annotate the dangerous behavior of the first target based on the danger level and occurrence probability corresponding to the dangerous behavior.

[0108] In an embodiment of the present disclosure, the image-based visual recognition system 20 for dangerous behaviors in mining areas further includes: a model training module, configured to determine the parameters of a region proposal network for targets; Determine the parameters of a fast region convolutional neural network for targets; Based on the region proposal network for targets and the fast region convolutional neural network for targets, obtain a target fast region recursive convolutional neural network model; Among them, the region proposal network for targets is a trained region proposal network, and the fast region convolutional neural network for targets is a trained fast region convolutional neural network.

[0109] In an embodiment of the present disclosure, the model training module is specifically configured to input a first image set into a backbone network to obtain a feature map; Input the feature map into the region proposal network to obtain initial candidate regions, and the initial candidate regions include anchor box categories and anchor box positions; Determine a first loss function based on the anchor box categories, determine a second loss function based on the anchor box positions, and perform weighted calculation on the first loss function and the second loss function to obtain a first comprehensive loss function; Update the parameters of the region proposal network based on the first loss value of the first comprehensive loss function until the first loss value is less than a first threshold or the number of iterations reaches a second threshold to obtain the region proposal network for targets.

[0110] In an embodiment of the present disclosure, the model training module is specifically further configured to determine a feature vector of a fixed size; Input the feature vector into the fully connected layer of the fast region convolutional neural network to obtain the category and position of each candidate region; Determine a third loss function based on the category of the candidate region, determine a fourth loss function based on the position of the candidate region, and perform weighted calculation on the third loss function and the fourth loss function to obtain a second comprehensive loss function; Update the parameters of the fast region convolutional neural network based on the second loss value of the second comprehensive loss function until the second loss value is less than a third threshold or the number of iterations reaches a fourth threshold, to obtain a target fast region convolutional neural network.

[0111] In an embodiment of the present disclosure, the target behavior recognition model includes two parts: the histogram of oriented gradients algorithm and the support vector machine; The behavior recognition module 22 is specifically configured to determine the local features of the first target based on the histogram of oriented gradients algorithm; Classify the local features of the first target based on the support vector machine to determine the dangerous behavior of the first target.

[0112] In an embodiment of the present disclosure, the behavior annotation module 23 is specifically configured to determine a first similarity based on the danger level and occurrence probability of the dangerous behavior, and the first similarity is the danger similarity between the dangerous behavior of the first target and multiple standard dangerous behaviors; Annotate the dangerous behavior of the first target based on the first similarity.

[0113] In an embodiment of the present disclosure, the behavior annotation module 23 is further specifically configured to determine a target annotation strategy corresponding to the dangerous behavior of the first target in response to the danger similarity being greater than or equal to a first similarity threshold; Annotate the dangerous behavior of the first target based on the target annotation strategy.

[0114] See Figure 3 , Figure 3 is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 3 shown, the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned system embodiments, such as Figure 2 the functions of the modules 21 to 23 shown.

[0115] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0116] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0117] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0118] In specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present disclosure may implement the implementation manners described in the first embodiment and the second embodiment of the method for visual recognition of dangerous behaviors in a mining area based on images provided by the embodiments of the present disclosure, and may also implement the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated herein.

[0119] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It can also be completed by instructing related hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0120] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the electronic device. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0121] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present disclosure.

[0122] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0123] In several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces or units, or can also be in the form of electrical, mechanical or other connections.

[0124] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can also be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.

[0125] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0126] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A method for visually identifying dangerous behaviors in mining areas based on images, characterized in that: include: Processing the mining area image based on the target fast regional recursive convolutional neural network model to obtain a first target, wherein the first target is an operator; Identify the first target based on the target behavior recognition model to determine the dangerous behavior of the first target; The dangerous behavior of the first target is marked based on the danger level and occurrence probability corresponding to the dangerous behavior.

2. The method for visually identifying dangerous behaviors in mining areas based on images according to claim 1, characterized in that: Also includes: Determine the parameters of the target region proposal network; Determine the parameters of the target fast regional convolutional neural network; Based on the target region proposal network and the target fast region convolutional neural network, obtaining the target fast region recursive convolutional neural network model; The target region proposal network is a trained region proposal network, and the target fast region convolutional neural network is a trained fast region convolutional neural network.

3. The method for visually identifying dangerous behaviors in mining areas based on images according to claim 2, characterized in that: The training process of the target region proposal network includes: Input the first image set into the backbone network to obtain a feature map; Inputting the feature map into a region proposal network to obtain an initial candidate region, wherein the initial candidate region includes an anchor box category and an anchor box position; Determine a first loss function based on the anchor box category, determine a second loss function based on the anchor box position, and perform weighted calculation on the first loss function and the second loss function to obtain a first comprehensive loss function; Based on the first loss value of the first comprehensive loss function, the parameters of the region proposal network are updated until the first loss value is less than a first threshold or the number of iterations reaches a second threshold, thereby obtaining a target region proposal network.

4. The method for visually identifying dangerous behaviors in mining areas based on images according to claim 3, characterized in that: The training process of the target fast regional convolutional neural network includes: Determine the fixed-size feature vector; Inputting the feature vector into the fully connected layer of the fast regional convolutional neural network to obtain the category and position of each candidate region; Determine a third loss function based on the category of the candidate area, determine a fourth loss function based on the position of the candidate area, and perform weighted calculation on the third loss function and the fourth loss function to obtain a second comprehensive loss function; Based on the second loss value of the second comprehensive loss function, the parameters of the fast regional convolutional neural network are updated until the second loss value is less than a third threshold or the number of iterations reaches a fourth threshold, thereby obtaining a target fast regional convolutional neural network.

5. The method for visually identifying dangerous behaviors in mining areas based on images according to claim 1, characterized in that: The target behavior recognition model includes two parts: a directional gradient histogram algorithm and a support vector machine; The identifying the first target based on the target behavior identification model and determining the dangerous behavior of the first target includes: Determining local features of the first target based on the histogram of oriented gradients algorithm; The local features of the first target are classified based on the support vector machine to determine the dangerous behavior of the first target.

6. The method for visually identifying dangerous behaviors in mining areas based on images according to claim 1, characterized in that: The marking of the dangerous behavior of the first target based on the danger level and occurrence probability of the dangerous behavior includes: Determining a first similarity based on the danger level and occurrence probability of the dangerous behavior, the first similarity being the danger similarity between the dangerous behavior of the first target and a plurality of standard dangerous behaviors; The dangerous behavior of the first target is marked based on the first similarity.

7. The method for visually identifying dangerous behaviors in mining areas based on images according to claim 6, characterized in that: The marking of the dangerous behavior of the first target based on the dangerous similarity includes: In response to the danger similarity being greater than or equal to a first similarity threshold, determining a target labeling strategy corresponding to the dangerous behavior of the first target; The dangerous behavior of the first target is labeled based on the target labeling strategy.

8. An image-based visual identification system for dangerous behaviors in mining areas, characterized in that: include: A target positioning module is used to process the mining area image based on a target fast regional recursive convolutional neural network model to obtain a first target, where the first target is an operator; A behavior recognition module, used to recognize the first target based on a target behavior recognition model and determine the dangerous behavior of the first target; The behavior labeling module is used to label the dangerous behavior of the first target based on the danger level and occurrence probability corresponding to the dangerous behavior.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Abnormal behavior monitoring method, apparatus, computer device, and storage medium

    CN109241946A

  • Target detection method and system based on deep learning

    CN111626349A

  • Intervertebral disc CT image detection method based on deep convolutional neural network

    CN112308822A

  • Mine personnel dangerous behavior identification method and system based on computer vision

    CN117593689A

  • Electric power operation risk behavior violation intelligent identification method based on machine vision

    CN119229526A