Rolling stone identification and tracking method based on improved yolk-deepsort
By improving the YOLO-Deepsort network model and combining it with MobileNetV4 and WTConv modules, a rolling stone target recognition model was constructed. Combined with the Deepsort algorithm, the problems of low rolling stone recognition efficiency and insufficient extraction of motion trajectory features were solved, achieving efficient and accurate rolling stone recognition and tracking.
Patent Information
- Application Number
- CN202511221518.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-12-05
AI Technical Summary
Existing technologies have low efficiency in identifying falling rocks in rockfall disaster monitoring, insufficient ability to extract features of movement trajectory and speed changes, and insufficient research on target tracking technology, leading to monitoring difficulties.
An improved YOLO-Deepsort network model was constructed by obtaining a rolling stone image dataset for improvement. The YOLOv11 network model was improved by combining the MobileNetV4 module and the WTConv module to construct a rolling stone target recognition model. The rolling stone recognition and tracking network was trained and validated by combining the deepsort tracking algorithm.
It improves the efficiency of rolling stone recognition, enhances the detection accuracy of small, distant or partially obscured rolling stones, and improves the ability to extract motion trajectory and speed change features, thus achieving efficient and accurate rolling stone recognition and tracking.
Smart Images

Figure CN121074792A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of disaster monitoring, and particularly relates to a rockfall recognition and tracking method based on an improved yolo-deepsort. BACKGROUND
[0002] The special geological environment, fragile ecological system and abundant rainfall in the mountainous areas of China cause rockfall geological disasters to ravage, causing significant losses to infrastructure, major projects and the lives and property of residents in the mountainous areas of China. In addition, rockfall disasters have the characteristics of suddenness and randomness, which makes active monitoring and early warning particularly difficult.
[0003] The current research status of rockfall monitoring can be divided into two categories: early dangerous rock mass monitoring and visual recognition technology, wherein the visual recognition technology includes traditional image processing technology and deep learning methods. The early dangerous rock mass monitoring method has problems such as high cost, difficult data processing, and high false alarm rate. The traditional image processing method usually relies on manually designed and extracted features for analysis and processing, which not only consumes time but also easily leads to information loss or limitation. Deep learning algorithms can automatically learn feature representations from raw data and have broad application prospects in the field of rockfall disaster monitoring. However, current research on intelligent recognition of rockfall disasters using deep learning methods mostly stays at the target detection stage, and the exploration of target tracking technology is relatively insufficient, resulting in a lack of extraction ability for rockfall motion trajectory, speed change and other features. In addition, since the target tracking based on detection has more advantages than single tracking algorithm, and the performance of the target tracking method based on detection directly depends on the accuracy and effect of target detection, it is urgent to improve the rockfall target detection effect and then combine the target tracking algorithm to complete rockfall recognition and tracking. SUMMARY
[0004] In view of the above deficiencies in the prior art, the application provides a rockfall recognition and tracking method based on an improved yolo-deepsort, which constructs a rockfall picture dataset and improves the neural network structure of yolov11, and then combines the deepsort tracking algorithm on the improved rockfall target recognition model to train and verify a rockfall recognition and tracking network model, solving the problems of low rockfall recognition efficiency in mountainous areas and weak extraction ability for motion trajectory, speed change and other features.
[0005] In order to achieve the above-mentioned application purposes, the technical scheme adopted by the application is as follows:
[0006] The application provides a rockfall recognition and tracking method based on an improved yolo-deepsort, which includes the following steps:
[0007] S1, obtain a rockfall picture dataset and perform preprocessing to obtain a rockfall picture training set and a validation set;
[0008] S2, improving the yolov11 network model based on the MobileNetV4 module and the WTConv module, constructing a rockfall target recognition model, and then obtaining a rockfall target recognition network;
[0009] S3, constructing a deepsort rockfall tracking model based on the obtained rockfall target recognition network, and then obtaining an improved recognition and tracking network;
[0010] S4, training and verifying the improved recognition and tracking network by using the rockfall picture training set and the rockfall image verification set, and obtaining a rockfall recognition and tracking model;
[0011] S5, obtaining a rockfall picture or video to be detected, and predicting the rockfall picture or video to be detected by using the obtained recognition and tracking model, and obtaining the recognition and tracking result of the rockfall picture or video to be detected.
[0012] The rockfall recognition and tracking method based on the improved yolo-deepsort is provided, the rockfall picture dataset is obtained and preprocessed, the yolov11 network model is improved based on the MobileNetV4 module and the WTConv module, and the rockfall target recognition model is constructed. The deepsort rockfall tracking model is constructed based on the rockfall target recognition network. The calculation complexity can be effectively reduced and the calculation efficiency can be improved through the MobileNetV4 module network, the small size, remote or partially occluded rockfall can be more accurately detected based on the WTConv module, so that the detection precision is improved and the rockfall recognition effect is improved. The improved yolo-deepsort network model is constructed, the rockfall picture training set and the rockfall picture verification set divided by the rockfall picture dataset are used to train and verify the completed network model, and finally the network model capable of efficiently and accurately recognizing and tracking the rockfall situation is obtained.
[0013] Further, the S1 comprises the following steps:
[0014] S11, obtaining a rockfall picture dataset;
[0015] S12, selecting the pictures in the rockfall picture dataset one by one, and performing data enhancement preprocessing to expand the rockfall picture dataset;
[0016] S13, randomly select a part of the rolling stone pictures in the expanded rolling stone picture dataset to form a rolling stone picture training set, and form the remaining rolling stone pictures in the rolling stone picture dataset into a rolling stone picture verification set.
[0017] The beneficial effect of the further scheme is that the model robustness is effectively enhanced by data augmentation of the rolling stone pictures in the rolling stone picture dataset.
[0018] Further, the improved yolo-deepsort rolling stone recognition and tracking model comprises a MobileNetV4 backbone network, a WTConv Neck network connected with the MobileNetV4 backbone network, a Head network connected with the WTConv Neck network, and a deepsort network connected with the Head network.
[0019] The MobileNetV4 backbone network comprises a rolling stone picture input module, a first CBS module, a second CBS module, a first UIB module, a second UIB module, and a third CBS module connected in sequence; the output end of the rolling stone picture input module is used as the image input end of the MobileNetV4 backbone network; the first rolling stone feature output end of the first UIB module, the second rolling stone feature output end of the second UIB module, and the third rolling stone feature output end of the third CBS module are all connected with the WTConv Neck network.
[0020] The beneficial effect of the further scheme is that the calculation path in the feature extraction process is optimized by the UIB module, the calculation complexity and the parameter amount of the model are reduced, and the MobileNetV4 structure is used in the backbone network to maintain good feature expression ability.
[0021] The WTConv Neck network comprises an SPPF module, a C2PSA module, a first WTC3 module, a second WTC3 module, a third WTC3 module, a fourth WTC3 module, a fourth CBS module, a fifth CBS module, a first up-sampling module, a second up-sampling module, a first Concat splicing module, a second Concat splicing module, a third Concat splicing module and a fourth Concat splicing module; an output end of the SPPF module is connected with a first input end of the C2PSA module; a first output end of the C2PSA module is connected with a first input end of the fourth Concat splicing module; a second output end of the C2PSA module is connected with an input end of the first up-sampling module; an output end of the first up-sampling module is connected with a first input end of the first Concat splicing module; an output end of the first Concat splicing module is connected with an input end of the first WTC3 module; a first output end of the first WTC3 module is connected with a first input end of the third Concat splicing module; a second output end of the first WTC3 module is connected with an input end of the second up-sampling module; an output end of the second up-sampling module is connected with a first input end of the second Concat splicing module; an output end of the second Concat splicing module is connected with an input end of the second WTC3 module; a first output end of the second WTC3 module is connected with an input end of the fourth CBS module; an output end of the fourth CBS module is connected with a second input end of the third Concat splicing module; an output end of the third Concat splicing module is connected with an input end of the third WTC3 module; a first output end of the third WTC3 module is connected with an input end of the fifth CBS module; an output end of the fifth CBS module is connected with a second input end of the fourth Concat splicing module; an output end of the fourth Concat splicing module is connected with an input end of the fourth WTC3 module.
[0022] The beneficial effects of the further scheme are that the WTConv Neck network adopts dynamic weight distribution and multi-scale feature fusion, optimizes the limitations of traditional convolution on the receptive field, enables the network to more accurately capture the subtle texture and edge features of the rolling stone target, and improves the detection accuracy of small target rolling stones,
[0023] The Head network comprises a first prediction head connected with a second output end of the second WTC3 module, a second prediction head connected with a second output end of the third WTC3 module, and a third prediction head connected with an output end of the fourth WTC3 module.
[0024] The deepsort network comprises a feature extraction module connected with the first prediction head, the second prediction head and the third prediction head, an association matching module, and a periodical updating module.
[0025] The beneficial effects of the above further scheme are: the improved target recognition model is combined with the deepsort network to extract the motion characteristics of the rolling stone, and through motion prediction and appearance similarity, the rolling stone trajectory description and parameter extraction can be realized and enhanced.
[0026] Further, the S4 comprises the following steps:
[0027] S41, randomly selecting a rolling stone picture in a rolling stone picture training set, and inputting the selected rolling stone picture into an improved yolo-deepsort network model;
[0028] S42, using a MobileNetV4 backbone network to extract features of the selected rolling stone picture, and sequentially outputting a first rolling stone feature output end, a second rolling stone feature output end and a third rolling stone feature output end to obtain a first MobileNetV4 backbone rolling stone feature map, a second MobileNetV4 backbone rolling stone feature map and a third MobileNetV4 backbone rolling stone feature map;
[0029] S43, using a WTConv Neck network to perform upsampling, feature splicing and feature extraction on the first MobileNetV4 backbone rolling stone feature map, the second MobileNetV4 backbone rolling stone feature map and the third MobileNetV4 backbone rolling stone feature map, to obtain a first rolling stone prediction feature map, a second rolling stone prediction feature map and a third rolling stone prediction feature map;
[0030] S44, using a first prediction head in a Head network to perform rolling stone prediction on the first rolling stone prediction feature map, using a second prediction head in the Head network to perform rolling stone prediction on the second rolling stone prediction feature map, and using a third prediction head in the Head network to perform rolling stone prediction on the third rolling stone prediction feature map, to obtain a rolling stone detection map with target frames and confidence;
[0031] S45, using a deepsort network to perform feature extraction, correlation matching and cycle updating on the rolling stone detection map with target frames and confidence, to obtain a rolling stone tracking map with time sequence information and motion trajectory;
[0032] S46, repeating the S41-S45 stage training a plurality of times, using a rolling stone picture verification set to verify the yolo-deepsort network model after the stage training, and saving the hyperparameters at the time of the verification;
[0033] S47, repeating the S46 a predetermined number of times, and based on the saved hyperparameters, obtaining the hyperparameters corresponding to the optimal network depth, and obtaining a rolling stone recognition and tracking network model.
[0034] The beneficial effect of adopting the further scheme is that the improved yolo-deepsort network model is trained and verified based on the rolling stone picture training set and the rolling stone picture verification set, and after the hyperparameters corresponding to the optimal network depth are determined, the robustness of the rolling stone recognition and tracking network model is enhanced.
[0035] Further, the first UIB module and the second UIB module each include a first DepthConv module, a second DepthConv module, an ExpandConv module and a ProjectConv module; the input end of the first DepthConv module serves as a first rolling stone picture input end; the output end of the first DepthConv module is connected with the input end of the ExpandConv module; the output end of the ExpandConv module is connected with the input end of the second DepthConv module; the output end of the second DepthConv module is connected with the input end of the ProjectConv module; and the first DepthConv module, the ExpandConv module, the second DepthConv module and the ProjectConv module sequentially perform convolution on the input pictures to obtain a UIB rolling stone feature map.
[0036] The beneficial effect of adopting the further scheme is that the UIB module is adopted in the backbone network in the application to effectively reduce the calculation amount of the model and effectively improve the balance ability of the model in terms of the efficiency and precision of rolling stone recognition.
[0037] Further, the first WTC3 module, the second WTC3 module, the third WTC3 module and the fourth WTC3 module each include a first WTConv module, a Split module, n Bottleneck modules connected in sequence, a fifth Concat splicing module and a second WTConv module.
[0038] The output end of the first WTConv module is connected with the input end of the Split module; the first output end of the Split module is connected with the input end of the first Bottleneck module; the second output end of the Split module is connected with the first input end of the fifth Concat splicing module; the output end of the nth Bottleneck module is connected with the second input end of the fifth Concat splicing module; the output end of the fifth Concat splicing module is connected with the input end of the second WTConv module; and the output end of the second WTConv module obtains a WTC3 rolling stone feature map, wherein n is a positive integer.
[0039] The beneficial effect of adopting the further scheme is that the model receptive field is expanded based on the WTConv module, and the monitoring precision for small target rolling stones is improved.
[0040] Further, the feature extraction module adopts Kalman filtering to extract features and output the predicted position of the rolling stone; the correlation matching module adopts IOU cascade matching to complete the optimal matching of the predicted position of the rolling stone; and the periodic updating module manages the life cycle of the rolling stone track and updates the predicted state of the rolling stone through data correlation.
[0041] The beneficial effects of the above further scheme are: the prediction-matching-updating closed-loop architecture based on Kalman filtering is adopted to realize high-precision tracking and state estimation of the rolling stone track.
[0042] Further, the Kalman filtering calculation expression is as follows:
[0043]
[0044] wherein, is the current frame state estimation of the rolling stone, is the last frame state estimation of the rolling stone, F k is the current frame rolling stone state transition matrix, K k is the rolling stone Kalman gain, z k is the rolling stone bounding box observation value, H k is the rolling stone observation matrix;
[0045] The IOU cascade matching calculation expression is as follows:
[0046] D(A, B) = S(A∩B) / S(A∪B)
[0047] C = 1 - D(A, B)
[0048] M = min_π∑_{(i,j)∈π}C_{ij}
[0049] wherein, A is the bounding box coordinate of the rolling stone predicted track, B is the bounding box coordinate of the rolling stone detection result, S(A∩B) is the area of the overlapping region of the two boxes, S(A∪B) is the total coverage area of the two boxes, D(A, B) is the ratio of the two areas, C is the rolling stone matching cost in the interval [0, 1], π is all possible rolling stone matching combinations, C_{ij} is the cost of the i-th rolling stone predicted box and the j-th rolling stone detected box, and M is the minimum total cost under the optimal matching of the rolling stone;
[0050] The data correlation calculation expression is as follows:
[0051] S = α·D + β·M + γ·A
[0052]
[0053] Wherein, S is a rolling stone comprehensive matching score, alpha, beta, gamma is the weight coefficient in the interval [0, 1], D is the rolling stone spatial coincidence degree, M is the rolling stone similarity, A is the rolling stone cosine similarity, H is the rolling stone confirmation threshold, L is the rolling stone deletion threshold, F is the rolling stone front number of times of not matching, N is the maximum number of tolerated lost frames of rolling stones, T is the trajectory state, 1 is the confirmation state, 2 is the tentative state, and 3 is the deletion state.
[0054] The beneficial effects of the above further scheme are: based on the improved target recognition network, the motion state of the rolling stone is dynamically predicted through Kalman filtering, high-precision target association is realized by combining IOU cascade matching, and the trajectory management strategy of multi-feature fusion is used to adaptively update the rolling stone trajectory life cycle, which significantly improves the robustness of the rolling stone tracking and provides real-time and accurate motion data support for geological disaster monitoring.
[0055] Other advantages of the present application will be analyzed in more detail in the subsequent examples. BRIEF DESCRIPTION OF DRAWINGS
[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced below, and it should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0057] Figure 1 The step flow chart of the improved yolo-deepsort-based rolling stone recognition and tracking method in the embodiments of the present application.
[0058] Figure 2 The structure diagram of the improved yolo-deepsort network model in the embodiments of the present application.
[0059] Figure 3 The schematic diagram of the UIB structure in the embodiments of the present application.
[0060] Figure 4 The schematic diagram of the WTC3 structure in the embodiments of the present application.
[0061] Figure 5 The schematic diagram of the deepsort structure in the embodiments of the present application. DETAILED DESCRIPTION
[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0063] Example:
[0064] like Figure 1 As shown, in one embodiment of the present invention, the present invention provides a rolling stone identification and tracking method based on improved YOLO-Deepsort, comprising the following steps:
[0065] S1. Obtain the rolling stone image dataset and preprocess it to obtain the rolling stone image training set and validation set.
[0066] S1 includes the following steps:
[0067] S11. Obtain the Rolling Stones image dataset;
[0068] S12. Select images from the rockfall image dataset one by one and perform data augmentation preprocessing to expand the rockfall image dataset;
[0069] S13. Randomly select a portion of the rolling stone images from the expanded rolling stone image dataset to form the rolling stone image training set, and use the remaining rolling stone images from the rolling stone image dataset to form the rolling stone image validation set.
[0070] S2. Based on the MobileNetV4 module and WTConv module, the yolov11 network model is improved to construct a rolling stone target recognition model.
[0071] like Figure 2 As shown, the improved YOLO-Deepsort rolling stone recognition model includes a MobileNetV4 backbone network, a WTConv Neck network connected to the MobileNetV4 backbone network, and a Head network connected to the WTConv Neck network.
[0072] The MobileNetV4 backbone network comprises, in sequence, a rolling stone picture input module, a first CBS module, a second CBS module, a first UIB module, a second UIB module and a third CBS module; an output end of the rolling stone picture input module serves as an image input end of the MobileNetV4 backbone network; a first rolling stone feature output end of the first UIB module, a second rolling stone feature output end of the second UIB module and a third rolling stone feature output end of the third CBS module are all connected with the WTConv Neck network.
[0073] As shown in Figure 3 The first UIB module and the second UIB module each comprise a first DepthConv module, a second DepthConv module, an ExpandConv module and a ProjectConv module; an input end of the first DepthConv module serves as a first rolling stone picture input end; an output end of the first DepthConv module is connected with an input end of the ExpandConv module; an output end of the ExpandConv module is connected with an input end of the second DepthConv module; an output end of the second DepthConv module is connected with an input end of the ProjectConv module; the first DepthConv module, the ExpandConv module, the second DepthConv module and the ProjectConv module sequentially perform convolution on the input picture to obtain a UIB rolling stone feature map.
[0074] The WTConv Neck network comprises an SPPF module, a C2PSA module, a first WTC3 module, a second WTC3 module, a third WTC3 module, a fourth WTC3 module, a fourth CBS module, a fifth CBS module, a first up-sampling module, a second up-sampling module, a first Concat splicing module, a second Concat splicing module, a third Concat splicing module, and a fourth Concat splicing module; an output end of the SPPF module is connected with a first input end of the C2PSA module; a first output end of the C2PSA module is connected with a first input end of the fourth Concat splicing module; a second output end of the C2PSA module is connected with an input end of the first up-sampling module; an output end of the first up-sampling module is connected with a first input end of the first Concat splicing module; an output end of the first Concat splicing module is connected with an input end of the first WTC3 module; a first output end of the first WTC3 module is connected with a first input end of the third Concat splicing module; a second output end of the first WTC3 module is connected with an input end of the second up-sampling module; an output end of the second up-sampling module is connected with a first input end of the second Concat splicing module; an output end of the second Concat splicing module is connected with an input end of the second WTC3 module; a first output end of the second WTC3 module is connected with an input end of the fourth CBS module; an output end of the fourth CBS module is connected with a second input end of the third Concat splicing module; an output end of the third Concat splicing module is connected with an input end of the third WTC3 module; a first output end of the third WTC3 module is connected with an input end of the fifth CBS module; an output end of the fifth CBS module is connected with a second input end of the fourth Concat splicing module; an output end of the fourth Concat splicing module is connected with an input end of the fourth WTC3 module.
[0075] As shown in Figure 4 the first WTC3 module, the second WTC3 module, the third WTC3 module, and the fourth WTC3 module each comprise a first WTConv module, a Split module, n Bottleneck modules connected in sequence, a fifth Concat splicing module, and a second WTConv module.
[0076] The output end of the first WTConv module is connected with the input end of a Split module; the first output end of the Split module is connected with the input end of a first Bottleneck module; the second output end of the Split module is connected with the first input end of a fifth Concat splicing module; the output end of an nth Bottleneck module is connected with the second input end of the fifth Concat splicing module; the output end of the fifth Concat splicing module is connected with the input end of a second WTConv module; the output end of the second WTConv module obtains a WTC3 rolling stone feature map, wherein n is a positive integer.
[0077] The Head network comprises a first prediction head connected with the second output end of a second WTC3 module, a second prediction head connected with the second output end of a third WTC3 module, and a third prediction head connected with the output end of a fourth WTC3 module.
[0078] S3, based on the obtained rolling stone target recognition network, a deepsort rolling stone tracking model is constructed.
[0079] As shown in Figure 2 , the improved yolo-deepsort rolling stone tracking model comprises a deepsort network connected with a Head network;
[0080] The deepsort network comprises a feature extraction module, an association matching module and a periodical updating module connected with the first prediction head, the second prediction head and the third prediction head.
[0081] As shown in Figure 5 , the feature extraction module extracts the predicted position of the rolling stone by using Kalman filtering; the association matching module completes the optimal matching of the predicted position of the rolling stone by using IOU cascade matching; and the periodical updating module manages the life cycle of the rolling stone track and updates the predicted state of the rolling stone.
[0082] The calculation expression of the Kalman filtering is as follows:
[0083]
[0084] Among them, is the state estimation of the rolling stone in the current frame, is the state estimation of the rolling stone in the last frame, F k is the state transition matrix of the rolling stone in the current frame, K k is the Kalman gain of the rolling stone, z k is the detection box observation value of the rolling stone, H k is the observation matrix of the rolling stone;
[0085] The calculation expression of the IOU cascade matching is as follows:
[0086] D(A, B) = S(A∩B) / S(A∪B)
[0087] C = 1 - D(A, B)
[0088] M = min_π∑{(i, j)∈π}C_{ij}
[0089] Wherein, A is the bounding box coordinate of the rockfall prediction trajectory, B is the bounding box coordinate of the rockfall detection result, S(A∩B) is the area of the overlapping region of the two boxes, S(A∪B) is the total coverage area of the two boxes, D(A, B) is the ratio of the two areas, C is the rockfall matching cost in the interval [0, 1], π is all possible rockfall matching combinations, C_{ij} is the cost of the ith rockfall prediction box and the jth rockfall detection box, and M is the minimum total cost under the optimal rockfall matching;
[0090] The data association calculation expression is as follows:
[0091] S = α·D + β·M + γ·A
[0092]
[0093] Wherein, S is the comprehensive matching score of the rockfall, α, β, γ are weight coefficients in the interval [0, 1], D is the spatial coincidence degree of the rockfall, M is the rockfall similarity, A is the rockfall cosine similarity, H is the rockfall confirmation threshold, L is the rockfall deletion threshold, F is the number of previous unmatched times of the rockfall, N is the maximum tolerance loss frame number of the rockfall, T is the trajectory state, 1 is the confirmed state, 2 is the tentative state, and 3 is the deletion state.
[0094] S4, training and verifying the improved recognition and tracking network by using the rockfall picture training set and the rockfall image verification set to obtain a rockfall recognition and tracking model; in the embodiment, the training parameters adopt the cosine annealing learning rate scheduling (initial value 0.01, minimum 0.001), cooperate with the tracking threshold mechanism (the threshold value of the appearance feature correlation is 0.2, and the threshold value of the intersection-over-union correlation is 0.7), and realize the stable improvement of the model performance in 300 training rounds through the combination of the integrated SGD optimizer and the weight decay.
[0095] The S4 includes the following steps:
[0096] S41, randomly selecting a rockfall picture in the rockfall picture training set, and inputting the selected rockfall picture into the improved yolo-deepsort network model;
[0097] S42, feature extraction is performed on the selected rolling stone picture by using the MobileNetV4 backbone network, and first MobileNetV4 backbone rolling stone feature maps, second MobileNetV4 backbone rolling stone feature maps and third MobileNetV4 backbone rolling stone feature maps are extracted from the first rolling stone feature output end, the second rolling stone feature output end and the third rolling stone feature output end in turn;
[0098] S43, the first MobileNetV4 backbone rolling stone feature maps, the second MobileNetV4 backbone rolling stone feature maps and the third MobileNetV4 backbone rolling stone feature maps are up-sampled, feature-spliced and feature-extracted by using the WTConv Neck network, to obtain first rolling stone prediction feature maps, second rolling stone prediction feature maps and third rolling stone prediction feature maps;
[0099] S44, the first rolling stone prediction feature maps are predicted by using the first prediction head in the Head network, the second rolling stone prediction feature maps are predicted by using the second prediction head in the Head network, and the third rolling stone prediction feature maps are predicted by using the third prediction head in the Head network, to obtain a rolling stone detection map with target boxes and confidence;
[0100] S45, the rolling stone detection map with target boxes and confidence is feature-extracted, correlation-matched and periodically updated by using the deepsort network, to obtain a rolling stone tracking map with time sequence information and motion trajectory;
[0101] S46, the S41-S45 stage is repeated for a training number of times, the yolo-deepsort network model after the stage training is verified by using a rolling stone picture validation set, and the hyperparameters at the time of the verification are saved;
[0102] S47, the S46 is repeated for a preset number of times, and based on the saved hyperparameters, the hyperparameters corresponding to the optimal network depth are obtained, and a rolling stone recognition and tracking network model is obtained.
[0103] In the embodiment, part of the rolling stone picture data set is taken as a test set, and the rolling stone recognition and tracking network model is tested based on the test set. When testing, the yolo-deepsort model is improved and compared for testing. The test results are shown in Table 1: Table 1
[0104] In Table 1, MAP:@0.5 represents the average precision of the rockfall recognition category at an intersection-over-union value of 0.5, parameters represent the parameter amount of the recognition model; MOTA is used to evaluate the tracking effect, which measures the overall accuracy of rockfall tracking, and comprehensively considers the missed detection, false detection and ID switching of the target. The results show that after the model is added to perform MobileNetV4 and WTCONV optimization, the precision is greatly improved, the parameter amount is also decreased to a certain extent, and the deepsort rockfall tracking algorithm based on the improved YOLO model also has certain stability.
[0105] In the embodiment, the output result is a rockfall recognition and tracking video or picture with an ID and a target frame and corresponding time sequence information, the target frame is used to describe the position and size of the recognized rockfall target, the ID is used to describe the individual measure of each tracked rockfall target in the rockfall video or picture, such as 1, 2 and the like, the time sequence information is used to describe the pixel point position information of the rockfall target in the video or picture; at the same time, the model outputs a track line for the rockfall video data to describe the track of each tracked rockfall.
[0106] The rockfall recognition and tracking method based on the improved YOLO-deepsort provided by the application solves the problems of low rockfall recognition efficiency and weak extraction ability of features such as motion track and speed change.
[0107] The above is only a specific embodiment of the application, but the protection scope of the application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the application, which should be covered within the protection scope of the application.
Claims
1. A rockfall identification and tracking method based on improved yolo-deepsort, characterized in that, Comprising the following steps: S1, obtain the rockfall picture dataset and preprocess to obtain the rockfall picture training set and the validation set; S2, improve the yolov11 network model based on the MobileNetV4 module and the WTConv module, construct the rockfall target recognition model, and then obtain the rockfall target recognition network; S3, based on the obtained rockfall target recognition network, construct the deepsort rockfall tracking model, and then obtain the improved recognition and tracking network; S4, train and verify the improved recognition and tracking network using the rockfall picture training set and the rockfall image validation set, and obtain the rockfall recognition and tracking model; S5, obtain the rockfall picture and video to be detected, and use the obtained recognition and tracking model to predict it, and obtain the recognition and tracking result of the rockfall picture and video to be detected.
2. The rockfall identification and tracking method based on YOLO-DeepSort according to claim 1, characterized in that, The S1 comprises the following steps: S11, obtain the rockfall picture dataset; S12, select the pictures in the rockfall picture dataset one by one, and perform data enhancement preprocessing to expand the rockfall picture dataset; S13, randomly select a part of the expanded rockfall picture dataset to form a rockfall picture training set, and the remaining rockfall picture dataset to form a rockfall picture validation set.
3. The rockfall identification and tracking method based on improved yolo-deepsort according to claim 1, characterized in that, The improved yolo-deepsort rockfall recognition and tracking model comprises a MobileNetV4 backbone network, a WTConv Neck network connected with the MobileNetV4 backbone network, a Head network connected with the WTConv Neck network, and a deepsort network connected with the Head network; The MobileNetV4 backbone network comprises a rockfall picture input module, a first CBS module, a second CBS module, a first UIB module, a second UIB module, and a third CBS module connected in sequence; the output end of the rockfall picture input module is connected with the image input end of the MobileNetV4 backbone network; the first rockfall feature output end of the first UIB module, the second rockfall feature output end of the second UIB module, and the third rockfall feature output end of the third CBS module are all connected with the WTConv Neck network; The WTConv Neck network comprises an SPPF module, a C2PSA module, a first WTC3 module, a second WTC3 module, a third WTC3 module, a fourth WTC3 module, a fourth CBS module, a fifth CBS module, a first up-sampling module, a second up-sampling module, a first Concat splicing module, a second Concat splicing module, a third Concat splicing module, and a fourth Concat splicing module; an output end of the SPPF module is connected with an input end of the C2PSA module; a first output end of the C2PSA module is connected with a first input end of the fourth Concat splicing module; a second output end of the C2PSA module is connected with an input end of the first up-sampling module; an output end of the first up-sampling module is connected with a first input end of the first Concat splicing module; an output end of the first Concat splicing module is connected with an input end of the first WTC3 module; a first output end of the first WTC3 module is connected with a first input end of the third Concat splicing module; a second output end of the first WTC3 module is connected with an input end of the second up-sampling module; an output end of the second up-sampling module is connected with a first input end of the second Concat splicing module; an output end of the second Concat splicing module is connected with an input end of the second WTC3 module; a first output end of the second WTC3 module is connected with an input end of the fourth CBS module; an output end of the fourth CBS module is connected with a second input end of the third Concat splicing module; an output end of the third Concat splicing module is connected with an input end of the third WTC3 module; a first output end of the third WTC3 module is connected with an input end of the fifth CBS module; an output end of the fifth CBS module is connected with a second input end of the fourth Concat splicing module; an output end of the fourth Concat splicing module is connected with an input end of the fourth WTC3 module; The Head network comprises a first prediction head connected with a second output end of the second WTC3 module, a second prediction head connected with a second output end of the third WTC3 module, and a third prediction head connected with an output end of the fourth WTC3 module; The deepsort network comprises a feature extraction module connected with the first prediction head, the second prediction head, and the third prediction head, an association matching module, and a periodical updating module.
4. The rockfall identification and tracking method based on improved yolo-deepsort according to claim 1, characterized in that, The S4 comprises the following steps: S41, randomly selecting a rockfall picture in a rockfall picture training set, and inputting the selected rockfall picture into the improved yolo-deepsort network model; S42, performing feature extraction on the selected rockfall picture by using a MobileNetV4 backbone network, and sequentially extracting a first MobileNetV4 backbone rockfall feature map, a second MobileNetV4 backbone rockfall feature map, and a third MobileNetV4 backbone rockfall feature map from a first rockfall feature output end, a second rockfall feature output end, and a third rockfall feature output end; S43, the first MobileNetV4 backbone rolling stone feature map, the second MobileNetV4 backbone rolling stone feature map and the third MobileNetV4 backbone rolling stone feature map are up-sampled, feature spliced and feature extracted by using the WTConv Neck network, and the first rolling stone prediction feature map, the second rolling stone prediction feature map and the third rolling stone prediction feature map are obtained; S44, the first rolling stone prediction feature map is predicted by using the first prediction head in the Head network, the second rolling stone prediction feature map is predicted by using the second prediction head in the Head network, and the third rolling stone prediction feature map is predicted by using the third prediction head in the Head network, and the rolling stone detection map with target frame and confidence is obtained; S45, the rolling stone detection map with target frame and confidence is extracted, matched and periodically updated by using the deepsort network, and the rolling stone tracking map with time sequence information and motion trajectory is obtained; S46, the training times of S41-S45 are repeated, the yolo-deepsort network model after the stage training is verified by using the rolling stone picture verification set, and the hyperparameters at this time are saved; S47, the S46 is repeated for a predetermined number of times, and based on the saved hyperparameters, the hyperparameters corresponding to the optimal network depth are obtained, and the rolling stone recognition and tracking network model is obtained.
5. The rockfall identification and tracking method based on improved yolo-deepsort according to claim 3, characterized in that, The first UIB module and the second UIB module each include a first DepthConv module, a second DepthConv module, an ExpandConv module and a ProjectConv module; an input end of the first DepthConv module is a first rolling stone picture input end; an output end of the first DepthConv module is connected with an input end of the ExpandConv module; an output end of the ExpandConv module is connected with an input end of the second DepthConv module; an output end of the second DepthConv module is connected with an input end of the ProjectConv module; the first DepthConv module, the ExpandConv module, the second DepthConv module and the ProjectConv module sequentially convolve the input picture to obtain a UIB rolling stone feature map.
6. The rockfall identification and tracking method based on improved yolo-deepsort according to claim 3, characterized in that, The first WTC3 module, the second WTC3 module, the third WTC3 module and the fourth WTC3 module each include a first WTConv module, a Split module, n Bottleneck modules connected in sequence, a fifth Contact splicing module and a second WTConv module; An output end of the first WTConv module is connected with an input end of a Split module; a first output end of the Split module is connected with an input end of a first Bottleneck module; a second output end of the Split module is connected with a first input end of a fifth Concat splicing module; an output end of an nth Bottleneck module is connected with a second input end of the fifth Concat splicing module; an output end of the fifth Concat splicing module is connected with an input end of a second WTConv module; and an output end of the second WTConv module obtains a WTC3 rockfall prediction feature map, wherein n is a positive integer.
7. The rockfall identification and tracking method based on improved yolo-deepsort according to claim 3, characterized in that, The feature extraction module extracts a predicted position of a rockfall by using Kalman filtering; the association matching module completes optimal matching of the predicted position of the rockfall by using IOU cascaded matching; and the periodical updating module manages a rockfall trajectory life cycle and updates a rockfall prediction state by using data association.
8. The rockfall identification and tracking method based on improved yolo-deepsort according to claim 7, characterized in that, The Kalman filtering calculation expression is as follows: ; wherein, is a current frame state estimate of the rolling stone, is a previous frame state estimate of the rolling stone, is a current frame state transition matrix of the rolling stone, is a Kalman gain of the rolling stone, is a bounding box observation of the rolling stone, is an observation matrix of the rolling stone; The IOU cascaded matching calculation expression is as follows: ; ; ; wherein, is the bounding box coordinate of the rock prediction trajectory, is the bounding box coordinate of the rock detection result, is the area of the overlapping region of the two boxes, is the total coverage area of the two boxes, is the ratio of the two areas, is the rock matching cost in the interval [0, 1], is all possible rock matching combinations, is the cost of the ith rock prediction box and the jth rock detection box, is the minimum total cost under the optimal rock matching. The data association calculation expression is as follows: ; ; wherein, is the rock comprehensive matching score, , , is the weight coefficient in the interval [0, 1], is the rock spatial coincidence degree, is the rock similarity degree, is the rock cosine similarity degree, is the rock confirmation threshold, is the rock deletion threshold, is the rock number of previous unmatched times, is the rock maximum tolerable loss frame number, is the track state, 1 is the confirmed state, 2 is the tentative state, and 3 is the deletion state.