Wind embracing identification method, system and device and storage medium
Through the improved YOLOv8m network, the horn rod and brake cylinder head are identified, their effectiveness is judged and length is calculated, and the safety hazards caused by the brake shoe holding the wheels during the railway hump slip operation are solved, and efficient brake holding risk warning and detection accuracy are achieved.
Patent Information
- Application Number
- CN202510498899.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
During the railway hump slip operation, the brake shoe holds the wheels and causes the vehicle to carry the gate, causing accidents such as stoppage, conflict, and derailment, which seriously affects the efficiency of railway operations and poses safety hazards.
Using the improved YOLOv8m network, by acquiring and preprocessing image data, the object detection model is trained to identify the horn rod and the head of the brake cylinder, judge the effectiveness of the horn rod and calculate its length, and issue a brake risk warning when the horn rod length exceeds the preset threshold.
It realizes efficient detection and positioning of horn rods, with detection accuracy reaching millimeter level and recognition rate as high as 99.9%, effectively warning and reducing railway accidents.
Smart Images

Figure CN120014568A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of railway brake detection, and in particular to a wind brake identification method, system, equipment and storage medium. Background Art
[0002] During the railway hump sliding operation, if the brake shoes hold the wheels and brake, the vehicle will operate with the brakes on. The sliding process of the vehicle with brakes will cause it to stop on the way, or even collide or derail, which will seriously affect the railway operation efficiency and may even cause safety accidents.
[0003] Intelligent identification of railway brakes is an important technology in the field of railway transportation safety. Braking refers to the situation when the vehicle to be released is not exhausted as required and the manual brake machine is not released as required during the hump dismantling and shunting operation, causing all the brake shoes of the vehicle to remain close to the wheel tread. Braking will cause the vehicle to stop on the way, resulting in frequent collisions, derailment and other accidents. The brakes used by railway freight cars to stop and control speed during operation are divided into two categories: one is air braking and the other is manual braking. Air braking is to inject air into the air cylinder to push the piston to extend the bellows rod to drive the brake shoe to hold the wheel, so that the brake shoe rubs against the wheel to perform railway braking; manual braking is to manually rotate the manual brake machine, so that the manual brake chain pulls the brake shoe to hold the wheel, so that the brake shoe rubs against the wheel to perform railway braking. Wind holding refers to the situation when the vehicle is released during the hump operation, the air is not exhausted as required, causing the brake shoe to hold the wheel, so that the vehicle is operated with the brake.
[0004] In recent years, with the continuous emergence of new models in the field of deep learning, vehicle intelligent detection technology based on computer vision has also made great progress. In the field of object detection, mainstream algorithms are mainly divided into two categories: (1) Single-stage model: This type of model simplifies the object detection task into a regression and classification problem, and directly predicts the category and bounding box of the object by inputting the image. Due to its simple structure, the single-stage model has a high detection speed and is suitable for real-time detection scenarios and devices with limited computing power. (2) Two-stage model: Unlike the single-stage model, the two-stage model divides the detection task into two stages. The first stage generates candidate regions, and the second stage further classifies and regresses these candidate regions. By gradually refining the target features, the model performs better in small target detection and complex scenes. However, the multi-stage calculation characteristics make the model structure more complex, which increases the computational overhead accordingly.
[0005] Therefore, it is worth studying how to use deep learning to dynamically detect and warn of railway brake risks in real time. Summary of the invention
[0006] In view of the above problems, the present invention proposes a wind resistance identification method, system, device and storage medium to try to solve or alleviate one or more of the above problems.
[0007] According to one aspect of the present invention, a wind resistance identification method is proposed, the method comprising: Get a training image dataset; Preprocess the training image data; The preprocessed training image dataset is input into the target detection model based on the improved YOLOv8m network for training to obtain the trained target detection model; Input the image to be detected into the trained target detection model for detection, and obtain the detection result; the detection result is whether the image to be detected contains the bellows rod and the brake cylinder head; If the image to be detected contains a bellows rod and a brake cylinder head, whether the bellows rod is a valid bellows rod is determined based on the positional relationship between the bellows rod and the brake cylinder head; the judgment standard of the valid bellows rod is: the brake cylinder head in the image is located within a certain pixel range around the bellows rod; When there is a valid bellows rod in the image to be detected, the bellows rod length is calculated; the bellows rod length is compared with a preset length threshold, and when the bellows rod length exceeds the preset length threshold, a brake risk warning is issued.
[0008] Furthermore, the improvements of the improved YOLOv8m network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m backbone network; and adding a coordinate attention mechanism after the neck network and before the head network.
[0009] Furthermore, the step of inputting the image to be detected into a trained target detection model for detection and obtaining the detection result includes: The image to be detected is input into the backbone network for feature extraction, including: using multiple C2f feature extraction modules in the backbone network to extract local features of the image and perform feature fusion; then, aggregating features of different scales through a spatial pyramid pooling module; then, through a partial self-attention module, the aggregated feature map is evenly divided into two parts, one part of the feature map enters the self-attention module for global information modeling, and key features are extracted through matrix operations between query vectors, key vectors and value vectors, and the other part of the feature map is fused with the output of the self-attention module through a jump connection; The neck network is used to fuse, enhance and process the features extracted by the backbone network; The coordinate attention mechanism is used to generate a feature map with enhanced position information, including: first, performing global average pooling on the features output by the neck network in the horizontal direction and the vertical direction respectively, and extracting compressed features in different directions; then, these features are fused through a set of convolutions with shared weights; the fused features are re-divided into two parts in the horizontal direction and the vertical direction, and the attention matrix is generated through 1×1 convolution and Sigmoid activation function respectively; finally, the features output by the neck network are weighted by the attention matrix through element-wise multiplication to obtain a feature map with enhanced position information; The head network is used to transform the feature map enhanced with position information into the final object detection result.
[0010] Furthermore, the aggregating features of different scales through the spatial pyramid pooling module includes: processing features of different scales through convolutional layers, batch normalization and SiLU activation functions to obtain feature maps with half the number of channels; passing through multiple maximum pooling layers in sequence to gradually reduce the size of the feature map; and splicing the feature maps after pooling of different scales in the channel dimension.
[0011] Furthermore, the dimension of the query vector and the key vector in the partial self-attention module is half of the value vector; batch normalization is used instead of layer normalization for normalization.
[0012] Furthermore, the preprocessing includes data enhancement, and the data enhancement includes rotation, flipping, scaling, and brightness adjustment.
[0013] Furthermore, the target detection model based on the improved YOLOv8m network introduces regularization technology during the training process, and the regularization technology includes weight decay and Dropout.
[0014] According to another aspect of the present invention, a wind resistance identification system is provided, the system comprising: The detection model training module includes a training data acquisition submodule, a preprocessing submodule, and a model training submodule; the training data acquisition submodule is configured to acquire a training image data set; the preprocessing submodule is configured to preprocess the training image data; the model training submodule is configured to input the preprocessed training image data set into a target detection model based on an improved YOLOv8m network for training, and acquire a trained target detection model; The target detection module is configured to input the image to be detected into the trained target detection model for detection and obtain a detection result; the detection result is whether the image to be detected contains the bellows rod and the brake cylinder head; The brake holding identification module is configured to determine whether the bellows rod is a valid bellows rod based on the positional relationship between the bellows rod and the brake cylinder head if the bellows rod and the brake cylinder head are included in the image to be detected; the judgment standard of the valid bellows rod is: the brake cylinder head in the image is located within a certain pixel range around the bellows rod; when there is a valid bellows rod in the image to be detected, the length of the bellows rod is calculated; the length of the bellows rod is compared with a preset length threshold, and when the length of the bellows rod exceeds the preset length threshold, a brake holding risk warning is issued.
[0015] According to another aspect of the present invention, an electronic device is also proposed, which includes a memory, a processor and a computer program; wherein the computer program is stored in the memory and is configured to be executed by the processor to implement the above-mentioned wind hugging identification method.
[0016] According to another aspect of the present invention, a computer-readable storage medium is further provided, wherein the storage medium stores a computer program; the computer program is executed by a processor to implement the above-mentioned wind resistance identification method.
[0017] The beneficial technical effects of the present invention are: The present invention proposes a wind-holding identification method, system, device and storage medium, which uses an improved YOLOv8m network to detect whether a bellows rod and a brake cylinder head are contained in an image; if the bellows rod and the brake cylinder head are contained, the length of the bellows rod is calculated; when the length of the bellows rod exceeds the preset length threshold, a brake risk warning is issued; wherein, a spatial pyramid pooling module and a partial self-attention module are introduced in the fourth stage of the YOLOv8m original backbone network; and a coordinate attention mechanism is introduced in the third stage of the neck network. The present invention adaptively adjusts the feature weights so that the detection model can focus on the target boundary more accurately, thereby realizing efficient detection and positioning of the bellows rod. Experiments show that the accuracy of the target detection of the present invention can reach the millimeter level, and the recognition rate of the bellows rod is as high as 99.9%. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, in which: Figure 1 It is a flow chart of a wind resistance identification method described in an embodiment of the present invention.
[0019] Figure 2 1 is another flow chart of a wind resistance identification method described in an embodiment of the present invention.
[0020] Figure 3 It is a schematic diagram of the structure of the improved YOLOv8m network in an embodiment of the present invention.
[0021] Figure 4 It is a structural schematic diagram of a wind hugging identification system described in an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0023] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, apparatus, method or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software. In this article, it is to be understood that any number of elements in the drawings is for illustration and not for limitation, and any naming is only for distinction and does not have any limiting meaning.
[0024] The wind holding identification system adopts automatic control technology to automatically collect panoramic color images of the bottom of the vehicle and build an efficient recognition model based on deep learning. It can automatically identify and locate the brake cylinder and its piston stroke components at the bottom for different models, and issue risk warnings based on the target shape changes of the brake cylinder piston stroke. The acquisition equipment is preset in the train track part. When the magnetic steel detects the signal of the incoming vehicle, the system obtains continuous video frame data as input. Subsequently, the input data is processed by the main program and the intelligent detection module. During the processing, the main program receives the image detection notification from the intelligent detection module, and if it is considered that there is a risk of holding the brake, a warning prompt is issued in time.
[0025] Therefore, the embodiment of the present invention proposes a wind resistance identification method, such as Figure 1~2 As shown, the method includes: S1, obtain a training image dataset; S2, preprocessing the training image data; S3, input the preprocessed training image data set into the target detection model based on the improved YOLOv8m network for training, and obtain the trained target detection model; S4, inputting the image to be detected into the trained target detection model for detection, and obtaining a detection result; the detection result is whether the image to be detected contains the bellows rod and the brake cylinder head; S5. If the image to be detected includes a bellows rod and a brake cylinder head, determining whether the bellows rod is a valid bellows rod based on the positional relationship between the bellows rod and the brake cylinder head; the criterion for determining whether the bellows rod is a valid bellows rod is that the brake cylinder head in the image is located within a certain pixel range around the bellows rod; S6. Calculating the length of the bellows rod when there is a valid bellows rod in the image to be detected; S7. Compare the length of the bellows rod with a preset length threshold, and when the length of the bellows rod exceeds the preset length threshold, issue a brake risk warning.
[0026] The method begins in S1. In S1, a training image dataset is obtained.
[0027] According to the embodiment of the present invention, the training and testing of the target detection model uses a self-made data set, which is a collection of field data collected from multiple stations and is divided into a training set and a test set in a ratio of 8:2. The data set contains images of two key components, the brake cylinder and the bellows rod, which are finely annotated by a professional team according to strict annotation standards, ensuring the high quality of the data and the accuracy of the annotation.
[0028] During the data collection process, the actual environmental differences of different stations were fully considered to enhance the model's adaptability to diverse scenarios. For example, the collection tasks covered a variety of light conditions, including strong light during the day, low light at night, and shadow areas; the diversity of equipment, including brake cylinders and bellows levers of various models; and the complexity of the background, including raindrops, leaves and debris obstructions.
[0029] Then S2 is executed, in which the training image data is preprocessed.
[0030] According to an embodiment of the present invention, in order to further improve the generalization ability of the model for diverse scenarios, data enhancement operations are performed on the training data. These enhancements include rotation, flipping, scaling, brightness adjustment, etc., so that the model can adapt to complex actual environments.
[0031] Then, S3 is executed. In S3, the preprocessed training image dataset is input into the target detection model based on the improved YOLOv8m network for training to obtain the trained target detection model.
[0032] According to an embodiment of the present invention, in order to balance the accuracy and reasoning speed of the model, YOLOv8m, which is more mature in technology in the single-stage model, is selected as the basic framework of the model. As a member of the YOLO series, YOLOv8m is also composed of three networks: the backbone, the neck, and the head. Specifically, the backbone network effectively captures the local and global information of the image by stacking the C2f feature extraction module designed with multiple gradient flows; the neck network adopts the idea of path aggregation network, and extracts effective information from feature maps at multiple levels by contracting and expanding the path, thereby enhancing the detection ability of small objects and multi-scale targets; the head network is responsible for completing the final target detection task, including category prediction, bounding box regression, and target confidence estimation.
[0033] In complex scenes, since the background and the target share similar texture or color features, the feature extraction process of the model is easily disturbed, resulting in false detection and missed detection. In order to better meet the detection requirements of the task target, this paper makes targeted improvements and optimizations based on YOLOv8m. The specific architecture of the improved model is as follows: Figure 3 As shown in the figure, the improved model structure still consists of three parts: the backbone, the neck and the head. The improvements include: introducing the spatial pyramid pooling module and the partial self-attention module in the fourth stage of the backbone network; introducing the coordinate attention mechanism in the third stage of the neck network. The purpose of this is to further enhance the model's ability to accurately locate the target boundary, so that it can better adapt to the actual needs of the system for detection accuracy while maintaining efficient reasoning speed.
[0034] The backbone network is responsible for extracting multi-level and multi-scale universal features from the input image to form a feature map with high-level semantic information to provide support for subsequent detection tasks. First, the C2f feature extraction module, as the basic unit of the network, is responsible for extracting local features at each stage and enhancing the expression of multi-scale information through feature fusion. This module adopts the CSP structure design and divides the input feature map into two branches through 1×1 convolution: one branch is connected through multiple residual modules. The first branch is processed, and the other branch directly performs convolution operations. Subsequently, the outputs of the two branches are spliced through the Concat operation, and after batch normalization (BN) and SiLU activation function processing, the features are finally sorted through the convolution operation to obtain the final output. By stacking multiple C2f modules, the network can learn rich feature expressions at different levels, while optimizing the fusion of high-level semantic information and low-level detail information at different stages. The formula of the C2f feature extraction module can be expressed as follows:
[0035] in, is the input feature, is the output feature, represents 1×1 convolution, Indicates the channel splitting operation. is the residual module, Represents a feature concatenation operation. It is the intermediate feature of the module processing process and is the output obtained after N times of residual module processing.
[0036] Subsequently, the spatial pyramid pooling module (Spatial Pyramid Pooling-Fast, SPPF) aggregates features of different scales, so that the network can better retain semantic information when processing complex scenes and generate richer feature representations. The SPPF module first processes the input feature map through convolutional layers, batch normalization (BN) and SiLU activation functions to obtain a feature map with half the number of channels. Then, the feature map passes through three 5×5 maximum pooling layers in sequence, and the size of the feature map gradually decreases after each pooling. The output of each pooling layer will be used as the input of the next pooling. This serial pooling process can effectively capture information at different scales and improve the model's ability to capture details and global information in the image. Finally, by splicing the feature maps after pooling at different scales in the channel dimension, the SPPF module can integrate feature information from different scales to form a multi-scale feature representation. The formula of the spatial pyramid pooling module can be expressed as follows:
[0037] in, is the input feature, is the output feature of the module, represents a 5×5 maximum pooling operation, Represents the output obtained after k times of maximum pooling.
[0038] When the model separates the target and background features, it often inevitably introduces unnecessary background information to interfere with the target features, resulting in false detection and missed detection. To solve this problem, the usual practice is to introduce an attention mechanism deep in the model to highlight the target features. However, the existing mainstream attention mechanisms are mostly implemented through convolution operators. The inherent limitations of convolution operators make it difficult to establish long-distance dependencies between features, which is precisely the key to accurately positioning the target. In contrast, the Transformer structure can efficiently capture global long-range dependencies with its self-attention mechanism. However, the Transformer has high computational complexity and memory usage, which often brings a large time overhead in real-time reasoning tasks.
[0039] Therefore, a partial self-attention module (PSA) is introduced into the backbone network. The design of the PSA module aims to effectively enhance the network's ability to model long-distance dependencies while avoiding the computational overhead caused by the global self-attention mechanism. Specifically, the input feature map is evenly divided into two parts, so as to control the complexity of the self-attention calculation within a low range. One part of the features is input into the self-attention module for global information modeling to capture long-distance dependencies and enhance the contextual understanding of the target; global information modeling is performed through matrix operations between the query vector (Query), key vector (Key) and value vector (Value) to capture the long-range correlation between features; the other part of the features is fused with the output of the self-attention module through jump connections. In addition, in order to further improve the inference efficiency, the PSA module optimizes the dimensions of the query vector and key vector in the self-attention mechanism, setting their dimensions to half of the value vector, thereby reducing the amount of calculation. At the same time, the module uses batch normalization (BN) instead of layer normalization (LN) for standardization, thereby improving the running speed and stability of the model. This module is placed after the fourth stage with the lowest resolution in the model, focusing on extracting and sorting out key information in high-level abstract features. Due to the low feature resolution of the fourth stage, the quadratic complexity of self-attention calculation is significantly reduced, so the overall inference speed can still meet the real-time requirements. The formula of some self-attention modules can be expressed as follows:
[0040] in, is the input feature, is the output feature, Represents the Transformer module. Through the above operations, the long-range dependencies between features can be effectively captured without significantly increasing the computational overhead, so as to better process the semantic information in the deep layer of the network.
[0041] The neck network is used to integrate the features extracted from the backbone network for fusion, enhancement and processing to help the model better detect objects of different scales. Specifically, the neck network receives the features of the P3, P4 and P5 levels in the backbone network as input, first enhances the low-level features through a bottom-up information transmission path, and then fuses them with high-level features to ensure that the network can obtain rich multi-level feature representations. Subsequently, the network diffuses the detail information to deeper levels of the network through a top-down information transmission path, enhances the interaction between high-level semantic features and low-level detail features, and further improves the model's sensitivity to details. The combination of bottom-up and top-down information flow ensures that the network still maintains efficient detection performance in complex scenarios, especially showing significant advantages in multi-scale object detection tasks.
[0042] In order to further improve the model's ability to extract target location information, the coordinate attention mechanism (CA) is introduced in the third-stage output part of the neck network and before the head network to enhance the ability to accurately locate the target boundary. The formula of the coordinate attention mechanism can be expressed as follows:
[0043] in, is the input feature, is the output feature, and Represents the average pooling operation in the horizontal and vertical directions of the feature, is the Sigmoid activation function.
[0044] The head network is responsible for converting the feature maps processed from the trunk and neck into the final target detection results. Specifically, the network contains 3 detection heads, each of which can be divided into two parts, each consisting of 2 3×3 convolutions and one 1×1 convolution, which are used to predict coordinate regression information and category confidence information respectively.
[0045] In order to prevent the model from overfitting, regularization techniques such as weight decay and Dropout are introduced during the model training process to improve the robustness of the model.
[0046] Then, S4 is executed. In S4, the image to be detected is input into the trained target detection model for detection to obtain a detection result; the detection result is whether the image to be detected contains the bellows rod and the brake cylinder head.
[0047] According to an embodiment of the present invention, it is determined whether a detection target (bellows rod and brake cylinder head) exists in a certain frame of image based on the output of the model, thereby deciding whether to perform subsequent calculations or end the detection of the frame.
[0048] Then execute S5~S7. In S5, if the image to be detected contains a bellows rod and a brake cylinder head, determine whether the bellows rod is a valid bellows rod based on the positional relationship between the bellows rod and the brake cylinder head; the judgment standard of the valid bellows rod is: the brake cylinder head in the image is located within a certain pixel range around the bellows rod; in S6, calculate the length of the bellows rod when there is a valid bellows rod in the image to be detected; in S7, compare the length of the bellows rod with a preset length threshold, and when the length of the bellows rod exceeds the preset length threshold, issue a brake risk warning.
[0049] According to an embodiment of the present invention, if the bellows rod and the brake cylinder head are detected, it is further determined whether the brake cylinder head is within 50 pixels around the bellows rod. If this condition is met, the subsequent calculations will continue; otherwise, the detection of this frame will end. When the positional relationship between the bellows rod and the brake cylinder meets the requirements, it is determined whether there is a complete brake cylinder head in the image based on the aspect ratio of the brake cylinder. If a complete brake cylinder head is detected, the brake cylinder size corresponding to the bottom component of the vehicle model is matched according to the model output, and the length of the bellows rod is automatically calibrated and calculated. Since the diameter of the brake cylinder head is a known fixed value, the actual physical length corresponding to a single pixel can be calculated based on the model output result. By further analyzing the number of detected bellows rod pixels, its length can be accurately calculated; the specific formula is:
[0050] in, and are the actual lengths of the bellows rod and the diameter of the brake cylinder head, respectively. and The number of pixels for the bellows rod length and the brake cylinder head diameter, respectively.
[0051] Subsequently, it is determined whether the length exceeds the set threshold, and a notification is sent to the main program based on the determination result. If it exceeds, an early warning notification is issued. If the complete brake cylinder head is not detected, the detection of the current frame will end.
[0052] Specifically, the bellows rod is usually longer than 50 mm, which is greater than 50 pixels when converted into pixels. The maximum range of the image recognition area is reduced by 50 pixels; for example, the image pixels , the effective range is 50 to 1230 in width and 50 to 670 in height. If the bellows rod appears in the effective range, it is identified as a valid bellows rod, otherwise it is invalid data, which can prevent the problem of other edge devices identifying it as a bellows rod. Furthermore, the width of the bellows rod is generally less than 1 / 2 of the air cylinder. If the width of the bellows rod is less than 1 / 2 of the air cylinder, it is a valid bellows rod, otherwise it is invalid data, which can prevent the problem of the bellows rod being disproportionate to the air cylinder. Furthermore, the air cylinder and the bellows rod are connected together and will not overlap. The overlap ratio is set not to exceed 50% of the bellows rod area, which can effectively prevent the problem of fuzzy identification of the brake cylinder and the bellows rod. Furthermore, in a carriage, the number of bellows rods exceeding the threshold / the total number of recognized brake cylinder data accounts for more than 50% of the valid data, which can prevent problems caused by angle recognition of the bellows rod being too long.
[0053] The technical effect of the present invention is further verified through experiments.
[0054] The experiment uses multiple datasets collected on site for training and testing. The dataset contains two types of images, gate cylinder and bellows, totaling 17,251 images. To ensure the effectiveness of the model in practical applications, the dataset has been strictly labeled and screened to cover targets in different scenes and lighting conditions. The model is built based on the PyTorch deep learning framework, with a batch size of 16, and an SGD optimizer with an initial learning rate of 0.01 and a momentum of 0.937, for a total of 300 cycles. To ensure the stability and convergence speed of the training process, a learning rate scheduling strategy is adopted to reduce the learning rate by 10 times every 100 cycles to achieve a balance between model exploration and convergence.
[0055] In order to achieve the best detection accuracy, detailed adjustments and optimizations were made for data enhancement. The final data enhancement parameters are shown in Table 1.
[0056] Table 1 Data augmentation parameters
[0057] In order to verify the advancedness of the detection model, the method proposed in the present invention is compared with advanced methods in the field, including Faster R-CNN, SSD, YOLOv5m, YOLOv8m, etc. Faster R-CNN is a two-stage target detection model, and the other methods are single-stage target detection models. In order to ensure a fair comparison, the public codes of these methods are used to reproduce their network structures, and the models are trained and evaluated under the same training environment. The experiment used the same hyperparameter settings, data sets and evaluation indicators to ensure the comparability of the results. As shown in Table 2, the model proposed in the present invention achieved the highest evaluation index in the comparison with other target detection methods, proving its advantage in target detection accuracy.
[0058] Table 2 Comparison results of the present invention with other target detection methods
[0059] In order to verify the effectiveness of the introduced modules, an ablation experiment was conducted. The experiment used YOLOv8m as the baseline model, and added partial self-attention modules to the backbone network and coordinate attention mechanisms to the head network in turn to verify the contribution and effectiveness of each module. The experimental results are shown in Table 3. Among them, the partial self-attention module effectively enhanced the model's understanding of the global context and improved the detection accuracy; the coordinate attention mechanism further improved the ability to locate the target position, especially in complex backgrounds. These results show that the design of each module has practical value, and its combination can significantly improve the overall model performance, verifying the effectiveness and rationality of the method.
[0060] Table 3 Experimental results
[0061] In order to improve the inference speed, the detection model was converted from PyTorch to TensorRT to make full use of hardware acceleration. The inference GPU deployed on site is NVIDIA GeForce RTX 4070, which supports FP16 and FP32 calculations, both of which provide 29.15 TFLOPS of computing performance. Through TensorRT optimization, the inference speed of the model has been significantly improved, and the inference time of a single image (including pre- and post-processing) is about 12ms. In order to further improve the processing efficiency, multi-threaded inference was adopted during the deployment process to enhance the concurrent processing capability. Finally, the optimized and deployed wind hug recognition system can process image data captured by multiple cameras in real time at a speed of 125 frames per second in actual operation on site, meeting the needs of real-time monitoring and detection.
[0062] After testing, the system can accurately identify and locate targets, quickly determine the status changes of key components, and meet the needs of industrial-grade high-precision detection; the detection accuracy of the Fengbao intelligent recognition system reaches ±5mm, demonstrating its good performance and reliability.
[0063] This paper proposes a wind-holding recognition method, which uses YOLOv8m as the core algorithm and is customized according to task requirements. It introduces a spatial pyramid pooling module, a partial self-attention module, and a coordinate attention mechanism to cope with the high requirements for precise positioning in the task; by adaptively adjusting the feature weights, the model can focus on the target boundary more accurately, thereby achieving efficient detection and positioning of the bellows. Thanks to the above optimizations and improvements, the target detection accuracy is at the millimeter level, and the recognition rate of the bellows is as high as 99.9%.
[0064] Another embodiment of the present invention provides a wind resistance identification system, such as Figure 4 As shown, the system includes: The detection model training module 410 includes a training data acquisition submodule 4110, a preprocessing submodule 4120, and a model training submodule 4130; the training data acquisition submodule 4110 is configured to acquire a training image data set; the preprocessing submodule 4120 is configured to preprocess the training image data; the model training submodule 4130 is configured to input the preprocessed training image data set into a target detection model based on an improved YOLOv8m network for training, and acquire a trained target detection model; The target detection module 420 is configured to input the image to be detected into the trained target detection model for detection and obtain a detection result; the detection result is whether the image to be detected contains the bellows rod and the brake cylinder head; The brake holding identification module 430 is configured to determine whether the bellows rod is a valid bellows rod based on the positional relationship between the bellows rod and the brake cylinder head if the image to be detected contains the bellows rod and the brake cylinder head; the judgment standard of the valid bellows rod is: the brake cylinder head in the image is located within a certain pixel range around the bellows rod; when there is a valid bellows rod in the image to be detected, the length of the bellows rod is calculated; the length of the bellows rod is compared with a preset length threshold, and when the length of the bellows rod exceeds the preset length threshold, a brake holding risk warning is issued.
[0065] For the undetailed parts of a wind resistance identification system according to an embodiment of the present invention, please refer to the above detailed description of the method embodiment.
[0066] The method of the present invention can be executed in an electronic device. The electronic device can be any device with storage and computing capabilities, which can be implemented as a server, a workstation, etc., or as a personal computer such as a desktop computer or a notebook computer, or as a terminal device such as a mobile phone, a tablet computer, a smart wearable device, an Internet of Things device, but is not limited thereto.
[0067] The electronic device may include: a processor, a memory, an input / output interface, a communication interface and a bus. The processor, the memory, the input / output interface and the communication interface are connected to each other through the bus in the electronic device. The processor may be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., for executing relevant programs to implement the technical solutions provided in the embodiments of this specification. The memory may be implemented in the form of ROM, RAM, static storage devices, dynamic storage devices, etc. The memory may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory and called and executed by the processor. The input / output interface is used to connect the input / output module to realize information input and output. The input / output / module may be configured in the electronic device as a component, or it may be externally connected to the electronic device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc. The communication interface is used to connect the communication module to realize the communication interaction between the electronic device and other devices. The communication module may realize communication by wire or by wireless. A bus comprises a pathway that transfers information between the various components of an electronic device.
[0068] The embodiment of the present invention also provides a non-transitory readable storage medium, which stores instructions, and the instructions are used to enable the electronic device to perform a method according to an embodiment of the present invention. The readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be a computer-readable instruction, a data structure, a module of a program, or other data. Examples of readable storage media include, but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage, etc.
[0069] It should be noted that the terms used in the present invention are only for describing specific embodiments, rather than limiting the scope of the present application. As shown in the present specification, unless the context clearly indicates an exception, the words "one", "a", "a kind of" and / or "the" do not specifically refer to the singular, but may also include the plural. The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of more restrictions, the elements defined by the sentence "include one..." do not exclude the presence of other identical elements in the process, method or device including the elements.
[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying wind resistance, characterized in that: include: Get a training image dataset; Preprocess the training image data; The preprocessed training image dataset is input into the target detection model based on the improved YOLOv8m network for training to obtain the trained target detection model; The improvements of the improved YOLOv8m network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m backbone network; adding a coordinate attention mechanism after the neck network and before the head network; The image to be detected is input into the trained target detection model for detection to obtain the detection result; the detection result is whether the image to be detected contains the bellows rod and the brake cylinder head, including: inputting the image to be detected into the backbone network for feature extraction, including: using multiple C2f feature extraction modules in the backbone network to extract local features of the image and perform feature fusion; then, aggregating features of different scales through the spatial pyramid pooling module; then, the aggregated feature map is evenly divided into two parts through the partial self-attention module, one part of the feature map enters the self-attention module to perform global information modeling through matrix operations between the query vector, the key vector and the value vector; the other part of the feature map is fused with the output of the self-attention module through a jump connection; using the neck The neck network fuses, enhances and processes the features extracted by the backbone network; the coordinate attention mechanism is used to generate a feature map with enhanced position information, including: first, performing global average pooling on the features output by the neck network in the horizontal and vertical directions respectively, and extracting compressed features in different directions; then, these features are fused through a set of convolutions with shared weights; the fused features are re-divided into two parts in the horizontal and vertical directions, and the attention matrix is generated through 1×1 convolution and Sigmoid activation function respectively; finally, the features output by the neck network are weighted by the attention matrix through element-wise multiplication to obtain a feature map with enhanced position information; the feature map with enhanced position information is converted into the final target detection result using the head network; If the image to be detected contains a bellows rod and a brake cylinder head, whether the bellows rod is a valid bellows rod is determined based on the positional relationship between the bellows rod and the brake cylinder head; the judgment standard of the valid bellows rod is: the brake cylinder head in the image is located within a certain pixel range around the bellows rod; When there is a valid bellows rod in the image to be detected, the bellows rod length is calculated; the bellows rod length is compared with a preset length threshold, and when the bellows rod length exceeds the preset length threshold, a brake risk warning is issued.
2. A method for identifying wind resistance according to claim 1, characterized in that: The aggregating features of different scales through the spatial pyramid pooling module includes: processing features of different scales through convolution layers, batch normalization and SiLU activation functions to obtain feature maps with half the number of channels; passing through multiple maximum pooling layers in sequence to gradually reduce the size of the feature map; and splicing the feature maps after pooling of different scales in the channel dimension.
3. A method for identifying wind resistance according to claim 2, characterized in that: The dimension of the query vector and the key vector in the partial self-attention module is half of the value vector; batch normalization is used instead of layer normalization for standardization.
4. A method for identifying wind resistance according to claim 1, characterized in that: The preprocessing includes data enhancement, and the data enhancement includes rotation, flipping, scaling, and brightness adjustment.
5. A method for identifying wind resistance according to claim 1, characterized in that: The target detection model based on the improved YOLOv8m network introduces regularization technology during the training process, and the regularization technology includes weight decay and Dropout.
6. A wind resistance identification system, characterized in that: include: A detection model training module, including a training data acquisition submodule, a preprocessing submodule, and a model training submodule; The training data acquisition submodule is configured to acquire a training image data set; The preprocessing submodule is configured to preprocess the training image data; The model training submodule is configured to input the preprocessed training image data set into the target detection model based on the improved YOLOv8m network for training, and obtain the trained target detection model; The improvements of the improved YOLOv8m network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m backbone network; adding a coordinate attention mechanism after the neck network and before the head network; The target detection module is configured to input the image to be detected into the trained target detection model for detection and obtain the detection result; the detection result is whether the bellows rod and the brake cylinder head are included in the image to be detected; including: inputting the image to be detected into the backbone network for feature extraction, including: using multiple C2f feature extraction modules in the backbone network to extract local features of the image and perform feature fusion; then, aggregating features of different scales through the spatial pyramid pooling module; then, through the partial self-attention module, evenly dividing the aggregated feature map into two parts, one part of the feature map enters the self-attention module to perform global information modeling through matrix operations between the query vector, the key vector and the value vector; the other part of the feature map is connected with the output of the self-attention module through a jump connection Fusion; using the neck network to fuse, enhance and process the features extracted by the backbone network; using the coordinate attention mechanism to generate a feature map with enhanced position information, including: first, performing global average pooling on the features output by the neck network in the horizontal and vertical directions respectively, extracting compressed features in different directions; then, these features are fused through a set of convolutions with shared weights; the fused features are re-divided into two parts in the horizontal and vertical directions, and the attention matrix is generated through 1×1 convolution and Sigmoid activation function respectively; finally, the features output by the neck network are weighted by the attention matrix through element-wise multiplication to obtain a feature map with enhanced position information; using the head network to convert the feature map with enhanced position information into the final target detection result; The brake holding identification module is configured to determine whether the bellows rod is a valid bellows rod based on the positional relationship between the bellows rod and the brake cylinder head if the bellows rod and the brake cylinder head are included in the image to be detected; the judgment standard of the valid bellows rod is: the brake cylinder head in the image is located within a certain pixel range around the bellows rod; when there is a valid bellows rod in the image to be detected, the length of the bellows rod is calculated; the length of the bellows rod is compared with a preset length threshold, and when the length of the bellows rod exceeds the preset length threshold, a brake holding risk warning is issued.
7. An electronic device, characterized in that: include: A memory, a processor and a computer program; wherein the computer program is stored in the memory and is configured to be executed by the processor to implement the wind hugging identification method described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that The storage medium stores a computer program; the computer program is executed by a processor to implement the wind resistance identification method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Intelligent pre-detection and alarm system for hump humping vehicle
CN115923875A
Submarine cable target detection method based on improved deep network
CN118447363A
Small target detection method, system and device for medical image and storage medium
CN119130891A