Train head detection and identification method, system, medium and equipment in complex weather

Through the feature extraction based on position coding and the axial compression branch and detail enhancement branch of the detection model, combined with the depth-wise separable convolution and multi-head attention mechanism, the problems of low accuracy and poor adaptability in train head and number recognition in complex weather conditions are solved, and high-precision detection and recognition are achieved.

CN120726580APending Publication Date: 2025-09-30NAT HIGH SPEED TRAIN QINGDAO TECH INNOVATION CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510897223.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing technologies for train locomotive detection and number recognition under complex weather conditions have problems of low accuracy and poor adaptability. Especially in cases of diverse background colors, complex locomotive shapes, and poor lighting, existing methods are difficult to achieve efficient and accurate recognition.

Method used

A position-coding-based feature extraction module and detection model are adopted. Through the axial compression branch and detail enhancement branch, combined with the depth-wise separable convolution and multi-head attention mechanism, multi-scale features are extracted and the attention of the detection model is enhanced, thereby improving the recognition ability of vehicle heads and numbers.

Benefits of technology

The detection accuracy of train locomotives and numbers has been significantly improved in low-light and complex weather environments, reducing missed detections and false detections, and improving the stability and recognition capabilities of the detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726580A_ABST
    Figure CN120726580A_ABST
Patent Text Reader

Abstract

The invention provides a train head detection and recognition method, system, medium and equipment in complex weather, and relates to the technical field of image processing, and the method comprises the steps: obtaining a to-be-recognized picture containing a train head; inputting a to-be-recognized picture into the detection model, extracting the to-be-recognized picture in the detection model to obtain feature maps of three resolutions, and performing convolution operation on the feature maps to obtain input features; on the axial compression branch, obtaining attention features of each pixel point; splicing features are obtained on the detail enhancement branches; local detail information is extracted from the splicing features, and detail enhancement features are obtained after compression; multiplying the detail enhancement feature by the attention feature subjected to the activation function, and adding the detail enhancement feature and the input feature to obtain a fusion feature; and obtaining a train head and serial number detection frame from the fused features through a detection model head network. According to the invention, the detection model can more accurately pay attention to the vehicle head and the numbering area in the to-be-identified picture, and the precision and stability of the detection model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method, system, medium and equipment for detecting and identifying a train head in complex weather conditions. Background Art

[0002] With the rapid development of railway transportation, train locomotive detection and number recognition technology plays an increasingly important role in railway safety management, train maintenance and vehicle scheduling. Train locomotive detection and number recognition are the primary links in automatic train maintenance. Efficient and accurate recognition functions can significantly improve the efficiency and accuracy of maintenance operations. However, compared with ordinary car detection and license plate number recognition, train locomotive detection and number recognition face problems such as diverse background colors, complex locomotive shapes, poor maintenance lighting conditions, and complex weather. Traditional detection methods rely on manual operations or simple image processing technologies, which have disadvantages such as low accuracy and poor adaptability. Therefore, those skilled in the art are in urgent need of solving the above technical problems. Summary of the Invention

[0003] The purpose of this application is to provide a method, system, computer-readable storage medium and electronic device for detecting and identifying train locomotives in complex weather conditions, which can improve the detection and identification capabilities of train locomotives in complex weather conditions.

[0004] To solve the above technical problems, this application provides a method for detecting and identifying train locomotives in complex weather conditions. The specific technical solutions are as follows:

[0005] Get the image to be identified that contains the train locomotive;

[0006] Inputting the image to be identified into a detection model, extracting the image to be identified in the detection model to obtain feature maps of three resolutions, and performing a convolution operation on the feature maps to obtain input features; the input features include a first feature, a second feature, and a third feature;

[0007] On the axial compression branch, the first feature, the second feature, and the third feature are sequentially compressed along the horizontal axis, along the vertical axis, and along the channel axis to obtain nine compressed features, and the multi-head attention of the compressed features is calculated to obtain the attention feature of each pixel;

[0008] In the detail enhancement branch, the first feature, the second feature, and the third feature are spliced ​​in the channel dimension to obtain a spliced ​​feature; local detail information is extracted from the spliced ​​feature using a depthwise separable convolution, and the local detail information is compressed to obtain a detail enhancement feature;

[0009] Multiplying the detail enhancement feature by the attention feature after the activation function, and adding the result to the input feature to obtain a fusion feature;

[0010] The fusion features are passed through the regression branch and the classification branch of the detection model to obtain the train head and number detection frame in the image to be identified; the train head and the number detection frame are used to obtain the recognition result of the train number through text recognition.

[0011] Optionally, before inputting the image to be identified into the detection model, the process further includes:

[0012] Obtaining a train head image, and marking the train head and number bounding box in the train head image;

[0013] Processing the train head image for severe weather noise, and expanding the number of images for different types of severe weather to obtain training data;

[0014] Performing position-coding-based feature extraction on the training data to obtain feature maps of three resolutions;

[0015] The feature map is input into the detection head for feature classification and bounding box regression to obtain a detection model.

[0016] Optionally, the train head image is subjected to severe weather noise processing, and the number of images is expanded for different types of severe weather to obtain training data, including:

[0017] Using Gaussian distribution to simulate raindrop brightness, and randomly generating the center position of raindrops, the raindrops are superimposed on the train head image to obtain rainy day train head data; the raindrop brightness does not exceed the pixel range of the train head image;

[0018] Randomly generate the center position of the snowflake, and generate the snowflake brightness distribution according to the snowflake cycle length combined with the snowflake brightness generation formula, and superimpose the snowflake brightness distribution on the train head image to obtain the snow train head data;

[0019] The center position of the fog cluster is simulated based on the atmospheric scattering effect, and the fog generation distribution is obtained according to the fog concentration and the fog generation formula. The fog generation distribution is superimposed on the train head image to obtain the foggy train head data;

[0020] The training data is obtained according to the train locomotive picture, the rainy day train locomotive data, the snowy day train locomotive data and the foggy day train locomotive data.

[0021] Optionally, feature extraction based on position encoding is performed on the training data to obtain feature maps of three resolutions, including:

[0022] Extracting basic features of the training data using a feature extraction module;

[0023] The Transformer position encoding module is used to add position information to the basic features to obtain features of different scales;

[0024] The feature fusion module is used to fuse features of different scales to obtain feature maps of three resolutions.

[0025] Optionally, the Transformer position encoding module is used to add position information to the basic features to obtain features of different scales, including:

[0026] Determine the position encoding of each pixel in the feature map;

[0027] The position code is superimposed on the basic features before adding the position information and encoded through a Transformer encoder to obtain features of different scales.

[0028] Optionally, determining the position code of each pixel in the feature map includes:

[0029] The position encoding of each pixel in the feature map is determined according to the encoding position, encoding dimension index, encoding dimension and scale factor.

[0030] Optionally, after the fusion features are passed through the regression branch and the classification branch of the detection model to obtain the train head and number detection frame in the image to be identified, the method further includes:

[0031] Determine a numbered area according to the numbered detection frame;

[0032] Performing histogram equalization on the numbered regions to obtain a first processed image;

[0033] performing wavelet decomposition, soft threshold denoising, and wavelet inverse transformation on the first processed image in sequence to obtain a second processed image;

[0034] performing contrast-limited adaptive histogram equalization on the second processed image to obtain a third processed image;

[0035] An edge detection operator is used to perform image sharpening processing on the third processed image to obtain a high-contrast image; the high-contrast image is used to perform numbered text recognition.

[0036] This application also provides a train head detection and identification system in complex weather conditions, comprising:

[0037] An image acquisition module is used to acquire an image to be identified that includes the locomotive head;

[0038] A feature extraction module is configured to input the image to be identified into a detection model, extract the image to be identified in the detection model to obtain feature maps of three resolutions, and perform a convolution operation on the feature maps to obtain input features; the input features include a first feature, a second feature, and a third feature;

[0039] a module configured to compress the first feature, the second feature, and the third feature in the axial compression branch in sequence along the horizontal axis, along the vertical axis, and along the channel axis to obtain nine compressed features, and calculate the multi-head attention of the compressed features to obtain an attention feature for each pixel;

[0040] a feature enhancement module, configured to, on the detail enhancement branch, concatenate the first feature, the second feature, and the third feature in the channel dimension to obtain a concatenated feature; extract local detail information from the concatenated feature using depthwise separable convolution, and compress the local detail information to obtain a detail enhancement feature;

[0041] A feature fusion module, configured to multiply the detail enhancement feature by the attention feature after the activation function, and add the result to the input feature to obtain a fused feature;

[0042] The image recognition module is used to obtain the train head and number detection frame in the image to be identified by passing the fusion feature through the regression branch and the classification branch of the detection model.

[0043] The present application also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above-described method when executed by a processor.

[0044] The present application also provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps of the above-mentioned method when calling the computer program in the memory.

[0045] The present application provides a method for detecting and identifying train locomotives in complex weather conditions, comprising: obtaining a picture to be identified containing a train locomotive; inputting the picture to be identified into a detection model, extracting the picture to be identified in the detection model to obtain feature maps of three resolutions, and subjecting the feature maps to a convolution operation to obtain input features; the input features include a first feature, a second feature, and a third feature; on an axial compression branch, compressing the first feature, the second feature, and the third feature in turn along the horizontal axis, along the vertical axis, and along the channel axis to obtain nine compressed features and calculating the multi-head attention of the compressed features to obtain an attention feature for each pixel. ; On the detail enhancement branch, the first feature, the second feature and the third feature are spliced ​​in the channel dimension to obtain a spliced ​​feature; local detail information is extracted from the spliced ​​feature using depthwise separable convolution, and the local detail information is compressed to obtain a detail enhancement feature; the detail enhancement feature is multiplied by the attention feature after the activation function, and is added to the input feature to obtain a fusion feature; the fusion feature is passed through the regression branch and the classification branch of the detection model to obtain the train head and number detection frame in the picture to be identified; the train head and the number detection frame are used to obtain the recognition result of the train number through text recognition.

[0046] When identifying the locomotive of a train, in order to further improve the recognition accuracy of key areas in the image to be identified, this application enhances the attention mechanism of the detection model by setting detail enhancement branches and compression branches, so that the detection model can focus more accurately on the locomotive and number areas in the image to be identified, thereby improving the accuracy and stability of the detection model. In particular, in low light and complex weather environments, the detection model's ability to recognize the locomotive and number can be significantly improved, reducing missed detections and false detections.

[0047] The present application also provides a train head detection and identification system, a computer-readable storage medium and an electronic device under complex weather conditions, which have the above-mentioned beneficial effects and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0049] Figure 1 A flow chart of a method for detecting and identifying a train locomotive in complex weather conditions provided by an embodiment of the present application;

[0050] Figure 2A schematic diagram of the network structure corresponding to the train locomotive detection and identification method provided in an embodiment of the present application;

[0051] Figure 3 A schematic diagram of the structure of a feature extraction module based on position coding provided in an embodiment of the present application;

[0052] Figure 4 A schematic diagram of the structure of the axial compression attention module provided in an embodiment of the present application;

[0053] Figure 5 Schematic diagram of the rain, snow and fog data enhancement effect of a train provided in an embodiment of the present application;

[0054] Figure 6 A flow chart of the train locomotive number identification process provided in an embodiment of the present application;

[0055] Figure 7 This is a diagram showing the train locomotive detection and number recognition provided by an embodiment of the present application;

[0056] Figure 8 A schematic diagram of the structure of a train head detection and identification system in complex weather conditions provided by an embodiment of the present application;

[0057] Figure 9 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0058] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0059] The object information involved in this application, including but not limited to the object device information, the object personal information, etc., and the data, including but not limited to data used for analysis, stored data, displayed data, etc., are all information and data authorized by the object or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards.

[0060] In real-world applications, train locomotives and serial numbers may appear at varying scales and positions within images, particularly in complex maintenance environments, where locomotive sizes and positions can vary significantly. Existing detection algorithms, typically based on fixed or a few preset scales, struggle to effectively handle multi-scale and small objects. This results in low detection rates and recognition accuracy for some small objects, such as serial numbers. Consequently, existing algorithms have limitations when handling multi-scale and small objects, preventing them from delivering high-precision detection and recognition results in real-world applications.

[0061] Furthermore, in complex weather conditions like rain, snow, and fog, image clarity and contrast can significantly decrease, and background noise can increase, directly leading to a significant drop in algorithm performance and reduced train locomotive positioning accuracy. Consequently, poor image quality in complex weather conditions creates a significant performance bottleneck for existing algorithms in train locomotive positioning, making them unable to meet the demands of practical applications.

[0062] To solve the above problems, see Figure 1 , Figure 1 A flow chart of a method for detecting and identifying a train locomotive in complex weather conditions provided by an embodiment of the present application, the method comprising:

[0063] S101: Obtaining a picture to be identified containing a train locomotive;

[0064] S102: Inputting the image to be identified into a detection model, extracting the image to be identified in the detection model to obtain feature maps of three resolutions, and performing a convolution operation on the feature maps to obtain input features; the input features include a first feature, a second feature, and a third feature;

[0065] S103: On the axial compression branch, compress the first feature, the second feature, and the third feature along the horizontal axis, the vertical axis, and the channel axis in sequence to obtain nine compressed features, and calculate multi-head attention of the compressed features to obtain an attention feature for each pixel;

[0066] S104: In the detail enhancement branch, the first feature, the second feature, and the third feature are concatenated in the channel dimension to obtain a concatenated feature; local detail information is extracted from the concatenated feature using a depthwise separable convolution, and the local detail information is compressed to obtain a detail enhancement feature;

[0067] S105: multiplying the detail enhancement feature by the attention feature after the activation function, and adding the result to the input feature to obtain a fusion feature;

[0068] S106: The fusion features are passed through the regression branch and the classification branch of the detection model to obtain the train head and number detection frame in the image to be identified.

[0069] The image to be identified collected in step S101 is an image containing the front of a train, but is not limited to the collected scene and the real-time weather of the scene, including various complex weather conditions such as rain, snow, fog, etc.

[0070] Thereafter, in step S102, the image to be identified is identified using the detection model. This embodiment assumes that the detection model has been generated or trained before step S102 is executed, and does not specifically limit how the detection model is trained.

[0071] In step S102, the detection model is used to extract features of the image to be identified. Specifically, it is necessary to extract feature maps of three different resolutions and obtain input features through convolution operations.

[0072] Before feature extraction, the image to be identified may also be preprocessed, including but not limited to operations such as scaling and normalization.

[0073] Afterwards, if Figure 2 As shown, Figure 2 The network structure diagram corresponding to the train head detection and recognition method provided in the embodiment of the present application. In the detection model, feature maps of three resolutions can be obtained through the feature extraction module based on position coding. The image to be identified is sent to the feature extraction module based on position coding PEBFE Module to extract features. Figure 3 As shown, Figure 3 This is a structural diagram of the position coding-based feature extraction module provided in an embodiment of the present application, which passes through the CSPDarknet module, the Transformer position coding module and the PAFPN module in sequence.

[0074] The CSPDarknet module extracts basic features, the Transformer positional encoding module enhances feature position information, and the PAFPN module fuses multi-scale features. Ultimately, three feature maps of different sizes are obtained: the primary, secondary, and tertiary features, corresponding to high, medium, and low resolutions, respectively. By extracting these three features at different resolutions, we can extract rich feature information, providing a solid foundation for subsequent object detection.

[0075] In the task of train and number target detection, deep features are rich in semantic information, while shallow features provide key spatial position information. For the detection of small targets, spatial position information is particularly important because when the resolution of the feature map is reduced, small targets are easily missed. Therefore, the effective combination of high-level semantic information and spatial position information is the core step of target detection. In this embodiment of the application, the train image is input into the feature extraction network CSPDarknet to obtain the feature Next, we fuse features through position encoding. Since the position information of the target in the high-level feature map is relatively small, adding position encoding and Transformer encoder can further enhance the position information of the features, helping the network better understand the spatial position of the target at different scales, thereby achieving accurate detection and positioning of small targets. The operation process is shown in the following formula:

[0076] ;

[0077] in, represents the Transformer encoder, Represents the feature after adding position information and passing through the encoder. PE represents the position encoding of each pixel in the feature map. The two-dimensional sine-cosine position encoding method is used to encode the position of each pixel in the feature map. The position encoding generation process can be expressed by the following formula:

[0078] ;

[0079] ;

[0080] ;

[0081] ;

[0082] ;

[0083] in, Indicates the encoding position, represents the encoding dimension index, represents the encoding dimension, Represents the scale factor.

[0084] .

[0085] Afterwards, the characteristics Generate feature pyramid through PAFPN operation , and then input into the head network for regression and classification.

[0086] The head network is used to perform train head and number detection. Specifically, the three feature maps of different sizes can be The detection head is input for classification and bounding box regression in turn to obtain the detection frame of the train head and number. The Axial Compression Attention Module (ACA Module) is added to the classification branch and the bounding box regression branch, as shown in the following example: Figure 4 As shown, Figure 4A schematic diagram of the axial compression attention module provided in an embodiment of this application. In low-light and complex environments, this module can guide the model to more accurately focus on the locomotive and numbered areas, reducing computational effort while improving detection accuracy of the target area, ultimately obtaining bounding box information for the locomotive and numbered area.

[0087] Among them, ACA Module includes axial compression branch and detail enhancement branch. First, the feature map obtained by feature extraction is The convolution operation obtains Then, by axially compressing the branches, Compression is performed along the horizontal, vertical, and channel axes, and multi-head attention is calculated to extract global semantic features. Finally, the detail enhancement branch is used to compensate for the local details lost during the compression process. The specific operation process is as follows:

[0088] In step S103, in the compression operation in the horizontal axis direction, first Dimensional query feature map The compression operation aggregates global information onto one axis, thus greatly reducing the pressure of global semantic extraction. and Repeat the compression operation above, and finally get , the process is expressed by the following formula:

[0089] ;

[0090] ;

[0091] ;

[0092] in, Represents a tensor The dimension of is converted. Is the length of A vector with all elements equal to 1 is used to implement the sum operation in dimension W, and then divided by W to achieve The dimension-wise mean operation.

[0093] The compression on the vertical axis and channel axis is similar, and we get 、 、 , calculate the multi-head attention separately, and get 、 、 The three attention features of the dimension are added and fused to obtain the attention feature of each pixel. .

[0094] In step S104, in the detail enhancement branch, although the axial compression operation effectively extracts global semantic information, it sacrifices local details. The embodiment of the present application uses depthwise separable convolution to enhance spatial details, such as Figure 4 As shown in the detail enhancement branch, first Splicing is performed on the channel dimension to obtain splicing features , then use Depthwise separable convolution from concatenated features Extract local detail information. And through Convolution will The features of the dimension are compressed to , generating detail enhancement features The process can be expressed by the following formula:

[0095] ;

[0096] ;

[0097] In step S105, the detail enhancement feature and Sutra Attention features of functions Multiply and add to the input features Add to get fusion features The process can be expressed by the following formula:

[0098] ;

[0099] in, is the batch normalization layer, yes The convolution, yes Depthwise separable convolution, express activation function, Represents the Sigmoid nonlinear activation function, Represents a pixel-by-pixel multiplication operation.

[0100] The obtained fusion features are processed by the regression branch and the classification branch of the detection model to obtain the train head and number detection frame in the image to be identified.

[0101] When detecting and identifying train locomotives, the embodiments of the present application further improve the recognition accuracy of key areas in the image to be identified by designing a feature extraction module based on position coding to enhance contextual information, and by setting detail enhancement branches and axial compression branches, enhance the attention mechanism of the detection model, so that the detection model can focus more accurately on the locomotive and number areas in the image to be identified, thereby improving the accuracy and stability of the detection model. In particular, in low light and complex weather environments, the detection model's ability to recognize locomotives and numbers can be significantly improved, reducing missed detections and false detections.

[0102] The following is a feasible detection model generation process provided by this application:

[0103] The first step is data collection and annotation:

[0104] High-resolution cameras can be used to monitor and collect data from trains in real time, covering different train types, shooting angles, lighting conditions, and backgrounds. Images of train locomotives are collected from both maintenance and operational scenes. A dataset of 500 images is used as an example. Labeling tools such as Labelme are used to annotate the locomotives and numbers in these images, with the bounding boxes labeled "locomotive" and "number." This step ensures dataset diversity and annotation accuracy, providing high-quality data support for subsequent model training.

[0105] The second step is data enhancement and division:

[0106] Configure the rain, snow and fog data enhancement method to process the rain, snow and fog noise of the 500 collected and annotated images, such as Figure 5 As shown, Figure 5 This diagram illustrates the enhanced effects of rain, snow, and fog on train data provided in this application example, expanding the diversity of the dataset. 2,000 labeled data images were generated using data enhancement techniques. It should be noted that the rain, fog, and snow weather simulations are independent of each other and can also be superimposed.

[0107] When adding rain noise, a Gaussian distribution is used to simulate the brightness of raindrops, and the center position of raindrops is randomly generated. The raindrops are superimposed on the train head image to obtain the rainy train head data; the raindrop brightness does not exceed the pixel range of the train head image. Specifically, rain noise is generated by simulating the shape, density, and movement trajectory of raindrops. The brightness of raindrops is generated using a Gaussian distribution and superimposed on the image to achieve the rain effect. The center position of raindrops is randomly generated. , the brightness distribution of raindrops is generated according to the formula , add raindrops to the image, ensuring that the brightness value does not exceed the pixel range of the image.

[0108] Raindrop brightness generation:

[0109] ;

[0110] in, is the raindrop's coordinate in the image The brightness value at . It is the center of the raindrop.

[0111] Overlay to image:

[0112] ;

[0113] in, is the pixel value of the original image.

[0114] When adding snow noise, the center position of the snowflake is randomly generated, and the snowflake brightness distribution is generated based on the snowflake cycle length combined with the snowflake brightness generation formula. The snowflake brightness distribution is superimposed on the train head image to obtain the snowy train head data. Specifically, the snow noise can be generated by simulating the high brightness and random distribution of snowflakes. Snowflakes are usually brighter than raindrops and have more irregular shapes. The center position of the randomly generated snowflake is , generate the brightness distribution of snowflakes according to the formula , overlay the snow onto the image, ensuring that the brightness value does not exceed the pixel range of the image.

[0115] Snow brightness generation:

[0116] ;

[0117] Overlay snowflakes onto an image:

[0118] ;

[0119] in, is the snowflake's coordinate in the image The brightness value at is the period length of the snowflake (controls the shape of the snowflake).

[0120] When adding fog noise, the center position of the fog cluster is simulated based on the atmospheric scattering effect, and the fog generation distribution is obtained based on the fog concentration and the fog generation formula. This fog generation distribution is superimposed on the train head image to obtain the foggy train head data. The generation of fog noise can be achieved by simulating the atmospheric scattering effect. Fog will blur the entire image and reduce the contrast. The center position of the fog is randomly generated. , generate the brightness distribution of fog according to the formula , overlay fog onto the image, ensuring that the brightness values ​​do not exceed the pixel range of the image.

[0121] Fog Generation:

[0122] ;

[0123] Overlay fog onto an image:

[0124] ;

[0125] in, is the fog in image coordinates The brightness value at is the width of the fog (the width of the fog is used to indicate and control the density of the fog).

[0126] Afterward, the dataset can be divided into training and validation sets. For example, the dataset can be divided into a training set and a validation set in an 8:2 ratio, with 1,600 images used as the training set and 400 images used as the validation set, thus forming the Complex Weather Train Head Detection and Number Recognition Dataset (CWT-HDNR Dataset). The purpose of this step is to increase the robustness and generalization ability of the model through data augmentation, and to ensure the accuracy of model training and validation through reasonable data partitioning.

[0127] Step 3: Feature extraction:

[0128] The images in the train head detection and number recognition dataset are input into the feature extraction module for feature extraction. First, image preprocessing is performed, including scaling, normalization, etc. Then, the images are sent to the position-encoded feature extraction module PEBFE Module to extract features, such as Figure 3 As shown, Figure 3 The schematic diagram of the structure of the feature extraction module based on position coding provided in the embodiment of the present application is sequentially passed through the CSPDarknet module, the Transformer position coding module and the PAFPN module. The CSPDarknet module is used to extract basic features. , Transformer position encoding module is used to enhance the position information of features and output ,PAFPN module is used to fuse multi-scale features. After processing, three feature maps of different sizes are obtained , which are high-resolution, medium-resolution and low-resolution feature maps respectively.

[0129] Each scale feature map is sequentially fused through the axial compression attention module to obtain the fused feature These fused features will be input into the detection head for feature classification and bounding box regression to obtain the detection model.

[0130] The feature extraction and classification process will not be repeated here. The fused features output by the axial compression attention module are sent to the regression branch and the classification branch through the convolution operation to obtain the detection frame of the train head and number, and complete the training of the detection model.

[0131] The following is a number text recognition process provided by this application:

[0132] After obtaining the train head and number detection frame in the image to be identified, Figure 6 , Figure 6 The train locomotive number identification process flow chart provided in the embodiment of the present application can be implemented as follows:

[0133] S601: Determine a numbered area according to the numbered detection frame;

[0134] S602: performing histogram equalization on the numbered regions to obtain a first processed image;

[0135] S603: performing wavelet decomposition, soft threshold denoising, and wavelet inverse transformation on the first processed image in sequence to obtain a second processed image;

[0136] S604: performing contrast-limited adaptive histogram equalization on the second processed image to obtain a third processed image;

[0137] S605: Performing image sharpening processing on the third processed image using an edge detection operator to obtain a high-contrast image; the high-contrast image is used to perform numbered text recognition.

[0138] During the train number recognition process, the numbered area can be cropped based on the detected train number bounding box. An image interference removal and enhancement module (IDRE Module) is then used to process the numbered area to reduce the impact of noise such as rain, snow, and fog, thereby improving image quality. The goal is to optimize the image quality of the numbered area and provide clear input for subsequent text recognition. This IDRE module is independent of the train head detection and recognition process and is designed to handle images affected by weather interference such as rain, snow, and fog. Through a multi-step process, image quality is significantly improved. The core steps include:

[0139] Histogram Equalization: First, use the cv2.equalizeHist function in the OpenCV library to perform histogram equalization on the image to improve the overall contrast. The formula for histogram equalization can be expressed as:

[0140] ;

[0141] Where I(x,y) is the pixel value of the input image at position (x,y), and c is a normalization constant.

[0142] Next, the image is subjected to wavelet transform, and the soft threshold denoising method is used to remove noise and retain edge information. The main steps of wavelet transform denoising include wavelet decomposition, soft threshold denoising and inverse wavelet transform. The specific formula is as follows:

[0143] Wavelet decomposition:

[0144] ;

[0145] Where I is the input image and wavelet is the waveform function used.

[0146] Soft threshold denoising:

[0147] ;

[0148] in, is the threshold, determined according to the noise level.

[0149] Inverse wavelet transform:

[0150] ;

[0151] in, are the wavelet coefficients after denoising.

[0152] Then, use OpenCV's cv2.createCLAHE function to implement contrast-limited adaptive histogram equalization (CLAHE) to enhance local contrast. The formula of CLAHE can be expressed as:

[0153] ;

[0154] in, is the probability density function of the local histogram after contrast limitation.

[0155] Finally, convolution operation is performed using the Laplacian or Sobel operator to achieve image sharpening and enhance edges and details. Sharpening usually uses the Laplacian operator, and the formula is as follows:

[0156] ;

[0157] Where I(x,y) is the pixel value of the input image at position (x,y), is the sharpening strength parameter.

[0158] After being processed by the image interference removal and enhancement module, the image can be ensured to remain clear and high-contrast even in complex weather conditions.

[0159] When recognizing train number text, the image of the number region, processed by the Image Deinterference Enhancement (IDRE) module, is fed into a pre-trained text recognition network. Through vertical compression and a horizontal character-level attention mechanism, sequence prediction is performed, ultimately yielding the train number recognition result. This step aims to achieve high-precision number recognition, ensuring the accuracy and reliability of the final recognition result.

[0160] Through the above-mentioned detailed technical solution implementation steps, this application can effectively solve the shortcomings of the existing technology in train head detection and number identification, realize accurate detection and identification of train heads and numbers, and improve the overall efficiency and accuracy of railway safety management, train maintenance and vehicle scheduling.

[0161] See also Figure 7 , Figure 7 The train locomotive detection and number recognition renderings provided in the embodiment of the present application can clearly identify train locomotives and numbers in various complex environments.

[0162] See also Figure 8 , Figure 8 This is a schematic diagram of the structure of a train head detection and identification system in complex weather conditions provided by an embodiment of the present application. The system includes:

[0163] An image acquisition module is used to acquire an image to be identified that includes the locomotive head;

[0164] A feature extraction module is configured to input the image to be identified into a detection model, extract the image to be identified in the detection model to obtain feature maps of three resolutions, and perform a convolution operation on the feature maps to obtain input features; the input features include a first feature, a second feature, and a third feature;

[0165] a compression processing module, configured to compress the first feature, the second feature, and the third feature in the axial compression branch in sequence along the horizontal axis, along the vertical axis, and along the channel axis to obtain nine compressed features, and calculate multi-head attention of the compressed features to obtain an attention feature for each pixel;

[0166] a feature enhancement module, configured to, on the detail enhancement branch, concatenate the first feature, the second feature, and the third feature in the channel dimension to obtain a concatenated feature; extract local detail information from the concatenated feature using depthwise separable convolution, and compress the local detail information to obtain a detail enhancement feature;

[0167] A feature fusion module, configured to multiply the detail enhancement feature by the attention feature after the activation function, and add the result to the input feature to obtain a fused feature;

[0168] The image recognition module is used to obtain the train head and number detection frame in the image to be identified by passing the fusion feature through the regression branch and the classification branch of the detection model.

[0169] Based on the above embodiment, as a preferred embodiment, it also includes:

[0170] The detection model generation module is used to obtain a train head image and mark the head and number bounding box of the train head image; process the train head image for severe weather noise and expand the number of images for different types of severe weather to obtain training data; perform position coding-based feature extraction on the training data to obtain feature maps of three resolutions; input the feature map into the detection head for feature classification and bounding box regression to obtain the detection model.

[0171] Based on the above embodiment, as a preferred embodiment, the detection model generation module includes:

[0172] The severe weather simulation unit is used to perform the following steps:

[0173] Using Gaussian distribution to simulate raindrop brightness, and randomly generating the center position of raindrops, the raindrops are superimposed on the train head image to obtain rainy day train head data; the raindrop brightness does not exceed the pixel range of the train head image;

[0174] Randomly generate the center position of the snowflake, and generate the snowflake brightness distribution according to the snowflake cycle length combined with the snowflake brightness generation formula, and superimpose the snowflake brightness distribution on the train head image to obtain the snow train head data;

[0175] The center position of the fog cluster is simulated based on the atmospheric scattering effect, and the fog generation distribution is obtained according to the fog concentration and the fog generation formula. The fog generation distribution is superimposed on the train head image to obtain the foggy train head data;

[0176] The training data is obtained according to the train locomotive picture, the rainy day train locomotive data, the snowy day train locomotive data and the foggy day train locomotive data.

[0177] Based on the above embodiment, as a preferred embodiment, the detection model generation module includes:

[0178] The feature extraction unit is used to extract the basic features of the training data using the feature extraction module; add position information to the basic features using the Transformer position encoding module to obtain features of different scales; and fuse the features of different scales using the feature fusion module to obtain feature maps of three resolutions.

[0179] Based on the above embodiments, as a preferred embodiment, the feature extraction unit is a unit for determining the position code of each pixel in the feature map, superimposing the position code with the basic features before adding the position information, and encoding the result through a Transformer encoder to obtain features of different scales.

[0180] Based on the above embodiment, as a preferred embodiment, the feature extraction unit includes:

[0181] The position coding determination subunit is used to determine the position coding of each pixel in the feature map according to the coding position, coding dimension index, coding dimension and scale factor.

[0182] Based on the above embodiment, as a preferred embodiment, it also includes:

[0183] The number text recognition module is used to perform the following steps:

[0184] Determine a numbered area according to the numbered detection frame;

[0185] Performing histogram equalization on the numbered regions to obtain a first processed image;

[0186] performing wavelet decomposition, soft threshold denoising, and wavelet inverse transformation on the first processed image in sequence to obtain a second processed image;

[0187] performing contrast-limited adaptive histogram equalization on the second processed image to obtain a third processed image;

[0188] An edge detection operator is used to perform image sharpening processing on the third processed image to obtain a high-contrast image; the high-contrast image is used to perform numbered text recognition.

[0189] This application also provides a computer-readable storage medium having a computer program stored thereon. When executed, the computer program can implement the steps provided in the above embodiments. The storage medium may include: a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code.

[0190] This application also provides an electronic device, see Figure 9 , a structural diagram of an electronic device provided in an embodiment of the present application, such as Figure 9 As shown, a processor 1410 and a memory 1420 may be included.

[0191] The processor 1410 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1410 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1410 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1410 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1410 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0192] The memory 1420 may include one or more computer-readable storage media, which may be non-transitory. The memory 1420 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 1420 is at least used to store the following computer program 1421, wherein, after the computer program is loaded and executed by the processor 1410, it can implement the relevant steps in the method performed by the electronic device side disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 1420 may also include an operating system 1422 and data 1423, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 1422 may include Windows, Linux, Android, etc.

[0193] In some embodiments, the electronic device may further include a display screen 1430 , an input / output interface 1440 , a communication interface 1450 , a sensor 1460 , a power supply 1470 , and a communication bus 1480 .

[0194] certainly, Figure 9 The structure of the electronic device shown does not constitute a limitation on the electronic device in the embodiment of the present application. In actual applications, the electronic device may include Figure 9 More or fewer components than shown, or combinations of certain components.

[0195] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems provided in the embodiments, since they correspond to the methods provided in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0196] Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core concept of this application. It should be noted that, for those skilled in the art, without departing from the principles of this application, various improvements and modifications may be made to this application, and such improvements and modifications also fall within the scope of protection of this application.

[0197] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

Claims

1. A train head detection and identification method under complex weather conditions, characterized in that: include: Get the image to be identified that contains the train locomotive; Inputting the image to be identified into a detection model, extracting the image to be identified in the detection model to obtain feature maps of three resolutions, and performing a convolution operation on the feature maps to obtain input features; the input features include a first feature, a second feature, and a third feature; On the axial compression branch, the first feature, the second feature, and the third feature are sequentially compressed along the horizontal axis, along the vertical axis, and along the channel axis to obtain nine compressed features, and the multi-head attention of the compressed features is calculated to obtain the attention feature of each pixel; In the detail enhancement branch, the first feature, the second feature, and the third feature are spliced ​​in the channel dimension to obtain a spliced ​​feature; local detail information is extracted from the spliced ​​feature using a depthwise separable convolution, and the local detail information is compressed to obtain a detail enhancement feature; Multiplying the detail enhancement feature by the attention feature after the activation function, and adding the result to the input feature to obtain a fusion feature; The fusion features are passed through the regression branch and the classification branch of the detection model to obtain the train head and number detection frame in the image to be identified; the train head and the number detection frame are used to obtain the recognition result of the train number through text recognition.

2. The method according to claim 1, characterized in that Before inputting the image to be identified into the detection model, the method further includes: Obtaining a train head image, and marking the train head and number bounding box in the train head image; Processing the train head image for severe weather noise, and expanding the number of images for different types of severe weather to obtain training data; Performing position-coding-based feature extraction on the training data to obtain feature maps of three resolutions; The feature map is input into the detection head for feature classification and bounding box regression to obtain a detection model.

3. The method according to claim 2, characterized in that The train head image is processed for severe weather noise, and the number of images is expanded for different types of severe weather to obtain training data, including: Using Gaussian distribution to simulate raindrop brightness, and randomly generating the center position of raindrops, the raindrops are superimposed on the train head image to obtain rainy day train head data; the raindrop brightness does not exceed the pixel range of the train head image; Randomly generate the center position of the snowflake, and generate the snowflake brightness distribution according to the snowflake cycle length combined with the snowflake brightness generation formula, and superimpose the snowflake brightness distribution on the train head image to obtain the snow train head data; The center position of the fog cluster is simulated based on the atmospheric scattering effect, and the fog generation distribution is obtained according to the fog concentration and the fog generation formula. The fog generation distribution is superimposed on the train head image to obtain the foggy train head data; The training data is obtained according to the train locomotive picture, the rainy day train locomotive data, the snowy day train locomotive data and the foggy day train locomotive data.

4. The method according to claim 2, characterized in that The training data is subjected to feature extraction based on position coding to obtain feature maps of three resolutions: Extracting basic features of the training data using a feature extraction module; The Transformer position encoding module is used to add position information to the basic features to obtain features of different scales; The feature fusion module is used to fuse features of different scales to obtain feature maps of three resolutions.

5. The method according to claim 4, characterized in that The Transformer position encoding module is used to add position information to the basic features to obtain features of different scales, including: Determine the position encoding of each pixel in the feature map; The position code is superimposed on the basic features before adding the position information and encoded through a Transformer encoder to obtain features of different scales.

6. The method according to claim 5, characterized in that Determining the position encoding of each pixel in the feature map includes: The position encoding of each pixel in the feature map is determined according to the encoding position, encoding dimension index, encoding dimension and scale factor.

7. The method according to any one of claims 1 to 6, characterized in that After the fusion features are passed through the regression branch and the classification branch of the detection model to obtain the train head and number detection frame in the image to be identified, the method further includes: Determine the numbered area according to the numbered detection frame; Performing histogram equalization on the numbered regions to obtain a first processed image; performing wavelet decomposition, soft threshold denoising, and wavelet inverse transformation on the first processed image in sequence to obtain a second processed image; performing contrast-limited adaptive histogram equalization on the second processed image to obtain a third processed image; An edge detection operator is used to perform image sharpening processing on the third processed image to obtain a high-contrast image; the high-contrast image is used to perform numbered text recognition.

8. A train head detection and identification system under complex weather conditions, characterized by: include: An image acquisition module is used to acquire an image to be identified that includes the locomotive head; A feature extraction module is configured to input the image to be identified into a detection model, extract the image to be identified in the detection model to obtain feature maps of three resolutions, and perform a convolution operation on the feature maps to obtain input features; the input features include a first feature, a second feature, and a third feature; a compression processing module, configured to compress the first feature, the second feature, and the third feature in the axial compression branch in sequence along the horizontal axis, along the vertical axis, and along the channel axis to obtain nine compressed features, and calculate multi-head attention of the compressed features to obtain an attention feature for each pixel; a feature enhancement module, configured to, on the detail enhancement branch, concatenate the first feature, the second feature, and the third feature in the channel dimension to obtain a concatenated feature; extract local detail information from the concatenated feature using depthwise separable convolution, and compress the local detail information to obtain a detail enhancement feature; A feature fusion module, configured to multiply the detail enhancement feature by the attention feature after the activation function, and add the result to the input feature to obtain a fused feature; The image recognition module is used to obtain the train head and number detection frame in the image to be identified by passing the fusion features through the regression branch and the classification branch of the detection model; the train head and the number detection frame are used to obtain the recognition result of the train number through text recognition.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which implements the steps of the method according to any one of claims 1 to 7 when executed.