Railway ground signal image feature analysis and recognition method and system
By collecting visible light and thermal infrared image data, and combining multimodal image fusion and temporal variation pattern analysis, the problem of low recognition accuracy of railway ground signals in complex environments has been solved, and accurate signal recognition under different lighting and weather conditions has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CREC RAILWAY ELECTRIFICATION RAILWAY OPERATIONS MANAGEMENT
- Filing Date
- 2026-04-14
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, railway ground signals have low recognition accuracy in complex environments, especially in low visibility conditions where they are difficult to accurately reflect the true state of the signal lights, resulting in insufficient recognition reliability.
Visible light and thermal infrared image data of ground signals are collected. Combined with mileage coordinate information and signal type, the target area and status of the signal lights are determined through multimodal image fusion and temporal change pattern analysis. The status recognition model is then used for accurate identification.
Under different lighting and weather conditions, the visual and thermal radiation characteristics of traffic lights are obtained through complementary multimodal image data, the location of traffic lights is accurately located, and the color category and dynamic change status of traffic lights are accurately determined through time series analysis, thereby improving the reliability of signal recognition.
Smart Images

Figure CN122024209B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of railway signal recognition technology, and in particular to a method and system for analyzing and recognizing railway ground signal image features. Background Technology
[0002] The railway ground signal image feature analysis and recognition method is mainly applied to signal confirmation scenarios during shunting and section operations of railway catenary work vehicles. With the development of railway transportation towards intelligence and the increasing demand for all-weather operation, assisting drivers in recognizing the status of ground signals through image processing technology has become an important technical direction for improving operational safety.
[0003] Current common railway signal image recognition methods mainly use a single visible light camera to capture images of the signal lights, and then use image processing algorithms to extract the color and shape features of the signal lights for recognition. These methods can achieve good recognition results in daylight or under good lighting conditions. Some solutions also combine the fixed position information of the signal lights along the railway line to assist in judging the recognition results. These methods have been applied in some railway operation scenarios.
[0004] However, in the actual railway operation environment, the signal recognition process is easily affected by changes in light and weather factors such as day and night alternation, rain, snow, fog and haze. This makes it difficult for a single visible light image to accurately reflect the true state of the signal lights under low visibility conditions. At the same time, the color change pattern of the signal over a continuous period of time is directly related to traffic safety. Relying solely on the static features of a single frame image is insufficient to adapt to complex and ever-changing on-site conditions. Therefore, existing technologies suffer from insufficient reliability in recognizing railway ground signals in complex environments. Summary of the Invention
[0005] This application provides a method and system for analyzing and recognizing railway ground signal images to solve the problem of low recognition accuracy of ground signals by railway operating vehicles in complex environments in the prior art.
[0006] To address the aforementioned technical problems, in a first aspect, this application provides a method for analyzing and recognizing railway ground signal image features, comprising: Collect visible light image data and thermal infrared image data of ground signals, and obtain the mileage coordinate information and signal type information of ground signals along the railway line; Based on the visible light image data and the thermal infrared image data, the target area of the traffic light in the image is determined; Based on the real-time positioning data of the work vehicle, the predicted mileage coordinates of the ground signal corresponding to the target area along the railway line are calculated. The predicted mileage coordinates are matched with the mileage coordinate information, and combined with the signal type information, the target signal type corresponding to the target area is determined. The feature vector sequence corresponding to the target region in multiple consecutive frames of images and the target signal type are input into the state recognition model to retrieve the corresponding signal state change rules. The change pattern of the feature vector sequence is compared with the signal state change rules to determine the color category and signal state of the signal. The coordinates of the target area, the color category, and the signal status are combined into a recognition result, which is then encapsulated according to preset railway communication requirements and sent to the driver's cab display terminal.
[0007] Secondly, this application provides a railway ground signal image feature analysis and recognition system, comprising: The acquisition module is used to acquire visible light image data and thermal infrared image data of ground signals, and to obtain the mileage coordinate information and signal type information of the ground signals along the railway line; The determination module is used to determine the target area of the traffic light in the image based on the visible light image data and the thermal infrared image data; The calculation module is used to calculate the predicted mileage coordinates of the ground signal corresponding to the target area along the railway line based on the real-time positioning data of the work vehicle, match the predicted mileage coordinates with the mileage coordinate information, and combine the signal type information to determine the target signal type corresponding to the target area. The input module is used to input the feature vector sequence corresponding to the target region in multiple consecutive frames of images and the target signal type into the state recognition model, so as to retrieve the corresponding signal state change rules, and compare the change pattern of the feature vector sequence with the signal state change rules to determine the color category and signal state of the signal. The combination module is used to combine the coordinates of the target area, the color category, and the signal status into a recognition result, and then encapsulate it according to preset railway communication requirements before sending it to the driver's cab display terminal.
[0008] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor is configured to execute the computer program to implement the steps of the railway ground signal image feature analysis and recognition method as described in the first aspect above.
[0009] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps of the railway ground signal image feature analysis and recognition method described in the first aspect above.
[0010] The technical solution provided in this application has the following beneficial effects: This application acquires both visible light and thermal infrared image data, enabling complementary acquisition of visual and thermal radiation characteristics of traffic lights under different lighting and weather conditions, providing a more comprehensive data foundation for subsequent analysis. Then, based on the fused image data, the target area of the traffic light in the image is determined, effectively eliminating background interference and accurately locating the traffic light. Next, by combining the real-time positioning data of the work vehicle to calculate predicted mileage coordinates and matching them with pre-stored information, the specific type of the current traffic light can be quickly determined, providing an accurate type basis for subsequent identification. Subsequently, the feature vector sequence of the target area in multiple consecutive frames of images and the traffic light type are input into the state recognition model. By comparing the feature change patterns with the signal state change rules, the color category and dynamic change state of the traffic light can be accurately determined. Finally, the recognition results are encapsulated and sent to the driver's cab display terminal, enabling the driver to obtain signal information in a timely manner to assist in operational decisions.
[0011] Furthermore, the feature vectors of the target region in multiple consecutive frames are arranged in chronological order to form a feature vector sequence and input into the state recognition model. The feature encoding module generates a temporal feature map to characterize the trend of feature changes between consecutive frames. The rule matching module retrieves the corresponding color change temporal rules according to the target signal type. The temporal analysis module performs inter-frame difference detection on the temporal feature map to identify the frame position where the color feature value jumps and generates a set of jump time intervals. Finally, the state decision module compares the color jump order and each time interval with the color appearance order and duration range specified in the rules item by item. Based on the comparison results, the color category and signal state of the signal in the current frame image are generated.
[0012] Furthermore, by analyzing the variation patterns of feature vectors in multiple consecutive frames of images and comparing them item by item with the inherent color change timing rules of the traffic lights, the color category and current state of the traffic lights during dynamic changes can be accurately identified, effectively avoiding misjudgments that may occur with single-frame image recognition.
[0013] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 A flowchart illustrating a railway ground signal image feature analysis and recognition method provided in this application embodiment; Figure 2 This is a schematic diagram illustrating a specific implementation of a railway ground signal image feature analysis and recognition method provided in this application embodiment; Figure 3 This is a schematic diagram of the structure of a railway ground signal image feature analysis and recognition system provided in an embodiment of this application. Detailed Implementation
[0016] To address the problems existing in the prior art, this application proposes a method for analyzing and recognizing railway ground signal image features. This method enhances signal perception capabilities in complex environments through multimodal image fusion and introduces temporal change pattern analysis to accurately determine the dynamic state of the signal, effectively solving the problem of insufficient signal recognition reliability caused by environmental interference and static recognition limitations in the prior art.
[0017] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0018] The core of this application is to provide a method for feature analysis and recognition of railway ground signal images, and a flowchart of one specific implementation is shown below. Figure 1 As shown, the method includes: Step 101: Collect visible light image data and thermal infrared image data of the ground signal, and obtain the mileage coordinate information and signal type information of the ground signal along the railway line.
[0019] In step 101, visible light image data refers to image information collected by a visible light camera that reflects the appearance, color, and shape of the signal; thermal infrared image data refers to image information collected by an infrared thermal imager that reflects the temperature distribution of the signal's luminous body; ground signal refers to signaling equipment fixedly installed beside the railway line that transmits train driving instructions to the train driver through the color and display status of the lights; railway line refers to the strip-shaped area traversed by the railway track, including the track itself and the ancillary facilities on both sides related to train driving safety; mileage coordinate information refers to the specific mileage position of the ground signal on the railway line; and signal type information refers to the identifier used to distinguish the different functions or structures of the signal.
[0020] In this embodiment, the ground signal is synchronously exposed by an image acquisition device installed on a railway overhead contact line work vehicle to acquire visible light image data and thermal infrared image data. At the same time, the mileage coordinates of the ground signal along the railway line and the corresponding signal type information are read from a preset signal location database to complete the process of acquiring basic data.
[0021] Step 102: Based on the visible light image data and the thermal infrared image data, determine the target area of the traffic light in the image.
[0022] Here, a signal light refers to a light-emitting device on a ground signal machine used to emit light signals, and the target area refers to the exact image area occupied by the signal light in the image.
[0023] In this embodiment, step 102 includes the following process: Step 1021: Based on the pre-stored joint calibration parameters, perform spatial coordinate system registration on the visible light image data and the thermal infrared image data to obtain the registered multimodal image data.
[0024] In step 1021, the joint calibration parameters include the internal parameters of the visible light camera, the internal parameters of the infrared thermal imager, and the external parameters between the two sensors. The internal parameters include the lens focal length and principal point coordinates, and the external parameters include the rotation matrix and translation vector between the two sensors. The method for obtaining the joint calibration parameters is to simultaneously photograph a calibration board containing feature points, extract the coordinates of feature points in the visible light image and the thermal infrared image, establish and solve geometric constraint equations based on multiple sets of corresponding feature points, and obtain the spatial position correspondence data between the coordinate systems of the two sensors.
[0025] Registered multimodal image data refers to the combined data of visible light images and thermal infrared images after spatial coordinate system alignment.
[0026] In this embodiment, pre-stored joint calibration parameters are first read from the memory, and then the image coordinate system corresponding to the visible light image data and the image coordinate system corresponding to the thermal infrared image data are spatially transformed using the joint calibration parameters, so that the pixel positions corresponding to the same physical point in the two images achieve coordinate unification. The visible light image data and the thermal infrared image data after registration processing together constitute the registered multimodal image data.
[0027] Step 1022: Analyze the thermal infrared image data in the registered multimodal image data using a region segmentation algorithm to extract candidate signal regions, and map the coordinates of the candidate signal regions onto the visible light image data in the registered multimodal image data.
[0028] In step 1022, the candidate signal region refers to the image region that may contain the light source of the traffic light, which is initially identified through thermal infrared image analysis. The coordinates of the candidate signal region refer to the position information of the region in the image coordinate system.
[0029] In this embodiment, the thermal infrared image data in the registered multimodal image data is processed by a region segmentation algorithm. The algorithm divides adjacent pixels with similar temperature features into the same connected region based on the similarity of the temperature features of the pixels in the image, and extracts the connected region with a relatively concentrated temperature gradient as a candidate signal region. After obtaining the candidate signal region, the coordinate correspondence established in the previous registration process is used to map the coordinates of the candidate signal region in the thermal infrared image onto the visible light image data in the registered multimodal image data, thereby marking the corresponding candidate region position in the visible light image.
[0030] Step 1023: Verify the candidate signal region in the visible light image data, eliminate interference regions based on the verification results, and take the image region corresponding to the remaining candidate signal region as the target region of the traffic light in the image.
[0031] In step 1023, the interference area refers to the false candidate area that is determined not to belong to the traffic light after verification.
[0032] In this embodiment of the application, after the coordinate mapping is completed, each candidate signal region is verified in the visible light image data. The verification process includes analyzing whether the color distribution of the region conforms to the standard color characteristics of the traffic light and analyzing whether the contour shape of the region conforms to the typical shape characteristics of the traffic light. Candidate signal regions whose color and shape characteristics do not conform to the traffic light standard are identified as interference regions and removed from the candidate set. The image region corresponding to the candidate signal region that is retained after screening is determined as the target region of the traffic light in the image.
[0033] This application achieves accurate extraction of the target area of the traffic light by spatially registering visible light and thermal infrared images through joint calibration parameters, and by combining temperature feature segmentation of thermal infrared images with color and shape verification of visible light images, effectively eliminating high-temperature interference in the environment.
[0034] Step 103: Based on the real-time positioning data of the work vehicle, calculate the predicted mileage coordinates of the ground signal corresponding to the target area along the railway line, match the predicted mileage coordinates with the mileage coordinate information, and combine the signal type information to determine the target signal type corresponding to the target area.
[0035] Among them, real-time positioning data refers to the location information output by the Global Positioning System received by the work vehicle at the current moment, predicted mileage coordinates refer to the estimated mileage position of the ground signal corresponding to the target area on the railway line obtained by calculation, and target signal type refers to the specific category of the ground signal corresponding to the target area after matching and screening.
[0036] In this embodiment, step 103 includes the following process: Step 1031: Extract the current GPS coordinates of the work vehicle from the real-time positioning data, and convert the GPS coordinates into the railway mileage value of the work vehicle based on the preset railway line geographic information table.
[0037] In step 1031, the Global Positioning System coordinates refer to the latitude and longitude values of the current position of the work vehicle obtained through satellite positioning, the railway line geographic information table refers to the pre-established relational data table that maps latitude and longitude coordinates to railway mileage values one-to-one, and the railway mileage value refers to the mileage number value of the current position of the work vehicle on the railway line.
[0038] It should be noted that the specific representation of the railway line geographic information table in this application embodiment is not specifically limited, and can be set accordingly according to the actual situation.
[0039] In this embodiment of the application, the current GPS coordinates of the work vehicle are first extracted from the real-time positioning data. The coordinates include two values: longitude and latitude. Then, a preset railway line geographic information table is retrieved. The table is used to find the longitude and latitude record that is closest to the current GPS coordinates, and the corresponding railway mileage value is read. This railway mileage value is used as the railway mileage value where the work vehicle is currently located.
[0040] In practical applications, assuming the current GPS coordinates of the work vehicle are (118°46′52″E, 31°58′33″N), by searching the railway line geographic information table, the mileage record matching these coordinates is found, and the current railway mileage of the work vehicle is read as K152+300. It should be understood that the aforementioned longitude and latitude values, as well as other related values given subsequently, are virtual and hypothetical geographic information, not actual geographical information.
[0041] Step 1032: Based on the pixel coordinates of the target area in the image, combined with the internal parameters and external installation parameters of the image acquisition device, calculate the distance and azimuth of the ground signal corresponding to the target area relative to the work vehicle, and based on the distance, the azimuth, and the current railway mileage value of the work vehicle, calculate the predicted mileage coordinates of the ground signal corresponding to the target area along the railway line.
[0042] In step 1032, pixel coordinates refer to the row and column positions of the target area in the image; internal parameters refer to the inherent optical characteristics of the image acquisition device itself, including the lens focal length and principal point coordinates; external installation parameters refer to the installation position and attitude parameters of the image acquisition device on the work vehicle, including the height of the device from the rail surface, the pitch angle, and the horizontal deflection angle; distance refers to the straight-line length of space between the ground signal corresponding to the target area and the work vehicle; and azimuth refers to the horizontal deflection angle of the ground signal corresponding to the target area relative to the direction of travel of the work vehicle.
[0043] In this embodiment, the pixel coordinates of the target area in the image are first read, and the internal parameters and external installation parameters of the image acquisition device are obtained simultaneously. Then, based on the pinhole imaging principle, the pixel coordinates are combined with the internal parameters and external installation parameters to perform spatial coordinate transformation calculation, thereby obtaining the distance and azimuth values of the ground signal corresponding to the target area relative to the work vehicle. Finally, the railway mileage value of the work vehicle at its current location is superimposed with the projection component of the distance value in the direction of the railway line to obtain the predicted mileage coordinates of the ground signal corresponding to the target area along the railway line.
[0044] In practical applications, the pixel coordinates of the target area in the image are set to (320, 240), the focal length of the lens of the image acquisition device is 8 mm, the principal point coordinates are (240, 320), the device installation height is 4.5 meters from the rail surface, the pitch angle is 2 degrees downward, and the horizontal deflection angle is 0 degrees. After calculation, the distance between the ground signal corresponding to the target area and the work vehicle is 120 meters, and the azimuth angle is 3 degrees to the right. The current railway mileage of the work vehicle is K152+300. Combining the projection component of the distance of 120 meters in the track direction, the predicted mileage coordinates are calculated to be K152+420.
[0045] Step 1033: Read the mileage coordinate information and signal type information of all ground signals from the preset signal location database, and input the predicted mileage coordinates and each of the mileage coordinate information into the fast matching model based on hash coding. The fast matching model based on hash coding performs hash function mapping on the predicted mileage coordinates and each of the mileage coordinate information respectively to generate the predicted mileage hash value and the hash value of each mileage coordinate.
[0046] In step 1033, the signal location database refers to a pre-stored database containing all ground signal location and type information. The hash-based fast matching model refers to a mathematical model that uses a hash function to map continuous numerical values into binary codes. The hash function refers to a mapping rule that converts input data into fixed-length binary codes. The predicted mileage hash value refers to the binary code obtained after converting the predicted mileage coordinates through a hash function. The mileage coordinate hash value refers to the binary code obtained after converting the mileage coordinate information of each ground signal through a hash function.
[0047] This application does not impose specific limitations on the model type, internal structure design, parameter design, training process, etc. of the fast matching model based on hash encoding. The appropriate settings can be made according to the actual situation.
[0048] It should be noted that this embodiment does not specifically limit the specific expression used in the hash function; it can be set according to the actual situation.
[0049] In this embodiment, the mileage coordinates and corresponding signal type information of all ground signals are first read from a preset signal location database. Then, the predicted mileage coordinates and each read mileage coordinate information are input into a fast matching model based on hash encoding. The model uses a pre-trained hash function to map the input mileage values, generate a predicted mileage hash value for the predicted mileage coordinates, and generate a corresponding mileage coordinate hash value for each mileage coordinate information. All generated hash values are fixed-length binary encoded sequences.
[0050] In practical applications, the predicted mileage coordinates are set to K152+420. The signal location database stores the mileage coordinate information of 500 ground signals. The predicted mileage coordinates are input into a fast matching model based on hash encoding. After being mapped by a hash function, a 32-bit predicted mileage hash value is obtained: 01101001110100101100101011010010. The mileage coordinate information of each ground signal is then input into the model sequentially, resulting in 500 corresponding 32-bit mileage coordinate hash values.
[0051] Step 1034: Calculate the similarity ranking by the Hamming distance between the predicted mileage hash value and each of the mileage coordinate hash values, and select the signal type information corresponding to the mileage coordinate information with the smallest Hamming distance as the preliminary matching result.
[0052] In step 1034, the Hamming distance refers to the number of bits with different values in corresponding bits between two binary codes of the same length, and the preliminary matching result refers to the type information of the ground signal that is most likely to correspond to the target area after preliminary screening through hash value similarity comparison.
[0053] In this embodiment, the predicted mileage hash value is first compared bit by bit with each mileage coordinate hash value, and the number of times the two binary values of each bit are different is counted. This number is the Hamming distance between the two hash values. Then, the ground signals corresponding to all mileage coordinate hash values are sorted in ascending order of Hamming distance. The ground signal corresponding to the mileage coordinate hash value with the smallest Hamming distance is selected, and the signal type information of the ground signal is read. This signal type information is used as the preliminary matching result.
[0054] In practical applications, let the predicted mileage hash value be 01101001110100101100101011010010. Compare it bit by bit with the mileage coordinate hash value of ground signal No. 1, which is 01001001110100101100101011010010. We find that the second bit is different, so the Hamming distance is 1. Compare it with the mileage coordinate hash value of ground signal No. 2, which is 01101001110100101100101011010000. We find that the last two bits are different, so the Hamming distance is 2. Calculate 500 Hamming distances in this way, and then select the ground signal No. 1 with the minimum Hamming distance of 1. Read its signal type information as an entry signal and use this information as the preliminary matching result.
[0055] Step 1035: Input the mileage coordinate information and signal type information of the ground signal corresponding to the preliminary matching result into the dynamic tracker based on Kalman filtering. The dynamic tracker based on Kalman filtering iteratively updates and corrects the error of the predicted mileage coordinates of the ground signal corresponding to the preliminary matching result to generate the corrected target mileage coordinates.
[0056] In step 1035, the Kalman filter-based dynamic tracker refers to a computational model that uses the Kalman filter algorithm to optimally estimate the state variables in a dynamic system; iterative update refers to the process of repeatedly correcting the previous estimate using new observation data. The termination condition for iterative update is that the change in the corrected target mileage coordinates calculated by the Kalman filter-based dynamic tracker in consecutive iterations is less than a preset convergence threshold, or the number of iterations reaches the preset maximum number of iterations. Error correction refers to reducing the estimation deviation caused by measurement noise and system noise through filtering algorithms. The corrected target mileage coordinates refer to the more accurate mileage position of the ground signal on the railway line obtained after dynamic tracking by Kalman filtering.
[0057] This application does not impose specific limitations on the model type, internal structure design, parameter design, training process, etc. of the Kalman filter-based dynamic tracker, and corresponding settings can be made according to the actual situation.
[0058] This application does not impose specific limitations on the preset convergence threshold and the number of iterations; these can be set according to actual circumstances.
[0059] In this embodiment, the mileage coordinates of the ground signal corresponding to the preliminary matching result are first used as the initial state value, and the signal type information of the ground signal is used as auxiliary reference data. Both are input into a Kalman filter-based dynamic tracker. Then, as the work vehicle moves and multiple frames of images are acquired, the dynamic tracker continuously receives new predicted mileage coordinates as observations. Iterative calculations of the prediction and update steps of Kalman filtering are performed by combining the state estimate of the previous moment and the observation of the current moment. The deviation of the mileage coordinates is gradually corrected by minimizing the estimation error covariance. After multiple iterations, the dynamic tracker outputs a corrected target mileage coordinate that is closer to the true position than the initial prediction value.
[0060] In practical applications, the mileage coordinates of the first ground signal corresponding to the initial matching result are set to K152+418, and the signal type information is set to an entry signal. As the work vehicle moves forward, five new predicted mileage coordinates are calculated for five consecutive frames of images, namely K152+420, K152+421, K152+422, K152+423, and K152+424. These observations are then input into a dynamic tracker based on Kalman filtering. After five iterations and error corrections, the corrected target mileage coordinates are generated as K152+419.5.
[0061] Step 1036: Read the corresponding signal type information from the preset signal location database according to the corrected target mileage coordinates, and use it as the target signal type corresponding to the target area.
[0062] In this embodiment, the corrected target mileage coordinates are first used as the query condition to find the ground signal record closest to the mileage coordinates in the preset signal location database; then the corresponding signal type information is read from the record and used as the target signal type corresponding to the target area for subsequent signal status identification processing.
[0063] In practical applications, let the corrected target mileage coordinates be K152+419.5. Search the signal location database for the record closest to this mileage coordinate. Find the mileage coordinate information of ground signal No. 1 as K152+418, read its signal type information as an entry signal, and take this entry signal as the target signal type corresponding to the target area.
[0064] This application obtains the current mileage value of the work vehicle by converting real-time positioning data, calculates the predicted mileage coordinates of the signal by combining image coordinates, and uses hash coding for fast matching and Kalman filtering for dynamic tracking for dual screening and correction, thereby achieving accurate positioning of the target signal type and improving the reliability of signal identification in dynamic environments.
[0065] Step 104: Input the feature vector sequence corresponding to the target region in multiple consecutive frames of images and the target signal type into the state recognition model to retrieve the corresponding signal state change rules, and compare the change pattern of the feature vector sequence with the signal state change rules to determine the color category and signal state of the signal.
[0066] The feature vector refers to the multi-dimensional numerical information extracted from the target region to characterize the attributes of the traffic light. Specifically, it consists of two parts: color features and shape features. The color features are the average hue, average saturation, and average brightness extracted after converting the target region from the RGB color space to the HSV color space. These three values represent the main hue, color purity, and luminous intensity of the traffic light, respectively. The shape features are the roundness values calculated after edge detection of the target region, which characterize the degree to which the traffic light outline closely resembles a standard circle. In the actual extraction process, firstly, image blocks of the target region are extracted from the original image based on the pixel coordinates of the target region. Then, pixel-level statistical analysis is performed on these image blocks to obtain the aforementioned color feature values. Simultaneously, edge detection and geometric calculations are performed on these image blocks to obtain shape feature values. Finally, these values are combined into a multi-dimensional vector as the feature vector corresponding to the target region.
[0067] A feature vector sequence refers to a data set formed by arranging the feature vectors corresponding to multiple consecutive frames of images in chronological order; a signal state change rule refers to parameterized information pre-stored in a signal state rule base to describe the normal operating rules of each type of signal. This rule specifically includes the signal type identifier, color appearance order, duration range of each color, and flashing frequency parameter. The color appearance order specifies the fixed order in which colors such as red, yellow, and green change sequentially. The duration range gives the shortest and longest time each color should be maintained. The flashing frequency parameter describes the periodic time of the signal light's on-off alternation; the color category refers to the specific color classification currently displayed by the signal light; and the signal state refers to the current operating mode of the signal light, including on, off, or flashing.
[0068] The specific examples of the structural design of each module of the state recognition model are as follows: The state recognition model adopts a hybrid neural network structure based on a combination of long short-term memory network and attention mechanism. Its specific composition includes five parts: input layer, feature encoding module, rule matching module, temporal analysis module and state decision module. The input layer receives feature vector sequence and target signal type information. The feature encoding module consists of two stacked long short-term memory network units, each containing 128 hidden nodes, which are used to temporally encode the feature vector sequence and extract inter-frame dependencies, and output a 128-dimensional temporal feature map. The rule matching module adopts an embedding layer structure, which maps the target signal type information into a 32-dimensional rule embedding vector, and uses this to index the corresponding color change timing rules from the externally stored signal state rule library; the timing analysis module consists of a differential calculation submodule, a threshold comparison submodule, and a jump frame filtering submodule connected in series. The differential calculation submodule uses a first-order differential algorithm, the threshold comparison submodule uses learnable adaptive threshold parameters, and the jump frame filtering submodule uses rule-based redundancy removal logic. The state decision module consists of two fully connected network layers. The first layer contains 64 neurons and is activated by a linear rectified function. The second layer contains output neurons corresponding to color category and signal state and is classified and output using a normalized exponential function. The training process of this state recognition model includes pre-collecting a large amount of historical image data of continuous changes in signal lights under different lighting and weather conditions, labeling each frame of the image with the true color category and signal state, extracting the feature vector sequence of the target area as training samples and inputting it into the model, using the cross-entropy loss function to calculate the difference between the prediction results and the true labels, and iteratively updating the weights of the long short-term memory network, the attention mechanism parameters, and the fully connected layer parameters in the model through the backpropagation algorithm until the loss function converges, thus obtaining the trained state recognition model.
[0069] It should be noted that the above structure is exemplary, and this application does not impose specific limitations on the internal structure design of the state recognition model. It can be set according to the actual situation.
[0070] In this embodiment, step 104 includes the following process, such as... Figure 2 As shown: Step 1041: Arrange the feature vectors corresponding to the target region in multiple consecutive frames of images according to the acquisition time order to form a feature vector sequence.
[0071] In step 1041, the acquisition time order refers to the order in which the images were captured.
[0072] In this embodiment, feature vectors of the target regions corresponding to multiple consecutive frames of images are first read from the memory. Then, these feature vectors are arranged sequentially according to the order in which each frame of image was acquired, forming a sequence of feature vectors arranged in chronological order. Each feature vector in this sequence corresponds to the signal light attribute information at a specific moment.
[0073] In practical applications, let the acquisition times of five consecutive frames of images be 0 seconds, 0.2 seconds, 0.4 seconds, 0.6 seconds, and 0.8 seconds, respectively. The feature vectors extracted from the target region of these five frames of images are [0.85, 0.10, 0.05], [0.88, 0.08, 0.04], [0.12, 0.83, 0.05], [0.10, 0.85, 0.05], and [0.09, 0.86, 0.05], respectively. The feature vector sequence formed by arranging these feature vectors in the order of acquisition time is {[0.85, 0.10, 0.05], [0.88, 0.08, 0.04], [0.12, 0.83, 0.05], [0.10, 0.85, 0.05], [0.09, 0.86, 0.05]}.
[0074] Step 1042: Input the feature vector sequence and the target signal type into the state recognition model, and perform time-series encoding on the feature vector sequence through the feature encoding module of the state recognition model to generate a time-series feature map. The time-series feature map is used to characterize the feature change trend between consecutive frames.
[0075] In step 1042, the feature encoding module refers to the component in the state recognition model responsible for converting the original feature vector into an encoded representation containing temporal information. The temporal feature map refers to the data structure generated after temporal encoding that can reflect the change of features over time.
[0076] In this embodiment, the feature vector sequence obtained from step 1041 and the target signal type determined from step 103 are first input into the state recognition model. Then, the feature encoding module inside the state recognition model processes the feature vector sequence. This module analyzes the change relationship of feature vectors between adjacent frames, extracts information that can characterize the trend of feature evolution over time, and organizes this information into a time-series feature map for output. This time-series feature map retains the numerical information of the original feature vectors while highlighting the change relationship between frames.
[0077] In practical applications, the feature vector sequence {[0.85, 0.10, 0.05], [0.88, 0.08, 0.04], [0.12, 0.83, 0.05], [0.10, 0.85, 0.05], [0.09, 0.86, 0.05]} and the target signal type of the station signal are input into the state recognition model. After processing by the feature encoding module, the generated time-series feature map is a 5-row, 3-column matrix. Each row in the matrix corresponds to the feature vector at a certain time. The matrix also contains the difference information between adjacent rows to represent the changing trend.
[0078] Step 1043: The rule matching module of the state recognition model retrieves the corresponding color change timing rule from the preset signal state rule library according to the target signal type.
[0079] In step 1043, the rule matching module refers to the component in the state recognition model responsible for finding the corresponding rule according to the signal type; the signal state rule base refers to a pre-built database containing the normal working rules of various types of signals. The rule base stores the signal type identifier, color appearance order, duration range of each color, and flashing frequency parameters corresponding to each record in tabular form; the color change timing rule refers to the order of appearance of each color and the duration range of each color when a specific type of signal is working normally.
[0080] In this embodiment of the application, the rule matching module of the state recognition model first receives the target signal type information determined in step 103; then, the module uses the target signal type as the query condition to search in the preset signal state rule library to find the color change timing rule record corresponding to the type; finally, the found color change timing rule is retrieved from the rule library and prepared for subsequent comparison and analysis.
[0081] In practical applications, assuming the target signal type is an entry signal, the rule matching module searches the signal status rule base with the entry signal as the keyword and retrieves the corresponding color change timing rule: red for 30 to 60 seconds, then turns yellow; yellow for 3 to 5 seconds, then turns green; and green continues until the train passes.
[0082] Step 1044: The temporal analysis module of the state recognition model performs inter-frame difference detection on the temporal feature map to identify the frame position where the color feature value in the feature vector sequence changes, and the time interval calculation unit built into the temporal analysis module generates a set of jump time intervals based on the acquisition time interval between adjacent jump frames.
[0083] Among them, the time series analysis module refers to the component in the state recognition model that is responsible for analyzing the time series feature map and detecting feature change points. The jump frame position refers to the moment corresponding to the image frame in which the color-related values in the feature vector change significantly. The time interval calculation unit refers to the functional unit inside the time series analysis module that is specifically used to calculate the time difference. The jump time interval refers to the time length value between two adjacent jump frames. The jump time interval set refers to the data set formed by arranging all jump time intervals in the order of occurrence.
[0084] Step 1044 may specifically include the following steps: A1: The difference calculation submodule of the time series analysis module performs difference calculation between adjacent frames on each feature vector arranged in time order in the time series feature map to generate an inter-frame difference value sequence.
[0085] In step A1, the difference calculation submodule refers to the functional unit in the time series analysis module that is specifically used to calculate the differences between adjacent frames. The inter-frame difference value refers to the quantified value of the degree of difference between the feature vectors of two adjacent frames. The inter-frame difference value sequence refers to the data set formed by arranging the difference values of all adjacent frames in chronological order.
[0086] In this embodiment, the difference calculation submodule first obtains the feature vectors arranged in time order in the temporal feature map; then, starting from the first feature vector, the difference operation is performed on each feature vector and the feature vector adjacent to it in turn to calculate the degree of difference between the two vectors; each difference value is recorded in the order of the corresponding inter-frame position to finally form an inter-frame difference value sequence, where each value in the sequence corresponds to the feature change amplitude between a pair of adjacent frames.
[0087] In practical applications, the temporal feature map is assumed to be a 5x3 matrix. The first row feature vector is [0.85, 0.10, 0.05], the second row feature vector is [0.88, 0.08, 0.04], the third row feature vector is [0.12, 0.83, 0.05], the fourth row feature vector is [0.10, 0.85, 0.05], and the fifth row feature vector is [0.09, 0.86, 0.05]. The difference calculation submodule calculates the difference between the first and second rows to 0.06, the difference between the second and third rows to 1.32, the difference between the third and fourth rows to 0.05, and the difference between the fourth and fifth rows to 0.03. The generated inter-frame difference value sequence is [0.06, 1.32, 0.05, 0.03].
[0088] A2: The threshold comparison submodule of the time series analysis module compares each difference value in the inter-frame difference value sequence with a preset jump detection threshold. When the difference value is greater than the jump detection threshold, the image acquisition time of the next frame corresponding to the difference value is marked as the jump frame position where the color feature value jumps, and the frame index corresponding to the jump frame position is recorded. The jump detection threshold is set according to the color change characteristics of railway signals.
[0089] In step A2, the threshold comparison submodule refers to the functional unit in the time series analysis module that is responsible for comparing the difference value with the set threshold. The jump detection threshold refers to the pre-set critical value used to determine whether the feature has changed significantly. The frame index refers to the sequential number of the image frame in the sequence.
[0090] The embodiments of this application do not specifically limit the value of the preset transition detection threshold; it can be set according to the actual situation.
[0091] In this embodiment, the threshold comparison submodule first obtains the inter-frame difference value sequence generated in step A1; then, it reads the preset jump detection threshold and compares the threshold with each difference value in the inter-frame difference value sequence in turn; for each difference value greater than the jump detection threshold, the acquisition time of the next frame image corresponding to the difference value is marked as a jump frame position, and the frame index number of the frame in the image sequence is recorded; for difference values less than or equal to the jump detection threshold, no marking is performed.
[0092] In practical applications, the preset jump detection threshold is set to 0.5, and the inter-frame difference value sequence obtained from step A1 is [0.06, 1.32, 0.05, 0.03]. The threshold comparison submodule compares the first difference value 0.06 with 0.5, and 0.06 is less than 0.5 and is not marked. The second difference value 1.32 is compared with 0.5, and 1.32 is greater than 0.5. The image acquisition time of the next frame corresponding to this difference value, i.e., the third frame, is marked as the jump frame position at 0.4 seconds, and the frame index is recorded as 3. The third difference value 0.05 is compared with 0.5, and 0.05 is less than 0.5 and is not marked. The fourth difference value 0.03 is compared with 0.5, and 0.03 is less than 0.5 and is not marked.
[0093] A3: Through the jump frame filtering submodule of the time sequence analysis module, all marked jump frame positions are redundantly removed. Combined with the upper limit of the number of color switching times of the railway signal within a fixed period, only the first jump frame position among multiple consecutive jump frame positions is retained as the valid jump frame position, resulting in a simplified jump frame position list.
[0094] In step A3, the jump frame filtering submodule refers to the functional unit in the timing analysis module responsible for de-redundancy processing of the initially marked jump frames, and the upper limit of the number of color switching times refers to the maximum number of color changes that may occur per unit time, determined according to the working principle of the signal machine. This application does not specifically limit the maximum number of color switching times; the value can be set according to the actual situation.
[0095] A valid jump frame position refers to the actual color change point retained after redundancy removal. The criteria for judgment include two aspects: First, the difference value corresponding to the jump frame position is greater than the preset jump detection threshold, indicating that a significant feature change has occurred at the position; Second, combined with the upper limit of the number of color switching times of railway signals within a fixed period, when multiple consecutive frame positions are marked as jump frame positions, only the first jump frame position is judged as a valid jump frame position, and subsequent consecutive jump frame positions are judged as false jumps caused by noise and are removed. The simplified jump frame position list refers to the data set of jump frame positions arranged in chronological order after removing redundancy.
[0096] In this embodiment, the jump frame filtering submodule first obtains all jump frame positions marked in step A2 and their corresponding frame indices; then, considering the constraint of the upper limit of the number of color switching times of the railway signal within a fixed period, the obtained jump frame positions are analyzed. When multiple consecutive frames are found to be marked as jump frame positions, it is determined that these consecutive markings are false jumps caused by noise. Only the first jump frame position is retained as a valid jump frame position, and the remaining consecutive jump frame positions are removed; finally, all the retained valid jump frame positions are rearranged in chronological order to form a simplified jump frame position list.
[0097] In practical applications, assuming that the jump frame position marked in step A2 is only the third frame corresponding to 0.4 seconds, and there are no consecutive multiple jump frame positions, the jump frame filtering submodule directly retains this position as a valid jump frame position, and the generated simplified jump frame position list only contains the jump frame position of 0.4 seconds.
[0098] A4: The time interval calculation unit of the time sequence analysis module calculates the time difference between each adjacent jump frame based on the image acquisition timestamps corresponding to two adjacent jump frame positions in the simplified jump frame position list, and uses the time difference as the jump time interval.
[0099] In step A4, the image acquisition timestamp refers to the precise moment when each frame of the image is captured, and the time difference refers to the length of time between two moments.
[0100] In this embodiment of the application, the time interval calculation unit first obtains the simplified jump frame position list obtained in step A3. The list contains multiple jump frame positions arranged in chronological order and their corresponding image acquisition timestamps. Then, starting from the first jump frame position in the list, the difference between the timestamps of two adjacent jump frame positions is calculated sequentially. The time length obtained by subtracting the timestamp of the previous jump frame from the timestamp of the next jump frame is the jump time interval between the two jump frames. Each calculated jump time interval is recorded in the corresponding jump order.
[0101] In practical applications, if the simplified jump frame position list contains only one jump frame position at 0.4 seconds, since there is only one jump frame position, it is impossible to calculate the time difference between adjacent jump frames. Therefore, this step does not generate a jump time interval in this practical application. If there are multiple jump frame positions in the list, such as 0.4 seconds and 1.2 seconds, then calculate 1.2-0.4=0.8 seconds as the jump time interval.
[0102] A5: The time interval calculation unit performs validity screening on the transition time intervals according to the standard range of the duration of each color of the railway signal, removes time intervals that exceed the standard range, and arranges the remaining transition time intervals in the order in which the transitions occur to generate a set of transition time intervals.
[0103] In step A5, the standard range refers to the time interval during which each color signal light should be used, as specified in the railway signaling technical specifications.
[0104] In this embodiment of the application, the time interval calculation unit first obtains all the transition time intervals calculated in step A4; then it reads the standard range information of the duration of each color of the railway signal, compares each transition time interval with the corresponding standard range, and determines whether the time interval falls within the standard range; the transition time intervals that fall within the standard range are retained, and the transition time intervals that exceed the standard range are removed; finally, the retained transition time intervals are arranged in the order in which the transitions occur to form a transition time interval set, which is used for subsequent state determination.
[0105] Step 1045: Through the state decision module of the state recognition model, the color transition sequence corresponding to the frame position of the transition, each time interval in the transition time interval set, and the color appearance order and duration range specified in the color change timing rule are compared item by item, and the color category and signal state of the signal machine in the current frame image are generated based on the comparison results.
[0106] In step 1045, the state decision module refers to the component in the state recognition model that is responsible for integrating various information and making a final judgment, and the color transition order refers to the order in which the colors corresponding to each transition frame position appear in time.
[0107] In this embodiment, the state decision module first obtains the color transition order corresponding to the simplified transition frame position list obtained in step A3, the transition time interval set generated in step A5, and the color change timing rule retrieved in step 1043. Then, it compares the actual detected color transition order with the color appearance order specified in the rule, and compares each time interval in the transition time interval set with the corresponding color duration range in the rule. If the actual detected color transition order is completely consistent with the order specified in the rule, and each transition time interval falls within the standard range of the corresponding color duration, the state decision module determines the color category and signal state of the signal machine in the current frame image based on the current transition frame position. If the comparison results are inconsistent, the state decision module outputs either recognition failure or maintaining the state of the previous moment.
[0108] This application constructs a feature vector sequence and inputs it into a state recognition model for time-series encoding. It then combines inter-frame difference detection and jump frame screening to accurately identify color change points. Finally, it compares the detected change patterns with the standard rules corresponding to the signal type item by item, thereby achieving accurate judgment of the signal color category and dynamic change state, effectively avoiding misjudgments that may occur in single-frame image recognition.
[0109] Step 105: Combine the coordinates of the target area, the color category, and the signal status into a recognition result, encapsulate it according to preset railway communication requirements, and send it to the driver's cab display terminal.
[0110] The identification result refers to the final output data containing signal light location and status information; the preset railway communication requirements refer to the data transmission format and protocol standards stipulated by the railway industry; and the driver's cab display terminal refers to the display device installed in the cab of the work vehicle to display information to the driver.
[0111] In this embodiment, step 105 includes the following process: Step 1051: Convert the coordinates of the target area in the image into actual position coordinates in the railway coordinate system, and combine the actual position coordinates, the color category, and the signal status into a recognition result.
[0112] In step 1051, the actual position coordinates refer to the actual geographical location of the signal on the railway line.
[0113] In this embodiment, firstly, based on the installation parameters and imaging model of the image acquisition device, the pixel coordinates of the target area in the image are converted into relative spatial coordinates with the work vehicle as the origin. Then, combined with the current GPS coordinates and driving direction of the work vehicle, the relative spatial coordinates are converted into actual position coordinates in the railway coordinate system. Then, the actual position coordinates are combined with the color category and signal status determined in step 104 to form a complete recognition result data unit.
[0114] Step 1052: Encapsulate the identification results according to the preset railway communication requirements to generate communication data frames.
[0115] In step 1052, a communication data frame refers to a data packet organized according to a specific format, which includes fields such as frame header, data length, valid data payload, and checksum.
[0116] In this embodiment, the preset railway communication protocol specification is first read to determine the format requirements of the data frame; then, the actual position coordinates, color category and signal status in the identification result are filled into the effective data payload area of the data frame in the order specified by the protocol. At the same time, the protocol identifier is filled into the frame header, the number of bytes of the effective data payload is filled into the data length field, and the check value calculated based on the entire data frame is filled into the check code field, and finally a complete communication data frame that meets the communication requirements is generated.
[0117] Step 1053: The communication data frame is sent to the driver's cab display terminal through the vehicle network inside the work vehicle. The driver's cab display terminal parses the communication data frame, marks the signal position corresponding to the actual location coordinates on the electronic map in the form of an icon, displays the color category and the signal status in a designated area on the screen, generates corresponding voice prompt information based on the signal status, and plays it through the vehicle audio equipment.
[0118] In step 1053, the vehicle-mounted network refers to the communication network inside the work vehicle used to connect various electronic devices, the electronic map refers to the digital map of the railway line stored in the display terminal, and the voice prompt information refers to the audio data that converts the signal status into voice broadcast content.
[0119] In this embodiment, the generated communication data frame is first sent to the driver's cab display terminal via the vehicle network inside the work vehicle. After receiving the communication data frame, the driver's cab display terminal parses the data frame according to the same communication protocol to extract the actual position coordinates, color category, and signal status. Then, the location point corresponding to the actual position coordinates is found on the electronic map on the display terminal screen, and the position of the signal is marked with an icon of a specific color at the location point. At the same time, the color category and signal status are displayed in text form in the designated information display area on the screen. Finally, the corresponding voice prompt information is retrieved from the pre-stored voice library according to the signal status and broadcast through the vehicle audio equipment.
[0120] In this embodiment, after step 105, the following process is also included: B1: Obtain the current operating status parameters and the real-time distance between the work vehicle and the forward signal. The operating status parameters include real-time vehicle speed, train pipe pressure, and driver's controller handle position. The real-time distance is calculated based on the coordinates of the target area and the real-time positioning data of the work vehicle.
[0121] In step B1, the operating status parameters refer to various monitoring data that reflect the current operating status of the work vehicle, the real-time vehicle speed refers to the current driving speed of the work vehicle, the train pipe pressure refers to the gas pressure value in the braking system, the driver's control handle position refers to the gear position of the driver's control handle, and the real-time distance refers to the straight-line spatial length between the current position of the work vehicle and the position of the forward signal.
[0122] In this embodiment of the application, the current real-time vehicle speed, train pipe pressure, and driver controller handle position are first read from the vehicle's on-board monitoring system and used as operating status parameters. At the same time, based on the actual position coordinates of the signal determined in step 1051 and the current GPS coordinates of the work vehicle, the spatial straight-line distance between the two points is calculated as the real-time distance between the work vehicle and the signal ahead.
[0123] B2: Input the signal state and the operating state parameters from the identification results into the fuzzy logic-based control rule base, and perform fuzzification processing and fuzzy reasoning on the current signal state and operating state parameters through the fuzzy logic-based control rule base to generate the current security level label.
[0124] In step B2, the fuzzy logic-based control rule base refers to a pre-constructed set of fuzzy reasoning rules related to the safe operation of railway vehicles. Each rule consists of a condition part and a conclusion part. The condition part includes a fuzzy set of signal status, a fuzzy set of real-time vehicle speed, and a fuzzy set of real-time distance. The conclusion part is a fuzzy set of safety level. For example, if the signal status is prohibited, the real-time vehicle speed is high, and the real-time distance is short, the safety level is dangerous. If the signal status is permitted, the real-time vehicle speed is low, and the real-time distance is far, the safety level is safe. These rules are formulated according to railway traffic safety regulations and stored in the rule base for fuzzy reasoning.
[0125] Safety level labels are classification marks used to characterize the current level of driving safety.
[0126] In this embodiment, the signal state from the identification result obtained in step 105 and the operating state parameters obtained in step B1 are first input into a fuzzy logic-based control rule base. The rule base first performs fuzzification processing on the input signal state and operating state parameters, converting the real-time vehicle speed into fuzzy concepts such as fast and slow, the train pipe pressure into fuzzy concepts such as high and low, the driver's controller handle position into fuzzy concepts such as traction and braking, and the signal state into fuzzy concepts such as permitted and prohibited. Then, based on pre-stored fuzzy inference rules, such as if the signal state is prohibited, the real-time vehicle speed is fast, and the real-time distance is close, the safety level is dangerous, fuzzy logic deduction is performed. Finally, the inference result is defuzzified to obtain a specific safety level label such as dangerous, warning, or safe.
[0127] B3: The identification results, operating status parameters, real-time distance, and safety level labels at multiple consecutive moments are arranged in chronological order to form operating status time-series data. The operating status time-series data is then input into a long short-term memory network. Based on the historical change patterns in the operating status time-series data, the long short-term memory network predicts the relative position change trend between the work vehicle and the forward signal and the probability of signal status changes at the next moment.
[0128] In step B3, the running status time sequence data refers to the data set formed by arranging the running-related information at multiple times in chronological order, the relative position change trend refers to the direction of change of the distance between the work vehicle and the signal over time, and the probability of signal status change refers to the probability that the signal will change color at a future time.
[0129] In this embodiment, the identification results, operating status parameters, real-time distance, and safety level labels for multiple consecutive moments are first obtained from historical records. These data are then arranged in chronological order to form a time-series data of operating status. This time-series data of operating status is then input into a pre-trained long short-term memory network. The long short-term memory network analyzes the changing patterns in the historical data through its internal memory units and gating structures, extracting features such as the speed change trend of the work vehicle, the distance change trend, and the signal status change pattern. Based on these features, it predicts whether the relative position change trend between the work vehicle and the signal ahead will be closer or farther away, and whether the probability of a change in the signal status is high or low.
[0130] B4: When the Long Short-Term Memory Network predicts that a signal state switch will occur within a preset time threshold or that the work vehicle may cross the signal, a warning command is triggered.
[0131] In step B4, the preset time threshold refers to the pre-set time length used to determine whether an alarm is needed, and the alarm command refers to the control signal used to trigger the alarm function.
[0132] The embodiments of this application do not specifically limit the value of the preset time threshold; it can be set according to the actual situation.
[0133] In this embodiment of the application, a preset time threshold, such as 10 seconds, is first read; then the prediction result obtained from step B3 is judged. If the long short-term memory network predicts that the signal state will switch within the next 10 seconds, for example, from green to yellow, or predicts that the work vehicle will cross the signal position within the next 10 seconds, then an early warning instruction is immediately triggered. The instruction contains early warning type and early warning level information.
[0134] B5: Generate secondary warning information containing the warning level according to the warning instruction, encapsulate the secondary warning information according to the preset railway communication requirements and send it to the driver's cab display terminal for prominent display, and at the same time play the corresponding warning voice through the on-board audio equipment.
[0135] In step B5, secondary warning information refers to supplementary alarm information generated based on the prediction results, and highlighting means emphasizing the display on the screen using a special color or flashing method.
[0136] In this embodiment, a secondary warning message is first generated based on the warning type and warning level contained in the warning instruction triggered from step B4. This message includes warning content such as a change in the signal ahead and a warning level such as medium. Then, the secondary warning message is encapsulated according to the same preset railway communication requirements as step 1052 to generate a warning data frame. The warning data frame is then sent to the driver's cab display terminal via the vehicle network. After parsing, the display terminal highlights the warning message on the screen with a flashing red icon and enlarged font. At the same time, the corresponding warning voice is retrieved from the voice library according to the warning level and played through the vehicle audio equipment to remind the driver to take appropriate action in a timely manner.
[0137] This application combines the identification results with the operating status parameters to perform fuzzy logic reasoning and timing prediction, and issues early warnings when the signal status may change or there is a risk of signal transgression, thus realizing the functions of proactive early warning and auxiliary decision-making for driving safety.
[0138] Figure 3 This application provides a schematic diagram of the structure of a railway ground signal image feature analysis and recognition system, as shown in the embodiment of the present application. Figure 3 As shown, the system includes: The acquisition module 31 is used to acquire visible light image data and thermal infrared image data of the ground signal, and to obtain the mileage coordinate information and signal type information of the ground signal along the railway line.
[0139] The determination module 32 is used to determine the target area of the traffic light in the image based on the visible light image data and the thermal infrared image data.
[0140] The calculation module 33 is used to calculate the predicted mileage coordinates of the ground signal corresponding to the target area along the railway line based on the real-time positioning data of the work vehicle, match the predicted mileage coordinates with the mileage coordinate information, and combine the signal type information to determine the target signal type corresponding to the target area.
[0141] The input module 34 is used to input the feature vector sequence corresponding to the target region in multiple consecutive frames of images and the target signal type into the state recognition model, so as to retrieve the corresponding signal state change rules, and compare the change pattern of the feature vector sequence with the signal state change rules to determine the color category and signal state of the signal.
[0142] The combination module 35 is used to combine the coordinates of the target area, the color category, and the signal status into a recognition result, and then encapsulate it according to preset railway communication requirements and send it to the driver's cab display terminal.
[0143] The railway ground signal image feature analysis and recognition system of this application is used to implement the aforementioned railway ground signal image feature analysis and recognition method. Therefore, the specific implementation of the railway ground signal image feature analysis and recognition system can be found in the embodiment section of the railway ground signal image feature analysis and recognition method above. The specific implementation can be referred to the description of the corresponding embodiments, and will not be repeated here.
[0144] This application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described railway ground signal image feature analysis and recognition methods.
[0145] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the above-described railway ground signal image feature analysis and recognition methods.
[0146] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory, random access memory, portable hard drives, magnetic disks, or optical disks.
[0147] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the railway ground signal image feature analysis and recognition method.
[0148] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0149] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0150] The foregoing has provided a detailed description of a railway ground signal image feature analysis and recognition method and system provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A railway ground signal image feature analysis and recognition method, characterized in that, include: Collect visible light image data and thermal infrared image data of ground signals, and obtain the mileage coordinate information and signal type information of ground signals along the railway line; Based on the visible light image data and the thermal infrared image data, the target area of the traffic light in the image is determined; Based on the real-time positioning data of the work vehicle, the predicted mileage coordinates of the ground signal corresponding to the target area along the railway line are calculated. The predicted mileage coordinates are matched with the mileage coordinate information, and combined with the signal type information, the target signal type corresponding to the target area is determined. The feature vector sequence corresponding to the target region in multiple consecutive frames of images and the target signal type are input into the state recognition model to retrieve the corresponding signal state change rules. The change pattern of the feature vector sequence is compared with the signal state change rules to determine the color category and signal state of the signal. The coordinates of the target area, the color category, and the signal status are combined into a recognition result, which is then packaged according to preset railway communication requirements and sent to the driver's cab display terminal. The step of inputting the feature vector sequence corresponding to the target region in multiple consecutive frames of images and the target signal type into the state recognition model to retrieve the corresponding signal state change rules, and comparing the change pattern of the feature vector sequence with the signal state change rules to determine the color category and signal state of the signal, includes: The feature vectors corresponding to the target region in multiple consecutive frames of images are arranged in the order of acquisition time to form a feature vector sequence; The feature vector sequence and the target signal type are input into the state recognition model. The feature encoding module of the state recognition model performs time-series encoding on the feature vector sequence to generate a time-series feature map. The time-series feature map is used to characterize the feature change trend between consecutive frames. The rule matching module of the state recognition model retrieves the corresponding color change timing rule from the preset signal state rule library according to the target signal type. The temporal analysis module of the state recognition model performs inter-frame difference detection on the temporal feature map to identify the frame position where the color feature value in the feature vector sequence changes. The time interval calculation unit built into the temporal analysis module generates a set of jump time intervals based on the acquisition time interval between adjacent jump frames. The state judgment module of the state recognition model compares the color transition order corresponding to the frame position of the transition, each time interval in the transition time interval set, with the color appearance order and duration range specified in the color change timing rule, and generates the color category and signal state of the signal in the current frame image based on the comparison results.
2. The method of claim 1, wherein, The time-series analysis module of the state recognition model performs inter-frame difference detection on the time-series feature map to identify the frame positions where color feature values in the feature vector sequence change. The time-series analysis module's built-in time-series calculation unit generates a set of jump time intervals based on the acquisition time intervals between adjacent jump frames, including: The difference calculation submodule of the time series analysis module performs difference calculation between adjacent frames on each feature vector arranged in time order in the time series feature map to generate a sequence of inter-frame difference values. The threshold comparison submodule of the time series analysis module compares each difference value in the inter-frame difference value sequence with a preset jump detection threshold. When the difference value is greater than the jump detection threshold, the image acquisition time of the next frame corresponding to the difference value is marked as the jump frame position where the color feature value jumps, and the frame index corresponding to the jump frame position is recorded. The jump detection threshold is set according to the color change characteristics of railway signal lights. The timing analysis module uses the jump frame filtering submodule to remove redundancy from all marked jump frame positions. Combined with the upper limit of the number of color switching times of the railway signal within a fixed period, only the first jump frame position is retained as the valid jump frame position from multiple consecutive jump frame positions, resulting in a simplified jump frame position list. The time interval calculation unit of the time series analysis module calculates the time difference between each adjacent jump frame based on the image acquisition timestamps corresponding to two adjacent jump frame positions in the simplified jump frame position list, and uses the time difference as the jump time interval. The time interval calculation unit performs validity screening on the transition time intervals according to the standard range of the duration of each color of the railway signal, eliminates time intervals that exceed the standard range, and arranges the remaining transition time intervals in the order in which the transitions occur to generate a set of transition time intervals.
3. The method of claim 1, wherein, The step of calculating the predicted mileage coordinates of the ground signal along the railway line corresponding to the target area based on the real-time positioning data of the work vehicle, matching the predicted mileage coordinates with the mileage coordinate information, and combining the signal type information to determine the target signal type corresponding to the target area includes: The current GPS coordinates of the work vehicle are extracted from the real-time positioning data, and the GPS coordinates are converted into the railway mileage value of the work vehicle based on the preset railway line geographic information table. Based on the pixel coordinates of the target area in the image, combined with the internal parameters and external installation parameters of the image acquisition device, the distance and azimuth of the ground signal corresponding to the target area relative to the work vehicle are calculated. Based on the distance, the azimuth, and the railway mileage value where the work vehicle is currently located, the predicted mileage coordinates of the ground signal corresponding to the target area along the railway line are calculated. The mileage coordinates and signal type information of all ground signals are read from the preset signal location database. The predicted mileage coordinates and each of the mileage coordinate information are input into a fast matching model based on hash coding. The fast matching model based on hash coding performs hash function mapping on the predicted mileage coordinates and each of the mileage coordinate information respectively to generate the predicted mileage hash value and the hash value of each mileage coordinate. By calculating the Hamming distance between the predicted mileage hash value and each of the mileage coordinate hash values, the similarity is sorted, and the signal type information corresponding to the mileage coordinate information with the smallest Hamming distance is selected as the preliminary matching result. The mileage coordinates and signal type information of the ground signal corresponding to the preliminary matching result are input into the dynamic tracker based on Kalman filtering. The dynamic tracker based on Kalman filtering iteratively updates and corrects the error of the predicted mileage coordinates of the ground signal corresponding to the preliminary matching result, and generates the corrected target mileage coordinates. Based on the corrected target mileage coordinates, the corresponding signal type information is read from the preset signal location database and used as the target signal type corresponding to the target area.
4. The railway ground signal image feature analysis and recognition method according to claim 1, characterized in that, After being sent to the driver's cab display terminal, it also includes: The current operating status parameters and the real-time distance between the work vehicle and the forward signal are obtained. The operating status parameters include real-time vehicle speed, train pipe pressure, and driver's controller handle position. The real-time distance is calculated based on the coordinates of the target area and the real-time positioning data of the work vehicle. The signal state and the operating state parameters in the identification results are input into the control rule base based on fuzzy logic. The current signal state and operating state parameters are fuzzified and fuzzy inferred by the control rule base based on fuzzy logic to generate the security level label at the current moment. The identification results, operating status parameters, real-time distance, and safety level labels at multiple consecutive moments are arranged in chronological order to form operating status time series data. The operating status time series data is then input into a long short-term memory network. Based on the historical change patterns in the operating status time series data, the long short-term memory network predicts the relative position change trend between the work vehicle and the signal ahead, as well as the possibility of signal status changes at the next moment. When the long short-term memory network predicts that a signal state switch will occur within a preset time threshold or that the work vehicle may cross the signal, a warning command is triggered. The system generates secondary warning information containing the warning level according to the warning command, encapsulates the secondary warning information according to the preset railway communication requirements, and sends it to the driver's cab display terminal for prominent display. At the same time, the corresponding warning voice is played through the onboard audio equipment.
5. The railway ground signal image feature analysis and recognition method according to claim 1, characterized in that, The step of combining the coordinates of the target area, the color category, and the signal status into a recognition result, and then encapsulating it according to preset railway communication requirements before sending it to the driver's cab display terminal includes: The coordinates of the target area in the image are converted into actual position coordinates in the railway coordinate system, and the actual position coordinates, the color category, and the signal status are combined into a recognition result. The identification results are encapsulated according to the preset railway communication requirements to generate communication data frames; The communication data frame is sent to the driver's cab display terminal via the vehicle's internal network. The driver's cab display terminal parses the communication data frame, marks the signal position corresponding to the actual location coordinates on the electronic map in the form of an icon, displays the color category and the signal status in a designated area on the screen, generates corresponding voice prompt information based on the signal status, and plays it through the vehicle's audio equipment.
6. The railway ground signal image feature analysis and recognition method according to claim 1, characterized in that, Determining the target area of the traffic light in the image based on the visible light image data and the thermal infrared image data includes: Based on the pre-stored joint calibration parameters, the visible light image data and the thermal infrared image data are registered in spatial coordinate system to obtain the registered multimodal image data; A region segmentation algorithm is used to analyze the thermal infrared image data in the registered multimodal image data to extract candidate signal regions, and the coordinates of the candidate signal regions are mapped onto the visible light image data in the registered multimodal image data. The candidate signal regions are verified in the visible light image data. Interference regions are eliminated based on the verification results, and the image regions corresponding to the remaining candidate signal regions are taken as the target regions of the traffic lights in the image.
7. A railway ground signal image feature analysis and recognition system, characterized in that, include: The acquisition module is used to acquire visible light image data and thermal infrared image data of ground signals, and to obtain the mileage coordinate information and signal type information of the ground signals along the railway line; The determination module is used to determine the target area of the traffic light in the image based on the visible light image data and the thermal infrared image data; The calculation module is used to calculate the predicted mileage coordinates of the ground signal corresponding to the target area along the railway line based on the real-time positioning data of the work vehicle, match the predicted mileage coordinates with the mileage coordinate information, and combine the signal type information to determine the target signal type corresponding to the target area. The input module is used to input the feature vector sequence corresponding to the target region in multiple consecutive frames of images and the target signal type into the state recognition model, so as to retrieve the corresponding signal state change rules, and compare the change pattern of the feature vector sequence with the signal state change rules to determine the color category and signal state of the signal. The combination module is used to combine the coordinates of the target area, the color category, and the signal status into a recognition result, and then encapsulate it according to preset railway communication requirements and send it to the driver's cab display terminal. The step of inputting the feature vector sequence corresponding to the target region in multiple consecutive frames of images and the target signal type into the state recognition model to retrieve the corresponding signal state change rules, and comparing the change pattern of the feature vector sequence with the signal state change rules to determine the color category and signal state of the signal, includes: The feature vectors corresponding to the target region in multiple consecutive frames of images are arranged in the order of acquisition time to form a feature vector sequence; The feature vector sequence and the target signal type are input into the state recognition model. The feature encoding module of the state recognition model performs time-series encoding on the feature vector sequence to generate a time-series feature map. The time-series feature map is used to characterize the feature change trend between consecutive frames. The rule matching module of the state recognition model retrieves the corresponding color change timing rule from the preset signal state rule library according to the target signal type. The temporal analysis module of the state recognition model performs inter-frame difference detection on the temporal feature map to identify the frame position where the color feature value in the feature vector sequence changes. The time interval calculation unit built into the temporal analysis module generates a set of jump time intervals based on the acquisition time interval between adjacent jump frames. The state judgment module of the state recognition model compares the color transition order corresponding to the frame position of the transition, each time interval in the transition time interval set, with the color appearance order and duration range specified in the color change timing rule, and generates the color category and signal state of the signal in the current frame image based on the comparison results.
8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor is configured to execute the computer program to implement the steps of the railway ground signal image feature analysis and recognition method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, enables the implementation of the railway ground signal image feature analysis and recognition method as described in any one of claims 1 to 6.