Rail transit management system screen signal recognition and response method and system

By performing high-frequency image acquisition and parallel processing on the display screen of the rail transit management system, combining the scale-invariant feature transformation algorithm and multi-threaded parallel color analysis, scenario simulation and strategy construction are carried out, which solves the problem that manual monitoring is difficult to quickly and accurately identify complex traffic conditions, and achieves efficient and safe rail transit management.

CN119169515BActive Publication Date: 2025-05-06GUANGZHOU SINORAIL INFORMATION ENG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411191524.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2025-05-06
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

The existing rail transit management system relies on manual monitoring, making it difficult to quickly and accurately identify complex traffic conditions and emergencies, resulting in delays in decision-making or errors in judgment, affecting safety and operational efficiency.

Method used

By performing high-frequency image acquisition and parallel processing on multiple display screens in the control room, feature extraction is performed using the scale-invariant feature transformation algorithm, combining multi-threaded parallel color analysis and state matching, scenario simulation and strategy construction are carried out, and natural language interactive text and response control instructions are generated.

Benefits of technology

It improves the real-time and accuracy of signal recognition, enhances the ability to identify abnormal situations, improves the forward-looking decision-making and responds to emergencies, reduces the work burden of dispatchers, improves the accuracy and timeliness of instruction execution, and significantly improves the efficiency and safety of rail transit management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119169515B_ABST
    Figure CN119169515B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and discloses a method and system for identifying and responding to screen signals in a rail transit management system, which are used to improve the efficiency and accuracy of identifying and responding to screen signals in a rail transit management system. The method includes: performing high-frequency image acquisition on multiple display screens at the same time, obtaining multiple original image data streams, and performing parallel preprocessing and adaptive screen area positioning to obtain multiple standardized screen images; performing feature extraction on multiple standardized screen images to obtain a high-dimensional screen signal feature data set, and performing multi-threaded parallel color analysis and state matching to obtain multiple signal state information; performing scenario simulation on the multiple signal state information to obtain simulated scenario data and perform strategy construction to obtain a target response strategy; generating a natural language interactive text according to the target response strategy, generating a response control instruction and a response prompt information, and transmitting the response control instruction and the response prompt information to a data interaction terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a method and system for identifying and responding to screen signals in a rail transit management system. Background Art

[0002] Rail transit management systems are the core of ensuring safe and efficient operation of urban rail transit. Traditional rail transit management systems mainly rely on manual monitoring and operation, identifying signal status and making dispatching decisions by observing multiple display screens in the control room. This approach can still cope with simple daily operations, but it is often difficult to respond quickly and accurately when faced with complex traffic conditions or emergencies.

[0003] However, with the continuous expansion of the urban rail transit network and the continuous increase in passenger flow, manual monitoring alone can no longer meet the needs of modern rail transit management. Manual identification and processing of large amounts of real-time information are prone to fatigue and errors, especially during peak hours or under abnormal circumstances, which may lead to decision delays or misjudgments, affecting the safety and operational efficiency of rail transit. Summary of the invention

[0004] In view of this, an embodiment of the present invention provides a method and system for identifying and responding to screen signals in a rail transit management system, which are used to improve the efficiency and accuracy of identifying and responding to screen signals in a rail transit management system.

[0005] The present invention provides a screen signal recognition and response method for a rail transit management system, comprising: simultaneously performing high-frequency image acquisition on multiple display screens in a rail transit management system control room to obtain multiple original image data streams, and performing parallel preprocessing and adaptive screen area positioning on the multiple original image data streams to obtain multiple standardized screen images; performing feature extraction on the multiple standardized screen images through a scale-invariant feature transformation algorithm to obtain a high-dimensional screen signal feature data set; performing multi-threaded parallel color analysis and state matching on the high-dimensional screen signal feature data set to obtain multiple signal state information; performing scenario simulation on the multiple signal state information to obtain simulated scenario data, and performing signal response strategy construction on the simulated scenario data to obtain a target response strategy; generating a natural language interactive text according to the target response strategy, generating a response control instruction and a response prompt information according to the natural language interactive text, and transmitting the response control instruction and the response prompt information to a preset data interaction terminal.

[0006] In the present invention, the steps of simultaneously performing high-frequency image acquisition on multiple display screens in the control room of the rail transit management system to obtain multiple original image data streams, and performing parallel preprocessing and adaptive screen area positioning on the multiple original image data streams to obtain multiple standardized screen image steps include: performing multi-channel video stream acquisition on the multiple display screens to obtain multiple original image data streams, and performing frame decomposition processing on the multiple original image data streams to obtain multiple groups of continuous image frame sequences; adjusting and screening the sampling frequency of the multiple groups of continuous image frame sequences to obtain an optimized image frame set, and performing multi-spectral fusion processing on the optimized image frame set to obtain Composite image data, and adaptively adjusting the exposure and white balance of the composite image data to obtain a light-balanced image; constructing a multi-scale image pyramid on the light-balanced image to obtain a multi-resolution image set, and performing parallel edge detection on the multi-resolution image set to obtain a plurality of screen contour candidate areas; screening the plurality of screen contour candidate areas by a deep learning target detection algorithm to obtain a screen position information set, and performing perspective transformation and geometric correction on the screen position information set to obtain a plurality of corrected screen images; and performing local contrast enhancement and noise suppression on the plurality of corrected screen images to obtain the plurality of standardized screen images.

[0007] In the present invention, the step of extracting features from the multiple standardized screen images through a scale-invariant feature transformation algorithm to obtain a high-dimensional screen signal feature data set includes: constructing a TS-SIFT Gaussian pyramid for the multiple standardized screen images to obtain a multi-scale image group of a rail transit screen, and performing Gaussian difference analysis on the multi-scale image group of a rail transit screen to obtain a differential image group; performing local extreme point detection on the differential image group to obtain an initial feature point candidate set, and performing edge response screening on the initial feature point candidate set to obtain a signal feature point candidate area; performing small-scale signal localization on the signal feature point candidate area; The invention relates to a method for detecting local extreme values ​​to obtain an initial track signal feature point set, and performing sub-pixel signal positioning on the initial track signal feature point set to obtain position track signal feature points; performing traffic signal main direction allocation on the position track signal feature points to obtain rotation invariant track signal feature points; performing local signal gradient information extraction on the rotation invariant track signal feature points to obtain a rail transit signal feature descriptor, and performing data enhancement on the rail transit signal feature descriptor to obtain a color sensitive track signal feature vector; performing traffic signal feature dimensionality reduction processing on the color sensitive track signal feature vector to obtain the high dimensional screen signal feature data set.

[0008] In the present invention, the step of performing multi-threaded parallel color analysis and state matching on the high-dimensional screen signal feature data set to obtain multiple signal state information includes: performing multi-threaded allocation on the high-dimensional screen signal feature data set to obtain a parallel processing task group; performing HSV color space conversion on each feature data in the parallel processing task group to obtain an HSV feature representation data set; defining a dynamic color range for the HSV feature representation data set to obtain a signal color dynamic threshold, and generating signal color judgment standard data based on the signal color dynamic threshold; performing feature vectorization processing on the signal color judgment standard data to obtain a color feature vector set, and performing similarity calculation on the color feature vector set and a preset signal state template to obtain a signal state candidate set; performing signal weight analysis on the signal state candidate set based on the pre-acquired historical state information to obtain a signal state weight set; performing state matching on the high-dimensional screen signal feature data set based on the signal state weight set to obtain the multiple signal state information.

[0009] In the present invention, the step of performing state matching on the high-dimensional screen signal feature data set based on the signal state weight set to obtain the multi-signal state information includes: performing time series analysis on the high-dimensional screen signal feature data set based on the signal state weight set to obtain a signal state change trend; performing frequency analysis on the signal state change trend to obtain a signal change periodic feature and a non-periodic signal, and performing pattern matching on the signal change periodic feature to obtain a periodic signal pattern; classifying and clustering the periodic signal pattern and the non-periodic signal to obtain a signal type grouping; performing spatiotemporal correlation analysis on the signal type grouping to obtain a logical relationship between signals, and constructing a signal state transition diagram based on the logical relationship between signals to obtain a complex signal combination pattern; performing flickering signal recognition and gradient color signal recognition on the complex signal combination pattern to obtain a target signal state; performing text content analysis on the target signal state to obtain text-color association information; and performing weighted fusion on the text-color association information based on the signal state weight set to obtain multi-signal state information.

[0010] In the present invention, the scenario simulation of the multi-signal status information is performed to obtain simulation scenario data, and the signal response strategy is constructed for the simulation scenario data to obtain the target response strategy step, including: performing line topology mapping on the multi-signal status information to obtain a track network status diagram, and performing signal linkage relationship analysis on the track network status diagram to obtain a signal interdependence list; generating train operation conflict detection rules based on the signal interdependence list to obtain a safety constraint condition set; performing statistical analysis on the pre-collected historical operation data to obtain a typical traffic flow pattern, and performing virtual train operation based on the typical traffic flow pattern and the multi-signal status information. The method comprises the following steps: constructing a scenario to obtain an initial simulation scenario; iterating the initial simulation scenario for multiple times to simulate different signal switching sequences to obtain multiple groups of simulation results, and performing safety assessment on the multiple groups of simulation results to screen out simulation scenarios that meet the safety constraint condition set to obtain a valid simulation scenario set; calculating the operating efficiency of the valid simulation scenario set to obtain an efficiency score table, and constructing a multi-level scheduling strategy based on the efficiency score table to obtain a preliminary response strategy library; performing simulation verification on the preliminary response strategy library to obtain a strategy impact assessment report, and performing strategy screening on the preliminary response strategy library according to the strategy impact assessment report to obtain the target response strategy.

[0011] In the present invention, the step of generating a natural language interactive text according to the target response strategy, generating a response control instruction and a response prompt information according to the natural language interactive text, and transmitting the response control instruction and the response prompt information to a preset data interaction terminal includes: classifying and analyzing the target response strategy to obtain a signal switching sequence, a train dispatching instruction, and a platform management suggestion; performing time-series expansion on the signal switching sequence to obtain a signal light change schedule; generating a train adjustment plan based on the train dispatching instruction to obtain a train operation diagram correction suggestion, and converting the platform management suggestion into passenger flow diversion measures to obtain a platform operation guide; performing time-series expansion on the signal light change schedule, the train operation diagram correction suggestion, and the platform operation guide The guide is used to integrate data to obtain a comprehensive dispatch plan; based on the comprehensive dispatch plan, a dispatcher's operation step list is generated and command conversion is performed to obtain a human-computer interaction instruction set, and the human-computer interaction instruction set is hierarchically processed and divided into emergency instructions, routine instructions and advisory instructions to obtain a hierarchical instruction library; the hierarchical instruction library is converted into a console operation sequence to obtain a response control instruction; based on the response control instruction, an operation prompt text conversion is performed to obtain response prompt information, and the response control instruction and the response prompt information are timestamped and prioritized to obtain an ordered data packet; data encapsulation is performed on the ordered data packet to obtain a system-compatible data stream, and the system-compatible data stream is transmitted to the data interaction terminal.

[0012] The present invention also provides a rail transit management system screen signal recognition and response system, comprising:

[0013] An acquisition module is used to simultaneously perform high-frequency image acquisition on multiple display screens in a rail transit management system control room to obtain multiple original image data streams, and perform parallel preprocessing and adaptive screen area positioning on the multiple original image data streams to obtain multiple standardized screen images;

[0014] An extraction module, used for performing feature extraction on the plurality of standardized screen images by using a scale-invariant feature transformation algorithm to obtain a high-dimensional screen signal feature data set;

[0015] A matching module, used for performing multi-threaded parallel color analysis and state matching on the high-dimensional screen signal feature data set to obtain multi-signal state information;

[0016] A simulation module, used to perform scenario simulation on the multi-signal state information to obtain simulation scenario data, and to construct a signal response strategy for the simulation scenario data to obtain a target response strategy;

[0017] A transmission module is used to generate a natural language interaction text according to the target response strategy, generate a response control instruction and a response prompt information according to the natural language interaction text, and transmit the response control instruction and the response prompt information to a preset data interaction terminal.

[0018] In the technical solution provided by the present invention, by performing high-frequency image acquisition and parallel processing on multiple display screens in the control room, the real-time and accuracy of signal recognition are greatly improved, and the problem of fatigue and error in manual monitoring is effectively solved; the scale-invariant feature transformation algorithm is used for feature extraction, combined with multi-threaded parallel color analysis and state matching, it is possible to quickly and accurately identify various complex signal states, including flashing signals and gradient color signals, greatly improving the ability to identify abnormal situations; by performing scenario simulation and strategy construction on multi-signal state information, it is possible to predict potential changes in traffic conditions and generate corresponding response strategies, which not only improves the foresight of decision-making, but also enhances the ability to respond to emergencies. In addition, natural language interactive texts, response control instructions and prompt information are automatically generated based on the target response strategy, which greatly reduces the workload of dispatchers and improves the accuracy and timeliness of instruction execution; the automation and intelligence of the entire process significantly improves the efficiency and safety of rail transit management, reduces human errors, and improves the efficiency and accuracy of screen signal recognition and response in the rail transit management system. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 This is a flow chart of a method for identifying and responding to screen signals in a rail transit management system in an embodiment of the present invention;

[0021] Figure 2 It is a schematic diagram of images collected in real time from a rail transit management system in an embodiment of the present invention.

[0022] Figure 3 It is a schematic diagram of a screen signal recognition and response system for a rail transit management system in an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.

[0025] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0026] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 , Figure 1 is a flow chart of a method for identifying and responding to screen signals in a rail transit management system according to an embodiment of the present invention. Figure 1 As shown, the following steps are included:

[0027] S101, simultaneously performing high-frequency image acquisition on multiple display screens in a rail transit management system control room to obtain multiple original image data streams, and performing parallel preprocessing and adaptive screen area positioning on the multiple original image data streams to obtain multiple standardized screen images;

[0028] S102, extracting features from a plurality of standardized screen images by using a scale-invariant feature transformation algorithm to obtain a high-dimensional screen signal feature data set;

[0029] S103, performing multi-threaded parallel color analysis and state matching on the high-dimensional screen signal feature data set to obtain multi-signal state information;

[0030] S104, performing scenario simulation on the multi-signal state information to obtain simulated scenario data, and constructing a signal response strategy for the simulated scenario data to obtain a target response strategy;

[0031] S105: Generate a natural language interaction text according to the target response strategy, generate a response control instruction and a response prompt information according to the natural language interaction text, and transmit the response control instruction and the response prompt information to a preset data interaction terminal.

[0032] Specifically, by installing high-resolution cameras in the control room, real-time images of multiple screens can be captured simultaneously. These cameras are usually configured to capture more than 30 frames per second, such as Figure 2 As shown in the figure, it is a real-time image collected by the camera, which can indicate the identification area of ​​the positioning. The identification information of each signal in the figure is displayed in real time at the corresponding position. The collected image data is then converted into multiple original image data streams, each of which corresponds to a real-time image sequence of a screen. These original image data streams are preprocessed in parallel and adaptively positioned in the screen area. Parallel preprocessing includes operations such as image denoising, contrast enhancement, and color correction, which can improve the accuracy of subsequent analysis. Adaptive screen area positioning uses an improved Canny edge detection algorithm and Hough transform to accurately locate the boundaries of each screen. The key to this step is to adapt to different lighting conditions and screen layouts to ensure accurate extraction of the screen area. After these processes, the original image is converted into multiple standardized screen images, laying the foundation for subsequent analysis.

[0033] The scale-invariant feature transform (SIFT) algorithm is applied to these standardized screen images for feature extraction. The advantage of the SIFT algorithm is that it can extract local features that are not affected by image scaling, rotation, and brightness changes. In this system, the SIFT algorithm is optimized to TS-SIFT (TrafficScreen-SIFT), which is specifically adapted to the characteristics of rail transit screens. The TS-SIFT algorithm first constructs the scale space of the image, then detects local extreme points at different scales, and finally generates feature descriptors. These feature descriptors constitute a high-dimensional screen signal feature dataset, which contains key information of various signal elements on the screen. Multi-threaded parallel color analysis and state matching are performed on the high-dimensional screen signal feature dataset. The color analysis uses a dynamic color perception system (DCPS), which can adaptively adjust the color judgment standard to adapt to different lighting conditions. The DCPS first converts the image to the HSV color space, and then defines a dynamic threshold based on a preset color range. State matching uses a machine learning model, such as a support vector machine (SVM), to match the extracted features with a pre-trained signal state template. Through this step, the system can identify various signal states on the screen, including signal light color, text information, etc., thereby obtaining multi-signal state information.

[0034] After obtaining the multi-signal status information, scenario simulation is performed. Using digital twin technology, a virtual rail transit operation scenario is constructed based on the currently identified signal status, combined with historical data and preset rules. During the simulation process, the system will consider a variety of possible situations, such as normal operation, peak hours, equipment failure, etc., and generate a series of simulated scenario data. These data contain information such as traffic flow, train operation status, and passenger distribution under different circumstances. Based on the simulated scenario data, the signal response strategy is constructed. This process uses reinforcement learning algorithms, such as deep Q network (DQN), to generate the optimal response strategy through repeated training and optimization. The strategy construction process considers multiple goals such as safety, efficiency, and energy consumption, and dynamically adjusts the strategy according to the priority of different situations. The final target response strategy is a complex structure with multiple decision points that can cope with various possible situations. Convert the target response strategy into natural language interactive text. By using natural language processing (NLP) technology, the professional terms and complex logic in the strategy are converted into easy-to-understand text descriptions. Based on these texts, specific response control instructions and response prompt information are generated. Control instructions include signal switching commands, train dispatching instructions, etc., while prompt information provides decision-making basis and operation suggestions for operators. This information is eventually transmitted to the pre-set data interaction terminal for use by dispatchers and other relevant personnel.

[0035] For example, during the peak hours of a certain rail transit line, a station platform screen showing crowded conditions was captured through high-frequency image acquisition. The TS-SIFT algorithm extracts a set of feature points from the image, including information such as passenger density and waiting area status. Color analysis identifies the red warning signal on the screen, indicating that the platform crowding exceeds 90%. The system then conducts a scenario simulation and predicts that if no measures are taken, the platform crowding will reach 110% in the next 30 minutes, posing a safety hazard. Based on this, the reinforcement learning algorithm generates a multi-step response strategy, including: increasing the frequency of up trains from 6 minutes / trip to 4 minutes / trip, temporarily closing some platform entrances, and starting the station broadcasting system to guide passengers.

[0036] This strategy is converted into natural language text: "Emergency: Platform A is too crowded. The following measures are recommended: 1. Increase the frequency of up trains and shorten the interval from 6 minutes to 4 minutes. 2. Temporarily close entrances 2 and 4 and guide passengers to enter the station from other entrances. 3. Start the station broadcast and ask passengers to disperse and avoid gathering. It is expected that after implementing these measures, the platform crowding will drop below 85% within 30 minutes." At the same time, corresponding control instructions are generated, such as "Adjustment of up train scheduling interval: 240 seconds", "Platform entrance control: Close gates 2 and 4", etc., and this information is transmitted to the data interaction terminal of the dispatching center.

[0037] By executing the above steps, high-frequency image acquisition and parallel processing are performed on multiple display screens in the control room, which greatly improves the real-time and accuracy of signal recognition, and effectively solves the problem of fatigue and error in manual monitoring. The scale-invariant feature transformation algorithm is used for feature extraction, combined with multi-threaded parallel color analysis and state matching, which can quickly and accurately identify various complex signal states, including flashing signals and gradient color signals, greatly improving the ability to identify abnormal situations. By performing scenario simulation and strategy construction on multi-signal state information, potential changes in traffic conditions can be predicted and corresponding response strategies can be generated, which not only improves the foresight of decision-making, but also enhances the ability to respond to emergencies. In addition, natural language interactive texts, response control instructions and prompt information are automatically generated based on the target response strategy, which greatly reduces the workload of dispatchers and improves the accuracy and timeliness of instruction execution. The automation and intelligence of the entire process significantly improves the efficiency and safety of rail transit management, reduces human errors, and improves the efficiency and accuracy of screen signal recognition and response in the rail transit management system.

[0038] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0039] (1) collecting multiple video streams from multiple display screens to obtain multiple original image data streams, and performing frame decomposition processing on the multiple original image data streams to obtain multiple groups of continuous image frame sequences;

[0040] (2) adjusting and screening the sampling frequencies of multiple groups of continuous image frame sequences to obtain an optimized image frame set, performing multi-spectral fusion processing on the optimized image frame set to obtain composite image data, and performing adaptive exposure and white balance adjustment on the composite image data to obtain a light-balanced image;

[0041] (3) constructing a multi-scale image pyramid for the illumination balanced image to obtain a multi-resolution image set, and performing parallel edge detection on the multi-resolution image set to obtain multiple screen contour candidate regions;

[0042] (4) screening multiple screen contour candidate areas through a deep learning target detection algorithm to obtain a screen position information set, and performing perspective transformation and geometric correction on the screen position information set to obtain multiple corrected screen images;

[0043] (5) Performing local contrast enhancement and noise suppression on the multiple corrected screen images to obtain multiple standardized screen images.

[0044] Specifically, multiple video streams are collected for multiple display screens in the control room. By installing high-definition cameras in the control room, each camera is responsible for capturing real-time images of one or more screens. The collected video streams usually have a high frame rate, such as 60 frames per second, to ensure that the rapidly changing information on the screen can be captured. These video streams constitute multiple original image data streams. These original image data streams are subjected to frame decomposition processing. Frame decomposition is the process of dividing a continuous video stream into discrete image frames. For example, a 60-frame-per-second video stream will be decomposed into 60 independent images within 1 second. In this way, the video stream of each screen is converted into a set of continuous image frame sequences. The sampling frequency of the obtained multiple sets of continuous image frame sequences is adjusted and filtered. The sampling frequency adjustment aims to balance processing efficiency and information integrity. For example, the original 60-frame-per-second video may be downsampled to 30 frames per second to reduce the amount of data while maintaining sufficient temporal resolution. The screening process removes blurred, overexposed or underexposed frames and retains only clear and information-rich frames. The optimized image frame set obtained in this step reduces the amount of data while ensuring image quality.

[0045] The optimized image frame set is subjected to multispectral fusion processing. Multispectral fusion is a technology that combines image information from different bands. In this solution, in addition to visible light images, near-infrared images are also collected. Fusion algorithms, such as wavelet transform-based fusion methods, integrate information from these different bands to generate composite image data with richer details. This step is particularly helpful in improving image quality under different lighting conditions. Adaptive exposure and white balance adjustments are performed on the composite image data. The adaptive exposure algorithm dynamically adjusts the exposure parameters by analyzing the brightness histogram of the image to ensure that the image is neither overexposed nor underexposed. White balance adjustment ensures the accuracy of image color by identifying neutral tones in the image (such as white or gray areas), calculating and correcting color deviations. The illumination-balanced images obtained after these processes can maintain consistent visual effects at different times and under different lighting conditions.

[0046] A multi-scale image pyramid is constructed for the illumination-balanced image. An image pyramid is a multi-resolution image representation method that is obtained by downsampling the original image level by level. For example, for a 1920x1080 original image, the first level of the pyramid may be 960x540, the second level is 480x270, and so on. This multi-scale representation helps with subsequent edge detection and object recognition.

[0047] Then, edge detection is performed on the multi-resolution image set in parallel. Edge detection is the process of identifying areas in an image where the brightness changes sharply. The commonly used Canny edge detection algorithm first performs a Gaussian filter on the image to reduce noise, then calculates the gradient amplitude and direction of the image, and obtains the edge through non-maximum suppression and double threshold processing. In this scheme, edge detection is performed in parallel at multiple resolution levels to improve efficiency and accuracy. The result of this step is multiple screen contour candidate regions. These candidate regions are screened by a deep learning target detection algorithm. This scheme uses an improved version of the YOLO (You Only Look Once) algorithm, which divides the image into grids, and predicts a bounding box and a category probability for each grid. By setting a confidence threshold (such as 0.8), high-probability screen areas are screened out to obtain an accurate set of screen location information.

[0048] The screened screen position information set is subjected to perspective transformation and geometric correction. Perspective transformation uses the mapping relationship between four corresponding point pairs to convert the tilted screen image into a front view. Geometric correction further corrects the distortion of the image to ensure the accurate presentation of the screen content. This step obtains multiple corrected screen images. Local contrast enhancement and noise suppression are performed on the corrected screen image. Local contrast enhancement uses adaptive histogram equalization technology to divide the image into multiple small areas (such as 8x8 grids), perform histogram equalization on each area separately, and then use bilinear interpolation to merge the processing results. Noise suppression uses a non-local mean denoising algorithm, which reduces noise by searching for similar image blocks in the entire image and performing weighted averaging on these blocks.

[0049] For example, in a rail transit control center, the monitoring system includes five high-definition display screens, each of which is captured by a camera with a 4K resolution (3840x2160 pixels) and 60 frames per second.

[0050] Multi-channel video stream acquisition: The amount of raw data captured by each camera per second: 3840×2160 pixels×3 bytes / pixel (RGB)×60 frames=1,492,992,000 bytes≈1.39GB The total data volume of 5 screens: 1.39GB×5=6.95GB / second

[0051] Frame decomposition processing: Decompose the 60-frame / second video stream into independent image frames to obtain 300 frames / second (5 screens × 60 frames / second).

[0052] Sampling frequency adjustment and screening: downsampling to 30 frames / second, data volume halved: 6.95GB / 2 = 3.475GB / second Assuming 10% of poor quality frames are screened out, the remaining data volume is: 3.475GB×0.9 = 3.1275GB / second

[0053] Multispectral fusion processing: Assuming the fusion of near-infrared images (single channel), the data volume increases by 33%: 3.1275GB×1.33=4.15957GB / second

[0054] Adaptive exposure and white balance adjustment: This step does not significantly change the amount of data, but improves image quality. For example, for an overexposed image area, the original RGB value is (255,255,255), which may become (220,225,230) after adjustment.

[0055] Multi-scale image pyramid construction: Build a 5-level pyramid with the resolution halved at each level:

[0056] Levela: 3840x2160 (original size);

[0057] Level b: 1920x1080;

[0058] Levelc:960x540;

[0059] Leveled: 480x270;

[0060] Level: 240x135;

[0061] The total data volume increased by about 33%: 4.15957GB × 1.33 = 5.53222GB / second

[0062] Parallel edge detection: Use the Canny edge detection algorithm. For a 3840x2160 image:

[0063] Gaussian filter: Using a 5x5 kernel, it requires about 3840×2160×25=207,360,000 floating point operations

[0064] Gradient calculation: 3840×2160×2=16,588,800 operations are required, non-maximum suppression and threshold processing: about 3840×2160×3=24,883,200 operations, a total of about 250 million operations / frame; Deep learning target detection: using YOLOv5s model, input size is 640x640: Scale 3840x2160 image to 640x640;

[0065] YOLOv5s has about 7.2 million parameters and requires about 1.44 billion floating point operations per forward pass;

[0066] Perspective transformation and geometric correction: For a 3840x2160 image, the perspective transformation requires: 3840×2160×9=74,649,600 floating point operations (9 multiplications and additions per pixel);

[0067] Local contrast enhancement and noise suppression: Adaptive histogram equalization using an 8×8 grid:

[0068] Divide the 3840x2160 image into 480×270 8×8 blocks;

[0069] Each block requires 256 accumulation operations to build the histogram;

[0070] A total of about 480×270×256=33,177,600 operations are required;

[0071] Non-local means denoising, assuming the search window is 21×21 and the block size is 7×7:

[0072] Each pixel needs to be compared (2121)(7×7)=21,609 times;

[0073] A total of 3840 × 2160 × 21,609 = 178,582,118,400 operations are required;

[0074] Through this series of processing, the original 6.95GB / second video data is converted into a set of high-quality, information-rich standardized screen images. Although the amount of data has increased, the information density and usability have been significantly improved. The processed images are clearer, less noisy, and have better contrast.

[0075] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0076] (1) Constructing a TS-SIFT Gaussian pyramid for multiple standardized screen images to obtain a multi-scale image group of rail transit screens, and performing Gaussian difference analysis on the multi-scale image group of rail transit screens to obtain a differential image group;

[0077] (2) Performing local extreme point detection on the differential image group to obtain an initial feature point candidate set, and performing edge response screening on the initial feature point candidate set to obtain a signal feature point candidate region;

[0078] (3) Performing small-scale signal local extrema detection on the candidate areas of signal feature points to obtain an initial track signal feature point set, and performing sub-pixel signal positioning on the initial track signal feature point set to obtain position track signal feature points;

[0079] (4) Allocating the main direction of the traffic signal to the position track signal feature points to obtain the rotationally invariant track signal feature points;

[0080] (5) Extracting local signal gradient information from the rotationally invariant rail signal feature points to obtain a rail transit signal feature descriptor, and performing data enhancement on the rail transit signal feature descriptor to obtain a color-sensitive rail signal feature vector;

[0081] (6) Perform traffic signal feature dimensionality reduction processing on the color-sensitive track signal feature vector to obtain a high-dimensional screen signal feature dataset.

[0082] Specifically, a TS-SIFT Gaussian pyramid is constructed for each standardized screen image. A series of images of different scales are generated by continuously Gaussian blurring and downsampling the original image. For example, for an original image of 1920×1080 pixels, five scale layers may be generated, with the resolution of each layer being 1920×1080, 960×540, 480×270, 240×135, and 120×68, respectively. Such a multi-scale representation enables the algorithm to detect features of different sizes. Gaussian difference analysis is performed on these multi-scale image groups. Gaussian difference is obtained by subtracting images of adjacent scales. Specifically, if G(x,y,K σ ) indicates the scale is K σ The Gaussian blurred image is then the Gaussian difference D(x,y,σ) can be expressed as: D(x,y,σ)=G(x,y,K σ )-G(x,y,σ). It enhances the features such as edges and corners in the image. It detects local extreme points on the differential image group. It finds the local maximum or minimum points by comparing the values ​​of each pixel with its 26 neighboring points (8 at the same scale and 9 at the upper and lower adjacent scales). These points constitute the initial feature point candidate set.

[0083] The initial feature point candidate set is screened for edge response to remove unstable edge response points. Here, the eigenvalue ratio of the Hessian matrix is ​​used for judgment. If the ratio of the maximum eigenvalue to the minimum eigenvalue exceeds a certain threshold (usually set to 10), the point is considered to be on the edge and is removed. The remaining points constitute the candidate region of the signal feature point. Small-scale signal local extreme value detection is performed in the candidate region of the signal feature point. In order to locate the feature points more accurately, a more accurate extreme value position is found by interpolating in a small range (such as a 3×3×3 three-dimensional space). The initial track signal feature point set obtained in this way has higher position accuracy. Sub-pixel signal positioning is performed on the initial track signal feature point set. This is achieved by Taylor expansion of the differential function D(x, y, σ) at the feature point position and then finding the extreme value point. The position of the feature point can be accurately determined to the decimal pixel level, which greatly improves the accuracy of subsequent matching.

[0084] The main direction of the traffic signal is assigned to the precisely located track signal feature points. This is achieved by calculating the gradient direction histogram of the area around the feature point. Usually a 16x16 window is selected around the feature point, the gradient amplitude and direction of each pixel is calculated, and then weighted with Gaussian weights. The direction corresponding to the main peak of the histogram is the main direction of the feature point. The rotationally invariant track signal feature points obtained after this processing have rotational invariance.

[0085] Extract local signal gradient information from rotationally invariant rail signal feature points. In the 16×16 area around the feature point, take the main direction of the feature point as the reference, calculate the 8-direction gradient histogram of the 4×4 sub-areas, and finally obtain a 128-dimensional feature vector. This vector is the rail transit signal feature descriptor.

[0086] In order to enhance the recognition ability of features, data enhancement is performed on the rail transit signal feature descriptor. Color information is introduced here to expand the original grayscale features into features containing color information. For example, the gradient can be calculated in the RGB or HSV color space to obtain a 384-dimensional (128x3) color-sensitive rail signal feature vector.

[0087] Finally, the color-sensitive track signal feature vector is subjected to traffic signal feature dimensionality reduction. This can be achieved through methods such as principal component analysis (PCA) or linear discriminant analysis (LDA). The purpose of dimensionality reduction is to reduce the amount of calculation and remove redundant information while retaining the most recognizable features. The high-dimensional screen signal feature dataset obtained after dimensionality reduction not only retains the key information of the original features, but also greatly improves the efficiency of subsequent processing.

[0088] For example, taking a standardized screen image of 1920x1080 pixels as an example, the processing of the TS-SIFT algorithm is described in detail:

[0089] TS-SIFT Gaussian pyramid construction: create 5 scale layers, initial σ = 1.6,

[0090] Layer 0: 1920x1080, σ=1.6;

[0091] Layer 1: 960x540,

[0092] Layer 2: 480x270, σ=1.6×2;

[0093] Layer 3: 240x135,

[0094] Layer 4: 120×68, σ=1.6×4.

[0095] Each layer of images is blurred using a (6σ+1)×(6σ+1) Gaussian kernel. For example, the 0th layer uses a 11×11 Gaussian kernel.

[0096] Gaussian difference analysis: Calculate the difference between adjacent scale images to obtain 4 sets of difference images.

[0097] For example, the difference between layer 0 and layer 1:

[0098] For the pixel (100,100), assume that: Then D(100,100,σ)=140-150=-10;

[0099] Local extreme point detection: Compare pixel values ​​in a 3x3x3 cube. Assume that at position (100,100), the difference value -10 is smaller than the surrounding 26 points, then it is a local minimum point.

[0100] Edge Response Screening:

[0101] Calculate the Hessian matrix: H = [D xx D xy ;D xy D yy ] Assume that at position (100,100): D xx =D(101,100)+D(99,100)-2D(100,100)=5D yy =D(100,101)+D(100,99)-2D(100,100)=4D xy =(D(101,101)-D(99,101)-D(101,99)+D(99,99)) / 4=1;

[0102] Eigenvalue ratio r=(D xx +D yy ) 2 / ((D xx +D yy -D xy ) 2 =81 / 19≈4.26, if r<10, keep this point.

[0103] Small-scale signal local extreme value detection: Quadratic function fitting is performed in the 3x3x3 space. Assuming that the obtained extreme point offset is (Δx, Δy, Δσ) = (0.3, -0.2, 0.1), the precise feature point position is (100.3, 99.8), and the scale is 1.6×2 (0+0.1) ≈1.72;

[0104] Main direction assignment:

[0105] The gradient is calculated in a 16x16 window around the feature point. For each pixel (x,y): (x,y)=tan -1 ((L(x,y+1)-L(x,y-1)) / (L(x+1,y)-L(x-1,y)). Create a 36-bin direction histogram, each bin represents 10°. Assuming the maximum bin value appears in the 80°-90° interval and the amplitude is 0.8, the main direction is 85°.

[0106] Extraction of local signal gradient information: Based on the main direction, the 8-directional gradient histogram of 4x4 sub-regions is calculated in a 16x16 window. Each sub-region has 4×4 pixels and 8 directions, resulting in a 4×4×8=128-dimensional vector. For example, the value of the first direction of the first sub-region is: sum(m(x,y)×w(x,y)), where w is the Gaussian weight and m is the gradient amplitude, assuming that it is calculated to be 0.5. Data enhancement: Expand to RGB space to obtain a 384-dimensional vector. For example, the first value 0.5 of the original 128-dimensional vector may become: R: 0.7, G: 0.4, B: 0.4 in RGB space;

[0107] Feature dimensionality reduction: PCA is used for dimensionality reduction. First, the covariance matrix of the 384-dimensional data is calculated, and then its eigenvalues ​​and eigenvectors are calculated. Assuming that the sum of the first 64 eigenvalues ​​accounts for 95% of the total, these 64 eigenvectors are retained. The final 64-dimensional eigenvector is obtained by taking the inner product of the original 384-dimensional vector and these 64 eigenvectors.

[0108] Through this process, the original 1920x1080 pixel image is converted into a set of 64-dimensional feature vectors, each of which represents a significant feature point in the image. For example, for an important traffic light in the image, its final 64-dimensional feature vector may be: [0.82, -0.15, 0.33, ..., 0.05]

[0109] This vector contains both the position and shape information of the traffic light (reflected in the larger values ​​of the first few dimensions) and the color information (possibly reflected in specific dimensions related to color). This compact and information-rich representation greatly improves the efficiency and accuracy of subsequent signal state recognition.

[0110] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0111] (1) Multi-thread allocation is performed on the high-dimensional screen signal feature data set to obtain a parallel processing task group;

[0112] (2) Perform HSV color space conversion on each feature data in the parallel processing task group to obtain an HSV feature representation data set;

[0113] (3) Dynamically define the color range of the HSV feature representation data set to obtain the signal color dynamic threshold, and generate signal color judgment standard data based on the signal color dynamic threshold;

[0114] (4) performing feature vectorization processing on the signal color judgment standard data to obtain a color feature vector set, and performing similarity calculation between the color feature vector set and a preset signal state template to obtain a signal state candidate set;

[0115] (5) performing signal weight analysis on the signal state candidate set based on the previously acquired historical state information to obtain a signal state weight set;

[0116] (6) Based on the signal state weight set, the state of the high-dimensional screen signal feature data set is matched to obtain multi-signal state information.

[0117] Specifically, the high-dimensional screen signal feature dataset is distributed to multiple threads to fully utilize the multi-core processing capabilities of modern computers. Assume that there is a dataset containing 10,000 feature vectors, each with 64 dimensions, which can be evenly distributed to 8 processing threads, with each thread processing 1,250 feature vectors.

[0118] Next, perform HSV color space conversion on each feature data in the parallel processing task group. The HSV (Hue, Saturation, Value) color space is closer to the way humans perceive color than RGB, and is particularly suitable for color analysis. The conversion formula is as follows:

[0119] V=max(R,G,B); S=(V-min(R,G,B)) / H={60×(GB) / (V-min(R,G,B)), if V=R120+60×(BR) / (V-min(R,G,B)), if V=G240+60×(RG) / (V-min(R,G,B)), ifV=B};

[0120] Where: R, G, B represent the values ​​of the red, green, and blue channels respectively, and the range is [0, 1]. V represents brightness, which is the maximum value in RGB. S represents saturation, which is calculated as the difference between the maximum and minimum values ​​divided by the maximum value. H represents hue, and different calculation formulas are selected depending on which channel has the largest value.

[0121] For example, for an RGB value of (0.8, 0.2, 0.1), the converted HSV value is approximately (5.7°, 0.875, 0.8).

[0122] Then, a dynamic color range is defined for the HSV feature representation dataset. This step takes into account the changes in color under different lighting conditions and defines a dynamic range for each signal color by analyzing a large amount of sample data. For example, the HSV range of a red signal light may be defined as: H: [0°, 10°] ∪ [350°, 360°], S: [0.7, 1.0], V: [0.6, 1.0]. This range is not fixed, but dynamically adjusted according to the current lighting conditions.

[0123] Based on the dynamic color range, the signal color judgment standard data is generated. This standard data contains the HSV range of each signal color and the corresponding confidence function. The confidence function f(h, s, v) can be defined as a three-dimensional Gaussian function:

[0124] f(h,s,v)=exp(-((h-μ h ) 2 / 2σ h 2 +(s-μ s ) 2 / 2σ s 2 +(v-μ v ) 2 / 2σ v 2 ))

[0125] Among them: h, s, v represent the HSV values ​​to be judged. h , μ s , μ v Respectively represent the standard HSV value of the color. h , σ s , σ v They represent the allowable range of variation of the three dimensions H, S, and V, namely the standard deviation.

[0126] For example, for the red signal, there may be μ h =5°,μ h =0.9,μ v =0.8,σ h =5°,σ s =0.1,σ v =0.1. Perform feature vectorization on the signal color judgment standard data. This step converts each color judgment standard into a vector to facilitate the subsequent similarity calculation. For example, the upper and lower bounds of the HSV range and the parameters of the confidence function can be combined into a vector: [Hmin, Hmax, Smin, Smax, Vmin, Vmax, μ h , μ s , μ v , σ h , σ s , σ v ].

[0127] Then, the color feature vector set is similar to the preset signal state template. The preset signal state template is a set of standard color feature vectors representing different signal states. The similarity calculation can use cosine similarity:

[0128] cos(θ)=(A·B) / (||A||||B||);

[0129] Where: A and B are two vectors to be compared. A·B represents the dot product of two vectors. ||A|| and ||B|| represent the Euclidean norm of two vectors. The signal weight analysis is performed on the candidate signal state set based on the historical state information obtained in advance. This step takes into account the time continuity and change law of the signal state. The weight analysis can use the Markov chain model to calculate the state transition probability matrix P, P through historical data. ij represents the probability of transitioning from state i to state j.

[0130] Finally, the state matching of the high-dimensional screen signal feature data set is performed based on the signal state weight set to obtain multi-signal state information. This step uses Bayesian inference to calculate the posterior probability of each possible state:

[0131] P(S|F)∝P(F|S)×P(S);

[0132] Where: S represents a specific signal state. F represents the observed feature. P(S|F) is the posterior probability that the signal state is S given the observed feature F. P(F|S) is the likelihood probability, which is obtained from the color similarity. P(S) is the prior probability, which is provided by the state transition probability.

[0133] The state with the highest posterior probability is selected as the final recognition result.

[0134] For example: Suppose that in a busy rail transit control center, there is a signal light that needs to be monitored in real time. After the previous processing, a 64-dimensional feature vector V is obtained. First, V is assigned to a processing thread. The thread first performs HSV conversion to obtain the HSV value (5°, 0.85, 0.75). Then, according to the current dynamic color range definition, it is determined that this HSV value falls into the range of the red signal. Next, the color confidence is calculated and 0.92 is obtained. This result is calculated with the preset red signal template for similarity, and the similarity is 0.95. Then, the historical state information is checked and it is found that the state sequence of this signal in the past 3 minutes is [red, red, red, yellow, red]. The state transition probability calculated using the Markov chain model shows that the probability of the next state being red is 0.8 and the probability of being yellow is 0.2. Finally, the posterior probability is calculated using Bayesian inference: P(red|F)∝0.95×0.8=0.76, P(yellow|F)∝0.05×0.2=0.01. Therefore, the final recognition result is a red signal with a confidence level of 0.76 / (0.76+0.01)≈0.987. This result is then added to the multi-signal status information.

[0135] In a specific embodiment, the process of performing the state matching step on the high-dimensional screen signal feature data set based on the signal state weight set may specifically include the following steps:

[0136] (1) Perform time series analysis on the high-dimensional screen signal feature data set based on the signal state weight set to obtain the signal state change trend;

[0137] (2) Performing frequency analysis on the signal state change trend to obtain the signal change periodic characteristics and non-periodic signals, and performing pattern matching on the signal change periodic characteristics to obtain the periodic signal pattern;

[0138] (3) Classify and cluster periodic signal patterns and non-periodic signals to obtain signal type groups;

[0139] (4) Performing spatiotemporal correlation analysis on signal type groups to obtain the logical relationship between signals, and constructing a signal state transition diagram based on the logical relationship between signals to obtain a complex signal combination pattern;

[0140] (5) Perform flashing signal recognition and gradient color signal recognition on complex signal combination patterns to obtain the target signal state;

[0141] (6) Analyze the text content of the target signal state to obtain text-color association information;

[0142] (7) Based on the signal state weight set, the text-color association information is weightedly fused to obtain multi-signal state information.

[0143] Specifically, the historical data of the signal state is modeled using time series analysis methods, such as the autoregressive integrated moving average model (ARIMA). The ARIMA model is defined by three parameters (p, d, q), where p is the number of autoregressive terms, d is the number of differences, and q is the number of moving average terms. By analyzing the deviation between the model's prediction results and the actual data, the trend of signal state changes can be obtained.

[0144] Next, the frequency analysis of the signal state change trend is performed using the Fast Fourier Transform (FFT) algorithm. FFT converts the time domain signal into a frequency domain representation, and the calculation formula is X(k) = Σ[n = 0 to N-1] x(n)e -j2πkn / N , where x(n) is the time domain signal, X(k) is the frequency domain signal, and N is the number of sampling points. By analyzing the spectrum, the main periodic components of the signal change can be identified and the periodic characteristics of the signal change can be obtained. For signals that cannot find obvious peaks in the spectrum, they are classified as non-periodic signals.

[0145] The dynamic time warping (DTW) algorithm is used to perform pattern matching on the periodic features of signal changes. DTW calculates the similarity of two time series. Its core is to find the best alignment path between the two sequences so that the total distance between corresponding points on the path is minimized.

[0146] The distance calculation formula of DTW is as follows:

[0147] DTW(i,j)=d(x i ,y j )+min{DTW(i-1,j),DTW(i,j-1),DTW(i-1,j-1)}, where

[0148] d(x i ,y j ) is the distance between two points. Through the DTW algorithm, similar periodic features can be classified to obtain periodic signal patterns.

[0149] K-means++ algorithm is used to classify and cluster periodic signal patterns and non-periodic signals. K-means++ is an improved version of K-means, and its initial center point selection is more reasonable, avoiding local optimal solutions. The objective function of the algorithm is to minimize the sum of the squares of the distances from all points to their nearest center point, that is, minΣ[i=1to k]Σ[x∈C i ]||x-μ i || 2 , where k is the number of clusters, 2 is the i-th cluster, μ i is the center point of the ith cluster. Through iterative optimization, the signal type grouping is finally obtained.

[0150] The signal types are grouped for spatiotemporal correlation analysis using spatial autocorrelation analysis methods, such as the Moran's I index. The calculation formula for Moran's I is as follows:

[0151] I=(N / W)×(Σ{i,j}W ij (x i -x)(x j -x)) / Σ{i}(x i -x) 2 ;

[0152] Where N is the number of spatial units, W is the spatial weight matrix and W ij is the weight between spatial units i and j, x i and x jis the attribute value, and x is the average value. By analyzing the spatial correlation between different signals, the logical relationship between signals is obtained. Based on the logical relationship between signals, a signal state transition diagram is constructed and represented by a directed graph. The nodes in the graph represent the signal states, and the edges represent the transition relationship between states. The weight of the edge can be represented by the transition probability, which is calculated by analyzing historical data. Using graph traversal algorithms, such as depth-first search (DFS) or breadth-first search (BFS), complex signal combination patterns can be identified.

[0153] Flickering signal recognition and gradient color signal recognition are performed on complex signal combination patterns. Flickering signal recognition uses the time window method to set a fixed-size time window (such as 2 seconds) and count the number of changes in the signal state within the window. If the number of changes exceeds the preset threshold (such as 4 times), it is determined to be a flickering signal. Gradient color signal recognition uses color gradient analysis to calculate the color difference between adjacent time points: Where L, a, and b are the coordinates of the CIELAB color space. If the ΔE value changes continuously within a certain range, it is judged as a gradient color signal.

[0154] The text content of the target signal state is analyzed using optical character recognition (OCR) technology. The OCR process includes steps such as image preprocessing, character segmentation, feature extraction, and character recognition. Convolutional neural network (CNN) is used for character recognition, and the network structure includes convolution layer, pooling layer, and fully connected layer. By training a large number of samples, CNN can learn the feature representation of characters and achieve high-accuracy text recognition.

[0155] Finally, based on the signal state weight set, the text-color association information is weighted fused. Using the weighted average method, the fusion formula is S = Σ(wi×si) / Σwi, where wi is the weight of the i-th information source and si is the corresponding state value. The weight can be dynamically adjusted according to the reliability of each information source. Through this weighted fusion, the final multi-signal state information is obtained.

[0156] For example, in a rail transit control center, there is a group of signal lights that need to be monitored in real time. First, the signal status data of the past 24 hours is modeled using ARIMA to obtain the model parameters (2,1,1) and predict the state change trend in the next hour. Then, FFT analysis is performed on these data, and it is found that the main period is 300 seconds and the amplitude is 0.8, which corresponds to a standard signal cycle. The DTW algorithm is used to compare the periodic characteristics of different signals, and the similarity threshold is set to 0.9. Signals with similarity higher than the threshold are classified into one category, and three typical periodic signal patterns are obtained.

[0157] These signal patterns and other non-periodic signals were clustered using the K-means++ algorithm, with k set to 5. After 10 iterations, five signal type groups were obtained. Next, the Moran's I index between these five groups was calculated to obtain a 5x5 correlation matrix, in which the highest correlation coefficient was 0.85, indicating that there was a strong spatial correlation between the two signal groups. Based on these correlations, a signal state transition graph with 15 nodes and 40 edges was constructed.

[0158] During the analysis, it was found that one signal changed state 6 times in 2 seconds with a frequency of 3Hz, which was determined to be a flickering signal. The color of another signal gradually changed from green (L=50, a=-40, b=30) to yellow (L=70, a=10, b=40) in 10 seconds, and the ΔELAB value increased from 0 to 36.6, which was determined to be a gradient color signal. The text next to the signal was OCR recognized using a 5-layer CNN model with a recognition accuracy of 98%, and the text content was obtained as "Speed ​​limit 30km / h".

[0159] Finally, all the obtained information is weighted fused. Assume that the weight of color information is 0.5, the weight of text information is 0.3, and the weight of flashing / gradient features is 0.2. For example, for the gradient yellow signal, the color state value is 0.8 (yellow), the text state value is 0.7 (speed limit), and the gradient feature state value is 0.9. The fused state value S = (0.50.8 + 0.30.7 + 0.2 × 0.9) / (0.5 + 0.3 + 0.2) = 0.79. This result indicates that the signal is in a high warning state and requires attention. These processed multi-signal state information provides accurate real-time conditions for the rail transit management system.

[0160] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0161] (1) Perform line topology mapping on multiple signal status information to obtain a track network status diagram, and perform signal linkage relationship analysis on the track network status diagram to obtain a signal interdependency list;

[0162] (2) Generate train operation conflict detection rules based on the signal interdependence list to obtain a set of safety constraint conditions;

[0163] (3) Statistically analyze the historical operation data collected in advance to obtain typical traffic flow patterns, and construct virtual train operation scenarios based on the typical traffic flow patterns and multiple signal status information to obtain initial simulation scenarios;

[0164] (4) The initial simulation scenario is iterated multiple times to simulate different signal switching sequences, obtain multiple sets of simulation results, and perform safety assessment on the multiple sets of simulation results to screen out simulation scenarios that meet the safety constraint condition set and obtain a valid simulation scenario set;

[0165] (5) Calculate the operating efficiency of the effective simulation scenario set to obtain an efficiency score table, and build a multi-level scheduling strategy based on the efficiency score table to obtain a preliminary response strategy library;

[0166] (6) Conduct simulation verification on the preliminary response strategy library to obtain a strategy impact assessment report, and screen the strategies in the preliminary response strategy library based on the strategy impact assessment report to obtain a target response strategy.

[0167] Specifically, using the adjacency matrix representation in graph theory, each signal point is regarded as a node of the graph, and the connection relationship between signal points is regarded as an edge. Assuming that there are n signal points, an n×n matrix A is formed, where Aij=1 indicates that there is a direct connection between signal points i and j, and Aij=0 indicates that there is no connection. In this way, the multi-signal state information is converted into a track network state graph. Next, the signal linkage relationship analysis is performed on the track network state graph. This step uses the depth-first search (DFS) algorithm to explore all possible paths from each node. During the search process, the dependencies between nodes are recorded, including direct dependencies and indirect dependencies. For example, if the state change of signal A causes the state change of signal B, then B is considered to be dependent on A. The signal interdependency list obtained in this way is an n×n matrix D, where Dij represents the degree of dependency of signal i on signal j, and the value range is [0,1].

[0168] Based on the signal interdependency list, train operation conflict detection rules are generated. This process uses the decision tree algorithm to train a conflict detection model with signal state combinations as features and whether a conflict occurs as a label. Each internal node of the decision tree represents a judgment condition for a signal state, and the leaf node represents the conclusion of conflict or no conflict. Through pruning optimization, a set of simplified rules is obtained to form a set of safety constraints. Statistical analysis is performed on the pre-collected historical operation data, and time series clustering methods such as the dynamic time warping (DTW) algorithm are used to cluster similar traffic flow patterns. The DTW algorithm calculates the similarity of two time series. Its core is to find an alignment method that minimizes the sum of the distances between corresponding points in the two sequences. Through clustering, several typical traffic flow patterns are obtained, each of which contains characteristics such as average flow and peak hours.

[0169] Based on typical traffic flow patterns and multi-signal status information, a virtual train operation scenario is constructed using discrete event simulation technology. In the simulation, each train acts as an independent entity and runs according to a preset schedule and route. The signal status acts as a constraint to affect the train operation. In this way, the initial simulation scenario is obtained.

[0170] Run the initial simulation scenario multiple times, randomly changing the switching sequence of some signals in each iteration. Use the Monte Carlo method to perform a large number of simulations (e.g., 10,000 times), and record indicators such as train operation status and delay time for each simulation. Perform a safety assessment on these multiple sets of simulation results and use the previously generated safety constraint condition set for screening. The assessment criteria include the minimum spacing between trains, the maximum waiting time on the platform, etc. The simulation scenarios that meet all safety constraints are retained to form a valid simulation scenario set.

[0171] The operational efficiency is calculated for the valid simulation scenario set. Efficiency indicators include average delay time, total passenger transportation volume, energy consumption, etc. A multi-objective optimization method, such as Pareto optimization, is used to calculate the comprehensive efficiency score for each scenario. Pareto optimization considers multiple objectives and finds the solution set where improvement in any one objective will lead to deterioration of other objectives. The resulting efficiency score table reflects the advantages and disadvantages of different scheduling strategies.

[0172] Construct a multi-level scheduling strategy based on the efficiency score table. Use a reinforcement learning algorithm, such as Q-learning, with the rail network state as the environment, the scheduling decision as the action, and the efficiency score as the reward. The core of Q-learning is to iteratively update the Q value: Q(s,a)←Q(s,a)+α[r+γmax(Q(s',a'))-Q(s,a)], where s is the current state, a is the action, r is the immediate reward, γ is the discount factor, and α is the learning rate. Through a large number of iterations, an optimized strategy mapping is obtained to form a preliminary response strategy library.

[0173] Finally, the preliminary response strategy library is simulated and verified. Using digital twin technology, a virtual environment that is highly similar to the actual rail transit system is constructed. In this environment, each strategy is tested multiple times to simulate different operating conditions and emergencies. The performance of each strategy, including safety indicators, efficiency indicators, and robustness indicators, is recorded to generate a strategy impact assessment report. Based on this report, the multi-criteria decision analysis (MCDA) method, such as the analytic hierarchy process (AHP), is used to comprehensively score and sort the strategies. Finally, several strategies with the highest comprehensive scores are selected to form the target response strategy. For example, assume that a rail transit network contains 50 signal points. First, a 50×50 adjacency matrix A is constructed, where A[1][2]=1 indicates that signal points 1 and 2 are directly connected. The DFS algorithm is used to analyze the signal chain relationship and obtain the dependency matrix D. For example, D[1][3]=0.8 indicates that signal 1 has a strong influence on signal 3. Based on the D matrix, a decision tree algorithm is used to generate conflict detection rules, such as "If signal 1 is red and signal 3 is green, there is a conflict."

[0174] By analyzing the operation data of the past year and using the DTW algorithm for time series clustering, we can get three typical traffic flow patterns: weekday pattern, weekend pattern and holiday pattern. Taking the weekday pattern as an example, the average passenger flow during the morning peak from 7:00 to 9:00 is 8,000 people / hour, and the average passenger flow during the evening peak from 17:00 to 19:00 is 7,500 people / hour.

[0175] Based on these data, a virtual operation scenario was constructed and 10,000 Monte Carlo simulations were performed. In each simulation, 80% of the signals were randomly selected to switch in a normal sequence, and 20% of the signals were randomly switched. The results of each simulation were recorded, such as the average delay time, the minimum train spacing, etc. Safety constraints were applied for screening, such as "the minimum train spacing is not less than 120 seconds", and 8,000 valid simulation scenarios were obtained.

[0176] The efficiency scores of these 8,000 scenarios were calculated using the Pareto optimization method, considering the two objectives of average delay time (the lower the better) and total passenger transportation volume (the higher the better). 100 Pareto optimal solutions were obtained to form the efficiency score table.

[0177] Using the Q-learning algorithm, we set the learning rate α = 0.1, the discount factor γ = 0.9, and performed 100,000 iterations of learning. We obtained a Q-value table, which represents the expected benefits of taking different scheduling actions under different network conditions. Based on this Q-value table, we built a preliminary response strategy library containing 1,000 strategies.

[0178] These 1,000 strategies were verified in the digital twin environment, and each strategy was tested 100 times to simulate different situations including normal operation, equipment failure, and passenger flow surge. The average delay time, safety accident rate, energy consumption and other indicators of each strategy were recorded. These indicators were weighted using the AHP method to obtain a comprehensive score for each strategy. Finally, the 10 strategies with the highest scores were selected as the target response strategies. These strategies cover a variety of scenarios from daily operations to emergency response.

[0179] In a specific embodiment, the process of executing step S105 may specifically include the following steps:

[0180] (1) Classify and analyze the target response strategies to obtain signal switching sequences, train dispatching instructions, and platform management suggestions;

[0181] (2) Expand the signal switching sequence in time to obtain a signal light change schedule;

[0182] (3) Generate a train adjustment plan based on the train dispatching instructions, obtain the train operation diagram correction suggestions, and convert the platform management suggestions into passenger flow diversion measures to obtain the platform operation guide;

[0183] (4) Integrate data on the signal change schedule, train operation diagram revision suggestions, and platform operation guide to obtain a comprehensive dispatching plan;

[0184] (5) Generate a list of dispatcher operation steps based on the comprehensive dispatch plan and perform command conversion to obtain a human-computer interaction command set, and then perform hierarchical processing on the human-computer interaction command set, dividing it into emergency commands, routine commands, and advisory commands to obtain a hierarchical command library;

[0185] (6) converting the hierarchical instruction library into a console operation sequence to obtain a response control instruction;

[0186] (7) converting the operation prompt text based on the response control instruction to obtain the response prompt information, and timestamping and prioritizing the response control instruction and the response prompt information to obtain an ordered data packet;

[0187] (8) Encapsulate the ordered data packets to obtain a system-compatible data stream, and transmit the system-compatible data stream to the data interaction terminal.

[0188] Specifically, the named entity recognition (NER) algorithm in natural language processing (NLP) technology is used to parse the target response strategy text. The NER algorithm uses the conditional random field (CRF) model to learn feature templates through training data to identify key entities in the text, such as signal switching sequences, train scheduling instructions, and platform management suggestions. For example, for the strategy text "Change the signal of station A from red to green, and adjust the departure of train No. 103 to be delayed by 5 minutes", the NER algorithm will identify key entities such as "station A signal", "red to green", "train No. 103", and "delay of 5 minutes". The identified signal switching sequence is time-series expanded, and the dynamic time warping (DTW) algorithm in time series analysis is used. The DTW algorithm calculates the similarity of two time series, and the core is to find an alignment method that minimizes the sum of the distances between the corresponding points of the two sequences. Through the DTW algorithm, the switching sequences of different signals are aligned to a unified time axis to obtain a signal light change schedule. This schedule is a two-dimensional matrix of time-signal state, where each row represents a time point and each column represents the state of a signal light.

[0189] Generate train adjustment plans based on train dispatch instructions, using the genetic algorithm (GA) in the heuristic algorithm. GA iteratively optimizes solutions by simulating the process of natural selection and inheritance. Here, each chromosome represents a train adjustment plan, and the gene position represents the adjustment time of each train. The fitness function takes into account factors such as total delay time and passenger waiting time. Through operations such as crossover and mutation, after multiple generations of evolution, the optimal train operation diagram correction suggestion is obtained.

[0190] The platform management suggestions are converted into passenger flow diversion measures using a decision tree algorithm. Each internal node of the decision tree represents an attribute test, such as "whether the platform congestion exceeds 80%", and the leaf nodes represent specific diversion measures. The decision tree is constructed using the ID3 algorithm, and the attribute with the largest information gain is selected as the split point. The subtree is recursively constructed, and finally a complete decision tree is obtained, i.e., the platform operation guide.

[0191] The signal light change schedule, train operation diagram amendment suggestions and platform operation guide are integrated with the Dempster-Shafer evidence theory in data fusion technology. This method can handle uncertain and incomplete information by calculating the basic probability distribution function (BPA) of each evidence source, and then fusing multiple BPAs using the Dempster combination rule to obtain a comprehensive judgment. The fusion result forms a comprehensive scheduling plan that includes information in multiple dimensions such as time, space and passenger flow.

[0192] Generate a list of dispatcher operation steps based on the comprehensive dispatch plan, using process mining technology. Process mining discovers the actual business process model by analyzing event logs. Apply the α algorithm to extract the operation sequence from the dispatch plan and build a complete workflow network, where each node represents an operation step and the edge represents the dependency between steps. Then, this workflow is converted into a specific human-computer interaction instruction set.

[0193] The human-computer interaction instruction set is graded and the analytic hierarchy process (AHP) in the multi-criteria decision analysis (MCDA) method is used. AHP establishes a hierarchical model, performs pairwise comparisons, calculates weight vectors, and finally obtains the importance score of each instruction. According to the score, the instructions are divided into emergency instructions (score>0.7), regular instructions (0.3<score≤0.7) and advisory instructions (score≤0.3) to form a graded instruction library.

[0194] The hierarchical instruction library is converted into a sequence of console operations using a finite state machine (FSM) model. FSM consists of states, events, and transitions, and each instruction corresponds to one or more state transitions. By defining the initial state, terminal state, and transition rules, the instruction sequence is mapped to the FSM to obtain a series of state transitions, i.e., the response control instructions.

[0195] The operation prompt text conversion is based on the response control instruction, using natural language generation (NLG) technology. Using a template-based approach, a text template is defined for each type of instruction, and then the template is filled with specific parameters to generate human-readable response prompt information. For example, the template "Please switch the {signal} signal at {station} to {color}" may become "Please switch the entry signal at station A to green" after filling.

[0196] The response control instructions and response prompt information are timestamped and prioritized using a priority queue data structure. Each instruction and prompt is an object that contains timestamp and priority attributes. These objects are inserted into the priority queue, which is automatically sorted according to priority and timestamp to obtain ordered data packets.

[0197] Finally, the ordered data packets are encapsulated using the Protocol Buffers technology. Define the .proto file to describe the data structure, and then use the protoc compiler to generate the corresponding serialization and deserialization code. In this way, the complex data structure is converted into a compact binary format, a system-compatible data stream is obtained, and it is transmitted to the data interaction terminal through the network.

[0198] Let's take an example to illustrate the whole process: suppose that in a rail transit system, the target response strategy is "to deal with sudden large passenger flow, change the platform signal of Line B1 at Station A from red to green, extend the green light time by 30 seconds, adjust the departure delay of Train No. 103 by 2 minutes, and start the diversion plan for Area B on the platform." First, the NER algorithm identifies the key entities: Station A, Line B1, platform signal, red to green, extension of 30 seconds, Train No. 103, delay of 2 minutes, and diversion plan for Area B.

[0199] By performing time-series expansion on the signal switching sequence, we can obtain the signal light change schedule:

[0200] Time point | Platform signal of Line B1 at Station A

[0201] 0s|red, 5s|yellow, 10s|green, 40s|yellow, 45s|red;

[0202] Genetic algorithm is used to generate train adjustment plan, assuming that the population size is 100 and the evolutionary number is 50. After iterative optimization, the optimal solution is obtained: Train 103 is delayed by 120 seconds, Train 104 is advanced by 60 seconds, and Train 105 is delayed by 30 seconds.

[0203] Station management recommendations are converted into specific measures through a decision tree:

[0204] If the platform congestion is >80%: activate the B area evacuation plan;

[0205] If the waiting time is > 10 minutes: open the backup channel;

[0206] else: increase platform guide personnel, else: normal operation;

[0207] Using the Dempster-Shafer theory to integrate this information, assuming that the credibility of signal adjustment is 0.8, the credibility of train scheduling is 0.7, and the credibility of platform management is 0.6. Through the calculation of Dempster combination rules, the comprehensive credibility is 0.928, forming a highly reliable comprehensive scheduling plan.

[0208] Based on this scheme, the α algorithm is used to generate the dispatcher operation steps:

[0209] Step 1: Switch the platform signal of Line B1 at Station A;

[0210] Step 2: Notify Train No. 103 of delayed departure;

[0211] Step 3: Activate the evacuation plan for area B;

[0212] Step 4: Adjust the timetable of train No. 104 and No. 105;

[0213] Use the AHP method to grade these steps, construct a judgment matrix, calculate the eigenvector, and obtain the weight score. Assume that the score of step 1 is 0.8, step 2 is 0.6, step 3 is 0.7, and step 4 is 0.5. Based on this, steps 1 and 3 are divided into emergency instructions, step 2 is a regular instruction, and step 4 is a recommended instruction. Convert these instructions into an FSM model and define the state transition: initial state -> switch signal -> notify train -> start diversion -> adjust time -> end state

[0214] Use NLG technology to generate operation prompt text, such as "Please immediately switch the platform signal of Line B1 at Station A to green for 30 seconds."

[0215] Finally, all instructions and prompts are timestamped and prioritized, for example: [Timestamp: 2023-05-0110:00:00, Priority: 1] Switch the platform signal of Line B1 at Station A; [Timestamp: 2023-05-0110:00:05, Priority: 2] Start the diversion plan for Area B; [Timestamp: 2023-05-0110:00:10, Priority: 3] Notify Train No. 103 to delay departure. These ordered data packets are encoded through ProtocolBuffers to generate a compact binary data stream, which is finally transmitted to the dispatcher's data interaction terminal to guide him to perform efficient and accurate dispatch operations.

[0216] The embodiment of the present invention also provides a rail transit management system screen signal recognition and response system, such as Figure 3 As shown, the screen signal recognition and response system of a rail transit management system specifically includes:

[0217] The acquisition module 201 is used to simultaneously perform high-frequency image acquisition on multiple display screens in the rail transit management system control room to obtain multiple original image data streams, and perform parallel preprocessing and adaptive screen area positioning on the multiple original image data streams to obtain multiple standardized screen images;

[0218] An extraction module 202, configured to extract features from the plurality of standardized screen images by using a scale-invariant feature transformation algorithm to obtain a high-dimensional screen signal feature data set;

[0219] A matching module 203, configured to perform multi-threaded parallel color analysis and state matching on the high-dimensional screen signal feature data set to obtain multi-signal state information;

[0220] The simulation module 204 is used to perform scenario simulation on the multi-signal state information to obtain simulation scenario data, and to construct a signal response strategy on the simulation scenario data to obtain a target response strategy;

[0221] The transmission module 205 is used to generate a natural language interaction text according to the target response strategy, generate a response control instruction and a response prompt information according to the natural language interaction text, and transmit the response control instruction and the response prompt information to a preset data interaction terminal.

[0222] Through the collaborative work of the above modules, high-frequency image acquisition and parallel processing of multiple display screens in the control room have greatly improved the real-time and accuracy of signal recognition, effectively solving the problem of fatigue and error in manual monitoring; the scale-invariant feature transformation algorithm is used for feature extraction, combined with multi-threaded parallel color analysis and state matching, it can quickly and accurately identify various complex signal states, including flashing signals and gradient color signals, greatly improving the ability to identify abnormal situations; through scenario simulation and strategy construction of multi-signal state information, it can predict potential changes in traffic conditions and generate corresponding response strategies, which not only improves the foresight of decision-making, but also enhances the ability to respond to emergencies. In addition, natural language interactive texts, response control instructions and prompt information are automatically generated based on the target response strategy, which greatly reduces the workload of dispatchers and improves the accuracy and timeliness of instruction execution; the automation and intelligence of the entire process significantly improves the efficiency and safety of rail transit management, reduces human errors, and improves the efficiency and accuracy of screen signal recognition and response in the rail transit management system.

[0223] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the embodiments, a person skilled in the art should understand that the specific implementation modes of the present invention can still be modified or replaced by equivalents, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be included in the scope of the claims of the present invention.

Claims

1. A method for identifying and responding screen signals in a rail transit management system, characterized in that: include: High-frequency image acquisition is performed on multiple display screens in a rail transit management system control room at the same time to obtain multiple original image data streams, and the multiple original image data streams are parallel preprocessed and adaptively positioned in the screen area to obtain multiple standardized screen images, including: multiple video stream acquisition is performed on the multiple display screens to obtain multiple original image data streams, including the acquisition of near-infrared images in addition to visible light images, and the multiple original image data streams are subjected to frame decomposition processing to obtain multiple groups of continuous image frame sequences; the sampling frequency of the multiple groups of continuous image frame sequences is adjusted and screened to obtain an optimized image frame set, and the optimized image frame set is subjected to multi-spectral fusion Processing to obtain composite image data, and adaptively adjusting the exposure and white balance of the composite image data to obtain a light-balanced image; constructing a multi-scale image pyramid on the light-balanced image to obtain a multi-resolution image set, and performing parallel edge detection on the multi-resolution image set to obtain a plurality of screen contour candidate areas; screening the plurality of screen contour candidate areas by a deep learning target detection algorithm to obtain a screen position information set, and performing perspective transformation and geometric correction on the screen position information set to obtain a plurality of corrected screen images; performing local contrast enhancement and noise suppression on the plurality of corrected screen images to obtain the plurality of standardized screen images; Extracting features from the plurality of standardized screen images by using a scale-invariant feature transformation algorithm to obtain a high-dimensional screen signal feature data set; Performing multi-threaded parallel color analysis and state matching on the high-dimensional screen signal feature data set to obtain multi-signal state information; Performing scenario simulation on the multi-signal state information to obtain simulated scenario data, and constructing a signal response strategy on the simulated scenario data to obtain a target response strategy; Generate a natural language interaction text according to the target response strategy, generate a response control instruction and a response prompt information according to the natural language interaction text, and transmit the response control instruction and the response prompt information to a preset data interaction terminal.

2. The screen signal recognition and response method of the rail transit management system according to claim 1 is characterized in that: The step of extracting features from the plurality of standardized screen images by using a scale-invariant feature transformation algorithm to obtain a high-dimensional screen signal feature data set comprises: Constructing a TS-SIFT Gaussian pyramid for the multiple standardized screen images, generating a series of images of different scales by continuously performing Gaussian blurring and downsampling on the original images, obtaining a multi-scale image group of the rail transit screen, and performing Gaussian difference analysis on the multi-scale image group of the rail transit screen to obtain a differential image group; Performing local extreme point detection on the differential image group to obtain an initial feature point candidate set, and performing edge response screening on the initial feature point candidate set to obtain a signal feature point candidate region; Performing small-scale signal local extrema detection on the signal feature point candidate area to obtain an initial track signal feature point set, and performing sub-pixel signal positioning on the initial track signal feature point set to obtain position track signal feature points; Allocating the main direction of the traffic signal to the position track signal feature point to obtain the rotationally invariant track signal feature point; Extracting local signal gradient information from the rotationally invariant rail signal feature points to obtain a rail transit signal feature descriptor, and performing data enhancement on the rail transit signal feature descriptor to obtain a color-sensitive rail signal feature vector; Traffic signal feature dimensionality reduction processing is performed on the color-sensitive track signal feature vector to obtain the high-dimensional screen signal feature data set.

3. The screen signal recognition and response method of the rail transit management system according to claim 1 is characterized in that: The step of performing multi-threaded parallel color analysis and state matching on the high-dimensional screen signal feature data set to obtain multi-signal state information includes: Performing multi-thread allocation on the high-dimensional screen signal feature data set to obtain a parallel processing task group; Performing HSV color space conversion on each feature data in the parallel processing task group to obtain an HSV feature representation data set; Defining a dynamic color range for the HSV feature representation data set to obtain a signal color dynamic threshold, and generating signal color judgment standard data according to the signal color dynamic threshold; Performing feature vectorization processing on the signal color judgment standard data to obtain a color feature vector set, and performing similarity calculation on the color feature vector set and a preset signal state template to obtain a signal state candidate set; The signal state candidate set is subjected to signal weight analysis based on the pre-acquired historical state information to obtain a signal state weight set; and the high-dimensional screen signal feature data set is subjected to state matching based on the signal state weight set to obtain the multi-signal state information.

4. The screen signal recognition and response method of the rail transit management system according to claim 3 is characterized in that: The step of performing state matching on the high-dimensional screen signal feature data set based on the signal state weight set to obtain the multi-signal state information comprises: Performing time series analysis on the high-dimensional screen signal feature data set based on the signal state weight set to obtain a signal state change trend; Performing frequency analysis on the signal state change trend to obtain signal change periodic characteristics and non-periodic signals, and performing pattern matching on the signal change periodic characteristics to obtain a periodic signal pattern; Classifying and clustering the periodic signal pattern and the non-periodic signal to obtain signal type groups; Performing spatiotemporal correlation analysis on the signal type groups to obtain logical relationships between signals, and constructing a signal state transition diagram based on the logical relationships between signals to obtain a complex signal combination pattern; Performing flashing signal recognition and gradient color signal recognition on the complex signal combination pattern to obtain a target signal state; Performing text content analysis on the target signal state to obtain text-color association information; Based on the signal state weight set, the text-color association information is weightedly fused to obtain multi-signal state information.

5. The screen signal recognition and response method of the rail transit management system according to claim 1 is characterized in that: The step of performing scenario simulation on the multi-signal state information to obtain simulated scenario data, and constructing a signal response strategy for the simulated scenario data to obtain a target response strategy comprises: Performing line topology mapping on the multi-signal status information to obtain a track network status diagram, and performing signal linkage relationship analysis on the track network status diagram to obtain a signal interdependency list; generating a train operation conflict detection rule based on the signal interdependence list to obtain a safety constraint condition set; Performing statistical analysis on the pre-collected historical operation data to obtain a typical traffic flow pattern, and constructing a virtual train operation scenario based on the typical traffic flow pattern and the multi-signal state information to obtain an initial simulation scenario; Iterate the initial simulation scenario multiple times to simulate different signal switching sequences to obtain multiple groups of simulation results, and perform safety assessment on the multiple groups of simulation results to screen out simulation scenarios that meet the safety constraint condition set to obtain a valid simulation scenario set; Calculating the operating efficiency of the effective simulation scenario set to obtain an efficiency score table, and constructing a multi-level scheduling strategy based on the efficiency score table to obtain a preliminary response strategy library; The preliminary response strategy library is simulated and verified to obtain a strategy impact assessment report, and the preliminary response strategy library is screened according to the strategy impact assessment report to obtain the target response strategy.

6. The screen signal recognition and response method of the rail transit management system according to claim 5 is characterized in that: The step of generating a natural language interaction text according to the target response strategy, generating a response control instruction and a response prompt information according to the natural language interaction text, and transmitting the response control instruction and the response prompt information to a preset data interaction terminal includes: Classify and analyze the target response strategies to obtain signal switching sequences, train dispatching instructions and platform management suggestions; Performing time-series expansion on the signal switching sequence to obtain a signal light change schedule; Generate a train adjustment plan based on the train dispatch instruction, obtain a train operation diagram correction suggestion, and convert the platform management suggestion into a passenger flow diversion measure to obtain a platform operation guide; integrate the signal light change plan, train operation diagram correction suggestion and platform operation guide to obtain a comprehensive dispatch plan; Based on the comprehensive dispatching plan, a dispatcher operation step list is generated and command conversion is performed to obtain a human-computer interaction instruction set, and the human-computer interaction instruction set is graded and divided into emergency instructions, routine instructions and advisory instructions to obtain a graded instruction library; Convert the hierarchical instruction library into a console operation sequence to obtain a response control instruction; perform operation prompt text conversion based on the response control instruction to obtain response prompt information, and timestamp and prioritize the response control instruction and the response prompt information to obtain an ordered data packet; The ordered data packets are data encapsulated to obtain a system-compatible data stream, and the system-compatible data stream is transmitted to the data interaction terminal.

7. A rail transit management system screen signal recognition and response system, used to execute the rail transit management system screen signal recognition and response method according to any one of claims 1 to 6, characterized in that: include: The acquisition module is used to simultaneously perform high-frequency image acquisition on multiple display screens in the control room of the rail transit management system to obtain multiple original image data streams, and perform parallel preprocessing and adaptive screen area positioning on the multiple original image data streams to obtain multiple standardized screen images, including: performing multi-channel video stream acquisition on the multiple display screens to obtain multiple original image data streams, including the acquisition of near-infrared images in addition to visible light images, and performing frame decomposition processing on the multiple original image data streams to obtain multiple groups of continuous image frame sequences; adjusting and screening the sampling frequency of the multiple groups of continuous image frame sequences to obtain an optimized image frame set, and performing multi-light processing on the optimized image frame set. Spectral fusion processing is performed to obtain composite image data, and adaptive exposure and white balance adjustment is performed on the composite image data to obtain a light-balanced image; a multi-scale image pyramid is constructed on the light-balanced image to obtain a multi-resolution image set, and parallel edge detection is performed on the multi-resolution image set to obtain multiple screen contour candidate areas; the multiple screen contour candidate areas are screened by a deep learning target detection algorithm to obtain a screen position information set, and perspective transformation and geometric correction are performed on the screen position information set to obtain multiple corrected screen images; local contrast enhancement and noise suppression are performed on the multiple corrected screen images to obtain the multiple standardized screen images; An extraction module, used for performing feature extraction on the plurality of standardized screen images by using a scale-invariant feature transformation algorithm to obtain a high-dimensional screen signal feature data set; A matching module, used for performing multi-threaded parallel color analysis and state matching on the high-dimensional screen signal feature data set to obtain multi-signal state information; A simulation module, used to perform scenario simulation on the multi-signal state information to obtain simulation scenario data, and to construct a signal response strategy for the simulation scenario data to obtain a target response strategy; A transmission module is used to generate a natural language interaction text according to the target response strategy, generate a response control instruction and a response prompt information according to the natural language interaction text, and transmit the response control instruction and the response prompt information to a preset data interaction terminal.

Citation Information

Patent Citations

  • Method for detecting occupation state of rail section of rail transit station

    CN116985873A

  • Traffic emergency command method and system based on digital twinning

    CN118247954A