Intelligent Processing Method and System for AI Multimodal Data in Rail Transit
By disassembling the feature dimension relationship of multimodal data in the rail transit system, establishing a scene feature analysis map, and performing granularity sensitivity evaluation and configuration, the problem of low matching between scenes and data in multimodal data processing is solved, and the accuracy and efficiency improvement of task scenario recognition is achieved.
Patent Information
- Application Number
- CN202510473714.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The multimodal data processing efficiency in the rail transit system is low, the scene and data matching degree are not high, and the feature enhancement granularity configuration is insufficient, resulting in low task scene recognition accuracy and efficiency.
Through the scene-data analysis relationship matching the corresponding modal data of the requested task scenario, the data feature dimension relationship is disassembled, the scene feature analysis map is established, and the feature enhancement granularity sensitivity evaluation and granularity configuration are performed, and the task scene recognition is finally achieved using enhanced features.
It improves the accuracy and efficiency of task scene recognition and improves the overall effect of multimodal data processing.
Smart Images

Figure CN119988895B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to an intelligent processing method and system for AI multi-modal data in rail transit. Background Art
[0002] With the acceleration of urbanization and the rapid development of rail transit, rail transit systems such as subways and light rails have become an important part of urban public transportation. However, the operation and maintenance of rail transit systems face huge challenges, especially in data management and fault diagnosis. Rail transit systems generate a large amount of multi-modal data every day, including visual data, sensor data, text data, voice data, and geographic information data, etc. These data have diverse sources and complex structures, and traditional single-modal data processing methods are difficult to effectively integrate and utilize these multi-modal data, resulting in problems such as low data processing efficiency, low matching degree between scenarios and data, and insufficient configuration of feature enhancement granularity, which in turn affect the accuracy and efficiency of task scenario recognition.
[0003] In actual operation, the fault diagnosis, resource allocation, and operation management of rail transit systems highly rely on the intelligent processing and analysis of multi-modal data. However, existing data processing technologies often cannot efficiently process multi-modal data, especially when facing complex task scenarios, it is difficult to accurately extract key features and perform effective scenario recognition. This not only increases the operation cost but also may lead to potential safety hazards. Therefore, how to efficiently process multi-modal data and improve the accuracy and efficiency of task scenario recognition has become a key technical problem in the intelligent development of rail transit. Summary of the Invention
[0004] This application provides an intelligent processing method and system for AI multi-modal data in rail transit, aiming to solve the technical problems of low recognition efficiency caused by low matching degree between scenarios and data and insufficient configuration of feature enhancement granularity in the process of processing rail transit multi-modal data.
[0005] In the first aspect disclosed by this application, an intelligent processing method for AI multi-modal data in rail transit is provided. The method includes: using the scene-data analysis relationship to match the corresponding modal data of the requested task scenario, disassembling the data feature dimension relationship of the matched modal data, and establishing a scene feature analysis map; performing a sensitivity evaluation of the feature enhancement granularity on the matched modal data according to the scene feature analysis map, and configuring the granularity according to the evaluation sensitivity; enhancing the features of the matched modal data in the scene feature analysis map according to the configured granularity, and using the enhanced features to perform recognition of the requested task scenario for the requested task scenario to obtain the recognition result of the requested task scenario.
[0006] Another aspect disclosed in this application provides an intelligent processing system for rail transit AI multimodal data. The system includes: a feature dimension decomposition module: using the scenario - data analysis relationship to match the corresponding modal data of the requested task scenario, decomposing the data feature dimension relationship of the matched modal data, and establishing a scenario feature analysis map; a granularity configuration module: evaluating the granularity sensitivity of the feature enhancement of the matched modal data according to the scenario feature analysis map, and performing granularity configuration according to the evaluation sensitivity; a task scenario recognition module: enhancing the features of the matched modal data in the scenario feature analysis map according to the configured granularity, and using the enhanced features to recognize the requested task scenario, obtaining the recognition result of the requested task scenario.
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0008] The above intelligent processing method for rail transit AI multimodal data uses the scenario - data analysis relationship to match the corresponding modal data of the requested task scenario, decomposes the data feature dimension relationship of the matched modal data, and establishes a scenario feature analysis map; then evaluates the granularity sensitivity of the feature enhancement of the matched modal data according to the scenario feature analysis map, and performs granularity configuration according to the evaluation sensitivity; finally, enhances the features of the matched modal data in the scenario feature analysis map according to the configured granularity, and uses the enhanced features to recognize the requested task scenario, obtaining the recognition result of the requested task scenario, achieving the technical effects of improving the recognition accuracy of the task scenario and enhancing the processing efficiency of multimodal data through the scenario feature analysis map, granularity sensitivity evaluation, and feature enhancement.
[0009] The above description is only an overview of the technical solutions of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the following specific embodiments of this application are specifically given. Description of the Drawings
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0011] Figure 1 It is a schematic flowchart of an intelligent processing method for rail transit AI multimodal data in an embodiment.
[0012] Figure 2It is an architecture diagram of an intelligent processing system for rail transit AI multi-modal data in an embodiment.
[0013] Description of the reference numerals: Feature dimension decomposition module 11, granularity configuration module 12, task scenario recognition module 13. Specific implementation manner
[0014] In an embodiment of the present application, by providing an intelligent processing method and system for rail transit AI multi-modal data, the technical problems of low matching degree between scenarios and data and low recognition efficiency caused by insufficient granularity configuration of feature enhancement in the process of processing rail transit multi-modal data are solved.
[0015] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0016] It should be noted that the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices.
[0017] Embodiment 1, as Figure 1 shown, the present application provides an intelligent processing method for rail transit AI multi-modal data, and the method includes:
[0018] Using the scenario-data analysis relationship to match the corresponding modal data of the requested task scenario, decomposing the data feature dimension relationship of the matched modal data, and establishing a scenario feature analysis map.
[0019] In the embodiments of the present application, first, according to the requirements of the task scenario, using the relationship between the scenario and data analysis, multimodal data related to the task scenario is matched. These multimodal data include visual data (such as images captured by cameras), sensor data (such as environmental data like temperature and humidity), text data (such as operation logs), etc. Subsequently, the multimodal data that has been matched is disassembled in terms of the data feature dimensions. Specifically, each modal data is decomposed into features of multiple dimensions. For example, for visual data, it can be disassembled into features such as color and shape; for sensor data, it can be disassembled into time series features, sampling frequency features, etc. Through this disassembly, the internal structure and information of each modal data can be understood more meticulously. After completing the disassembly of the feature dimensions, based on these feature dimension relationships, a scene feature parsing map is established. This map is a visual graphical structure, where nodes represent the task scenario, modal data, and their corresponding multi-dimensional features, and the connection relationships represent the feature quantization relationships between these nodes. For example, a certain task scenario may be related to multiple modal data, and certain feature dimensions of these modal data may jointly affect the recognition result of this scenario. By constructing such a map, the complex relationship between the task scenario and multimodal data can be clearly displayed, providing basic support for subsequent feature enhancement and scene recognition.
[0020] Exemplarily, assume the task scenario is track health monitoring. In this scenario, multiple modal data are required to comprehensively evaluate the health status of the track, including visual data, texture features, vibration data, etc. Among them, visual data refers to the images captured by cameras of the track surface, and these images have spatial resolution features and texture features, which can help identify physical damages such as cracks and wear on the track surface; vibration data refers to the vibration data collected by sensors when the train passes through, and these data have frequency features, which can help identify looseness or deformation of the track structure. In this example, the task scenario of track health monitoring is related to visual data and sensing data, and the spatial resolution and texture features of the visual data, as well as the frequency features of the vibration data of the sensing data, jointly affect the recognition result of the track health status. By analyzing the feature dimensions of these multimodal data, the health status of the track can be evaluated more comprehensively, thereby improving the accuracy and efficiency of monitoring.
[0021] Furthermore, the present application provides the corresponding modal data for matching the requested task scenario using the scenario-data analysis relationship, which includes:
[0022] Connect to the rail transit data acquisition platform to obtain all modal data and construct a full-modal database; according to the entire operation cycle of rail transit, perform task scenario parsing on each modal data in the full-modal database to determine the scenario-data analysis relationship.
[0023] Preferably, before matching the corresponding modal data for the request task scenario, it is first necessary to connect to the rail transit data collection platform to obtain modal data from different sources, including visual data (such as track images, platform cameras), sensor data (such as vibration, temperature, current), text data (such as social media public opinion, maintenance records), voice data (such as alarm broadcasts, fault audio), geographical information data (such as track topology, train position information), etc. By collecting these data in real-time or regularly, a full-modal database can be constructed to provide a basis for subsequent task analysis and processing. These data will be subjected to operations such as data cleaning, format conversion, and timestamp alignment during storage to ensure data quality and consistency. Subsequently, according to the entire operation cycle of rail transit, combined with actual business requirements and historical samples, basic events for multiple task scenarios are defined. For example, track health monitoring (detecting track cracks, deformations, foreign object intrusion), passenger flow prediction and scheduling optimization (predicting passenger flow density during peak hours and optimizing train departure intervals), emergency response (quickly locating the event location and generating emergency plans), equipment fault diagnosis (real-time monitoring of the status of train bearings, signal systems, etc.). Then, based on these basic events, through accident factor decomposition and data relationship sorting, a scenario-data analysis relationship is established. This scenario-data analysis relationship records the data and impacts involved in each basic event, providing basic support for subsequent data processing and task identification. For example, in the track health monitoring scenario, visual data (such as track images), sensing data (such as vibration data), and geographical information data (such as track topology) will be associated to detect track cracks or deformations; in the passenger flow prediction and scheduling optimization scenario, visual data (such as camera counting data), sensing data (such as turnstile passage data), and text data (such as social media public opinion) will be combined to predict the passenger flow during peak hours and optimize the train departure intervals based on these data; in the emergency response scenario, correlation analysis will be performed based on visual data (such as smoke images) and sensor data (such as temperature and smoke concentration) to quickly locate the event location and generate corresponding emergency plans; in the equipment fault diagnosis task, an association will be established based on sensor data (such as vibration and current signals) and text data (such as maintenance records) to monitor the status of train equipment. Through the above process, multi-modal data related to each task scenario can be extracted from the full-modal database according to the operation requirements of rail transit, and the relationship between the scenario and data analysis can be determined, providing basic support for subsequent data processing and task identification.
[0024] Furthermore, the present application provides task scenario parsing for each modal data in the full-modal database to determine the scenario-data analysis relationship, including:
[0025] Taking the task scenario as the top event, the top event is decomposed into basic events based on historical samples to obtain the basic events of the task scenario; the basic events of the task scenario are decomposed into accident factors according to historical samples to obtain the accident factors and event impact function relationships of the basic events; based on the accident tree algorithm, a causal relationship tree is constructed according to the top event-basic event-accident factor and its event impact function relationship; based on the causal relationship tree, the data relationships of each task scenario are sorted out, a scenario data association matrix is established, and the scenario-data analysis relationship is obtained.
[0026] Optionally, first, based on business needs and the entire operational cycle of rail transit, multiple task scenarios are defined and used as top-level events. This means starting the analysis from a high-level task scenario, such as equipment fault diagnosis or track health monitoring. These task scenarios represent important operational goals or issues in the entire rail transit system. Subsequently, based on historical sample data, the top-level events are decomposed into basic events. This process breaks down the task scenarios into multiple more basic events. For example, in the equipment fault diagnosis task, basic events such as sensor anomalies or mechanical damage are involved. By analyzing the event identifiers of each set of data in the historical sample data, multiple basic events for each task scenario can be identified, such as track surface crack events and track structure loosening events in the track health monitoring scenario. Subsequently, these basic events are decomposed into accident factors based on historical sample data. Accident factors refer to various factors that may cause a basic event to occur, which may include environmental factors, operational errors, or equipment failures. Data corresponding to each fundamental event is extracted, and the correlation between each factor in the data associated with each fundamental event and the fundamental event is quantified using the Pearson correlation coefficient, Spearman rank correlation coefficient, and other methods. For example, when the fundamental event is track crack size, data related to this event is obtained from historical sample data, such as temperature, humidity, wheel-rail friction, and material aging. The correlation between each factor and track crack size is quantified using the Pearson correlation coefficient, Spearman rank correlation coefficient, and other methods. The correlation of each factor is then checked using the lower limit of correlation. Factors with a correlation value below the lower limit are eliminated, and the remaining factors are considered accident factors. Using these accident factors as indices, the extracted data is analyzed to calculate the probability of each accident factor for each fundamental event (calculated by the ratio of the number of occurrences of the accident factor to the number of occurrences of the fundamental event). By comparing the probability of each accident factor with the lower limit of probability, the logical relationship between each accident factor and the basic event is determined (greater than is AND, less than is OR). Finally, each accident factor is associated with the basic event through correlation and logical relationship to form the event impact function relationship of each basic event. The event impact function relationship represents the degree of influence of each accident factor on the occurrence of the basic event.
[0027] After obtaining the accident factors and event impact function relationships for each basic event, the fault tree algorithm is applied to establish a causal relationship tree. The fault tree algorithm is a logic-based tool used to analyze the relationships between events and factors. By taking the top event as the top-level node of the tree, i.e., the root node, the basic events as intermediate nodes connected to the top event, and the accident factors as leaf nodes connected to the corresponding basic events, an initial causal relationship tree is constructed. Subsequently, according to the logical relationships in the event impact function relationships, the logical relationships between the nodes in the initial causal relationship tree are defined to represent how the accident factors jointly cause the basic events and how the basic events jointly cause the top event. Generally, the relationship between the basic event and the top event is OR. For example, basic event 1 (rail surface crack) is caused by accident factor 1 (wheel-rail friction) or accident factor 2 (material aging); basic event 2 (rail structure loosening) is jointly caused by accident factor 1 (bolt loosening) and accident factor 2 (foundation settlement); the top event (rail health monitoring) is caused by basic event 1 (rail surface crack) or basic event 2 (rail structure loosening). After the logical relationship configuration is completed, the causal relationship tree is obtained. Through this causal relationship tree, the dependence relationships and mutual influences between different events and factors can be clearly understood. Then, based on the causal relationship tree, the data relationships in each task scenario are sorted out to establish a scenario data association matrix. The data association matrix is a two-dimensional matrix used to describe the relationships between different modal data and the events and factors in the task scenario. In this matrix, the rows represent different data modalities (such as sensor data, visual data, text data, etc.), the columns represent different events or accident factors, and the values in the matrix represent the correlation and logical relationships between the data and the events or factors. This matrix helps to identify which data is most important for the events or factors in a specific task scenario, thereby further optimizing the use of data and the analysis of the task scenario. Through the above process, a complete scenario-data analysis relationship is finally obtained, clarifying the analysis relationships between each task scenario and multi-modal data, providing structured support for subsequent data processing and task identification.
[0028] Furthermore, the present application provides that the full-modal database at least includes: visual data, sensor data, text data, voice data, and geographic information data.
[0029] Optionally, a full-modal database refers to a database containing various types of data. In a rail transit system, this data comes from different monitoring devices and information systems and is used for comprehensive analysis and decision-making. Specifically, the full-modal database at least includes visual data, sensor data, text data, voice data, and geographic information data. Among them, visual data mainly comes from cameras and usually includes track images, carriage monitoring videos, platform monitoring pictures, etc. This data is used to detect track conditions, identify passenger behaviors, monitor the operating status of equipment, etc. For example, through image analysis, it can be detected whether there are cracks or other obstacles on the track; sensor data comes from various sensors installed on equipment and tracks, such as temperature sensors, vibration sensors, pressure sensors, etc. This data can provide real-time physical state monitoring and is used to evaluate equipment health, track stability, the smoothness of train operation, etc. For example, sensors can help monitor the vibration of trains or the temperature changes of tracks; text data comes from different documents, reports, sensor logs, maintenance records, operation manuals, social media, etc. This data is usually structured or unstructured and contains information such as system operation conditions, fault reports, and operation instructions. By analyzing text data, potential fault information, maintenance suggestions, or user feedback can be extracted; voice data includes voice instructions from station announcements, passengers or staff, alarm signals, communication records, etc. Through voice recognition and analysis, this data can help monitor and analyze passenger behaviors, station security, equipment failures, or emergencies. For example, fault alarm audio or emergency broadcasts can be used to determine whether an emergency has occurred; geographic information data includes the geographical location information of the rail transit network, station locations, train locations, road topologies, etc. This type of data is used for map display, route planning, traffic flow monitoring, fault location, etc. Through a geographic information system (GIS), precise positioning and efficient management of the rail transit system can be achieved. These different types of data form a multi-modal data system, which can provide a comprehensive view of all aspects of rail transit, thus helping with real-time monitoring, fault diagnosis, resource optimization, and intelligent decision-making. By integrating these modal data, the intelligent operation and maintenance of the system can be better achieved, the operation efficiency can be improved, and potential risks can be reduced.
[0030] Furthermore, this application provides a method for disassembling the relationship of data feature dimensions of the matching modal data and establishing a scene feature analysis map, including:
[0031] When the matching modal data is single-modal data, disassemble the data feature dimensions of the single-modal data to obtain multi-dimensional features; according to the single-modal data and its corresponding multi-dimensional features, perform visual graphic conversion to establish the scene feature analysis map, where the nodes are task scenarios, single-modal data, and multi-dimensional features, and the connection relationship represents the feature quantization relationship between the nodes.
[0032] Optionally, after matching the modal data of the current task scenario through the scenario-data analysis relationship, it is determined whether the matched modal data is single-modal data, that is, it is determined whether the current modal data comes from a single data source and only contains one type of information. If the matched modal data is single-modal data, the single-modal data will be disassembled into features, which are transformed into multi-dimensional features. These features are used to quantify different attributes of the data and provide support for subsequent analysis and visualization. For example, for visual data, the feature dimensions may include resolution, pixels, texture, edge strength, color, etc.; for vibration data, the feature dimensions may include sampling frequency, time-domain statistics, frequency-domain energy, multi-axis correlation, envelope features, etc.; for text data, the feature dimensions include word frequency distribution, entity density, syntactic complexity, sentiment polarity, topic distribution, etc.; for location data, its feature dimensions may be coordinate accuracy, timestamp synchronization, spatial topology, elevation features, motion trajectories, etc. After completing the disassembly of the feature dimensions of the single-modal data, these features are converted into visual graphs, that is, the relationships between the features are represented graphically to facilitate analysis and understanding. The scene feature parsing map generated in this process will help the subsequent recognition and analysis of the task scenario. In the scene feature parsing map, three types of nodes are defined, namely task scenario nodes, single-modal data nodes, and multi-dimensional feature nodes. Connection relationships are also defined between the nodes. The connection relationships represent the feature quantification relationships between the nodes, that is, the relevance between different data features and the task scenario. For example, connecting the task scenario node and the single-modal data node indicates the application of this modal data in this task scenario; connecting the single-modal data node and the multi-dimensional feature node indicates the multi-dimensional features of this modal data; connecting the quantification relationships between the multi-dimensional feature nodes indicates the mutual influence and dependence relationships between the features. These nodes and connection relationships together form the scene feature parsing map. Through the above process, the data feature dimensions of the single-modal data can be disassembled, and the scene feature parsing map can be established through visual graph conversion, clarifying the relationships between the task scenario, single-modal data, and multi-dimensional features, providing basic support for subsequent feature enhancement and task scenario recognition.
[0033] Furthermore, the present application provides a method for establishing a scene feature parsing map, including:
[0034] When the matched modal data is multiple modal data, the data feature dimensions of each modal data are disassembled respectively to obtain the multi-dimensional features of each modal data; according to the multi-dimensional features of each modal data, the fusion relationship of cross-modal data is analyzed with the task scenario as the target to establish cross-modal feature relationships; according to the multi-dimensional features of each modal data and the cross-modal feature relationships, visual graph conversion is performed to establish the scene feature parsing map.
[0035] Optionally, when the matching modal data is multiple modal data, the above-mentioned data feature dimension decomposition will be performed on each modal data first. After the data feature decomposition of each modal data is completed, the data of each modal will be converted into corresponding multi-dimensional features. Subsequently, aiming at the task scenario, the fusion relationship analysis of cross-modal data will be carried out. The purpose of cross-modal fusion is to combine data from different modalities (such as visual data and vibration data) to understand the task scenario more comprehensively. To achieve this, the different modal data will be aligned in time and space. For example, the timestamps of visual data, vibration data, and position data will be synchronized to ensure the corresponding relationship of the data at the same time point and spatial position. Then, according to the abnormal samples of the task scenario, the co-occurrence of the multi-dimensional features of the cross-modal data will be analyzed. For example, in the abnormal sample of track crack, the edge feature of visual data (edge intensity is 0.7) and the frequency domain energy of vibration data (low-frequency energy accounts for 60%) appear simultaneously; in the abnormal sample of track structure looseness, the envelope feature of vibration data (the peak of the envelope line appears at 50Hz) and the spatial topology of position data (the track topology structure is a straight line segment) appear simultaneously. After that, the confidence of the cross-modal dimension features will be calculated, and based on the calculation result of the confidence, the fusion relationship between the cross-modal data will be defined and the defined fusion relationship will be stored to obtain the cross-modal feature relationship, which serves as the basis for subsequent task scenario recognition. Then, according to the cross-modal feature relationship, the multi-dimensional features will be fused, and then according to the fused characteristics, the visualization graphic conversion will be carried out to establish a scene feature analysis map. This map includes task scenario nodes, single-modal data nodes, multi-dimensional feature nodes, fusion feature nodes, and the connection relationships between the nodes. Through the above process, the feature dimension decomposition of multi-modal data can be carried out, the fusion relationship of cross-modal data can be analyzed, and the scene feature analysis map can be established through visualization graphic conversion to clarify the relationship between the task scenario, multi-modal data, and their feature dimensions, providing basic support for subsequent feature enhancement and task scenario recognition.
[0036] Furthermore, the present application provides a method for analyzing the fusion relationship of cross-modal data aiming at the task scenario, including:
[0037] Aligning the cross-modal data in time and / or space; obtaining the co-occurrence data of the multi-dimensional features of the cross-modal data according to the abnormal samples of the task scenario, where the co-occurrence data is the data in which the dimension features of the cross-modal data appear simultaneously in the abnormal samples; calculating the confidence of the cross-modal dimension features according to the independent occurrence data of the single-modal data features in the abnormal samples of the cross-modal data and the co-occurrence data; determining the fusion relationship of the cross-modal data according to the confidence of the cross-modal dimension features.
[0038] Optionally, when processing data of multiple modalities, it is necessary to ensure that these data are aligned in time and space. For example, visual data (such as track images) and sensor data (such as vibration signals) may be collected at different times and spatial positions, so synchronization and alignment are required. During this process, different modality data will be synchronized according to timestamps to ensure that the timestamps of different modality data are aligned. If geographical location data (such as sensor positions and track image positions) are involved, it is also necessary to ensure the spatial consistency of different data sources. This is achieved by mapping data from different sources to the same coordinate system through coordinate transformation, thereby ensuring the consistency of data in physical space. Once the cross-modal data are aligned in time and / or space, abnormal samples of the current task scenario are retrieved from the abnormal sample library, such as track cracks, loose track structures, etc. Then, keyword comparison is performed between the abnormal samples and the multi-dimensional features of the cross-modal data to identify the co-occurrence data between the abnormal samples and the multi-dimensional features. This co-occurrence data represents the situation where the multi-dimensional features of different modality data appear simultaneously in the abnormal samples. For example, a certain frequency domain feature in the vibration signal may appear simultaneously with a certain feature in the visual data (such as the crack edge of the track image) in the abnormal sample. Subsequently, based on the co-occurrence data obtained above, the confidence of the cross-modal dimensional features is calculated. Confidence is an important indicator for measuring the strength and reliability of the relationship between different modality data features. During this process, it is analyzed whether the data features of each modality appear independently in the abnormal samples. For example, if the low-frequency energy of the vibration signal appears repeatedly in multiple abnormal samples but is independent of other modality features (such as visual data), it can be considered that the independence of this feature is relatively strong. Then, calculations are performed based on the co-occurrence data. The co-occurrence data represents the co-occurrence of multiple modality features in the same abnormal sample. By calculating the frequency of the co-occurrence data, etc., the understanding of the relationship between cross-modal features is further enhanced. After that, combining the independently occurring data and the co-occurrence data, the confidence of the cross-modal dimensional features is calculated. This confidence value reflects the strength of the association between cross-modal features and helps to determine whether there is an effective fusion relationship. Finally, based on the calculated confidence of the cross-modal dimensional features, the fusion relationship between different modality data can be further determined. If the confidence is relatively high, it indicates that there is a strong correlation between different modality data in the task scenario, and these data can be fused together to more accurately analyze the task scenario and make decisions. For example, in equipment fault diagnosis, if the visual data and vibration data have a relatively high confidence in abnormal situations, it means that there is a strong association between these features and they are suitable for closer fusion. At this time, the data of the two can be fused to form a more accurate fault diagnosis result. Through the above process, the system can complete the alignment of cross-modal data, co-occurrence data analysis, confidence calculation, and fusion relationship determination, providing basic support for the subsequent construction of the scenario feature parsing map and task scenario recognition, and improving the accuracy and efficiency of task scenario analysis.
[0039] Furthermore, the present application provides a method for calculating the confidence of cross-modal dimensional features, including:
[0040] Obtaining a single anomaly evaluation value of each single-modal data feature from the abnormal samples respectively according to the independently occurring data; obtaining a co-occurrence anomaly evaluation value according to the co-occurrence data; calculating a difference coefficient between the sum of the single anomaly evaluation values corresponding to the single-modal data features in the co-occurrence data and the co-occurrence anomaly evaluation value to obtain the confidence of the cross-modal dimensional features.
[0041] Optionally, in the process of cross-modal data processing, the anomaly evaluation of single-modal data features and the anomaly evaluation of co-occurrence data are very important steps for calculating the confidence of cross-modal dimensional features. Specifically, first, calculate the absolute difference between the independently occurring data in each single-modal data feature and the mean value of the normal samples of this data, and then calculate the ratio of the calculation result to the standard deviation of this data to obtain the single anomaly evaluation value of each single-modal data feature. At the same time, in the abnormal samples, count the number of times when different modal data features appear abnormally at the same time, and calculate the ratio of the counted co-occurrence times to the total number of feature appearances to obtain the co-occurrence anomaly evaluation value of the co-occurrence data. Subsequently, accumulate the single anomaly evaluation values of all independently occurring data, and then calculate the ratio of the accumulated result to the co-occurrence anomaly evaluation value of the co-occurrence data to determine the difference coefficient. The difference coefficient helps to measure the difference between the individual anomaly of each modal feature and its anomaly intensity when co-occurring in multi-modal data. Finally, by using the difference coefficient as the confidence of the cross-modal dimensional features, it provides a basis for determining the subsequent cross-modal data fusion relationship.
[0042] Performing a feature enhancement granularity sensitivity evaluation on the matching modal data according to the scene feature parsing map, and configuring the granularity according to the evaluation sensitivity.
[0043] In one embodiment, when performing the evaluation of the feature enhancement granularity sensitivity, it is first necessary to analyze the influence of the features of different modal data on the task scenario at different granularities. Granularity refers to the level and resolution of data features in terms of details. Fine-grained features usually contain more detailed information and can provide higher precision, while coarse-grained features may be more concise, cover a wider range, but may not be accurate enough in terms of details. During this process, multiple granularity levels will be defined, each level corresponding to different feature enhancement parameters. Then, according to the scene feature parsing map, the dimension features are enhanced at multiple granularity levels by using the feature enhancement parameters corresponding to the preset granularity levels. Subsequently, using the recognition result of the requested task scenario, the sensitivity evaluation is performed on the enhancement processing at each granularity level to obtain the feature recognition sensitivity, and the feature recognition sensitivity is judged according to the sensitivity requirement threshold, and the granularity level that meets the sensitivity requirement threshold is selected as the configured granularity to ensure that the task scenario can obtain appropriate information. Through the above process, it is possible to perform the evaluation of the feature enhancement granularity sensitivity on the matching modal data based on the scene feature parsing map, and perform the granularity configuration according to the evaluation sensitivity, thereby improving the accuracy and efficiency of the task scenario recognition.
[0044] Further, the present application provides a method for performing the evaluation of the feature enhancement granularity sensitivity on the matching modal data according to the scene feature parsing map and performing the granularity configuration according to the evaluation sensitivity, including:
[0045] Setting the granularity enhancement parameters corresponding to the granularity levels of the feature increase; according to the scene feature parsing map, performing the multi-granularity level enhancement processing on the dimension features by using the granularity enhancement parameters corresponding to the preset granularity levels; using the recognition result of the requested task scenario, performing the sensitivity evaluation on the enhancement processing at each granularity level to obtain the feature recognition sensitivity; based on the recognition target of the requested task scenario and the sensitivity requirement threshold, performing the granularity determination on the feature recognition sensitivity to determine the configured granularity.
[0046] Optionally, according to historical data or domain knowledge, multiple granularity levels are preset, and each level corresponds to different feature enhancement parameters. For example, low-granularity enhancement, global distribution adjustment, and the feature enhancement parameter is 0.1; medium-granularity enhancement: segmented processing, and the feature enhancement parameter is 0.5; high-granularity enhancement: local feature repair, and the feature enhancement parameter is 0.9. Subsequently, according to the scene feature parsing graph, multi-granularity level enhancement processing is performed on the dimensional features using the granularity enhancement parameters corresponding to the preset granularity levels. For example, in the track health monitoring scenario, high-granularity enhancement will be selected to sharpen the local crack area (such as Vision Transformer attention focusing); in the passenger flow prediction and scheduling optimization scenario, low-granularity enhancement will be selected to perform global illumination simulation (generate platform images at different time periods). Then, using the recognition result of the request task scenario, a sensitivity evaluation is performed on the enhancement processing of each granularity level to obtain the feature recognition sensitivity S. This process is obtained by calculating the proportion of correctly recognized samples in the total samples (i.e., the accuracy rate). Then, the sensitivity is compared with the scene threshold, and the scene threshold is determined based on the sensitivity requirement threshold θ, such as high-sensitivity scenario (S > θ), medium-sensitivity scenario (θ / 2 < S ≤ θ), low-sensitivity scenario (S ≤ θ / 2), and then the corresponding configured granularity is matched according to the comparison result. Through the above process, the system can perform feature enhancement granularity sensitivity evaluation on the matched modal data based on the scene feature parsing graph, and perform granularity configuration according to the evaluation sensitivity, thereby improving the accuracy and efficiency of task scenario recognition.
[0047] Perform feature enhancement on the matched modal data in the scene feature parsing graph according to the configured granularity, and use the enhanced features to perform request task scenario recognition on the request task scenario to obtain the recognition result of the request task scenario.
[0048] In one embodiment, according to the configured granularity, feature enhancement processing will be performed on the matched modal data in the scene feature parsing graph. The process of feature enhancement is to adjust the detail level of the feature according to different granularity requirements and improve the expression ability of the data. Specifically, according to the feature requirements of high, medium, and low granularity, the corresponding data features will be enhanced so that they can provide richer information in task scenario recognition. For example, local crack area sharpening is performed through Vision Transformer attention focusing, and high-frequency noise is processed according to the time window to perform segmented noise reduction on the vibration signal. These enhanced features are input into the task scenario recognition model to identify a specific task scenario. The task scenario recognition model analyzes the enhanced data features and outputs the recognition result of the task scenario. This result indicates whether the task scenario is successfully recognized and the accuracy of the recognition. In this way, effective task scenario recognition can be performed based on task requirements and feature granularity configuration, improving the accuracy and reliability of recognition.
[0049] For the task scenario recognition model, it is constructed based on a convolutional neural network. The training data includes sample modal features and sample task scenarios. During the training process, iterative training is performed through steps such as forward propagation, loss calculation, backpropagation, and parameter optimization until the maximum number of iterations is reached or the model converges. After the training is completed, the model is verified. If the accuracy meets the expected expectation, the current task scenario recognition model is output. Otherwise, hyperparameters such as the learning rate and the number of training batches are adjusted to further improve the recognition effect of the model.
[0050] In summary, the embodiments of the present application at least have the following technical effects:
[0051] The embodiments of the present application use the scenario - data analysis relationship to match the corresponding modal data of the requested task scenario, disassemble the data feature dimension relationship of the matched modal data, and establish a scenario feature parsing map; perform a feature enhancement granularity sensitivity evaluation on the matched modal data according to the scenario feature parsing map, and perform granularity configuration according to the evaluation sensitivity; perform feature enhancement on the matched modal data in the scenario feature parsing map according to the configured granularity, and use the enhanced features to perform recognition of the requested task scenario to obtain the recognition result of the requested task scenario. These technical effects together solve the technical problems of low matching degree between scenarios and data and insufficient feature enhancement granularity configuration in the process of rail transit multimodal data processing, and achieve the technical effects of improving the recognition accuracy of task scenarios and enhancing the efficiency of multimodal data processing through the scenario feature parsing map, granularity sensitivity evaluation, and feature enhancement.
[0052] Embodiment 2, based on the same inventive concept as the intelligent processing method of rail transit AI multimodal data in the foregoing embodiment, as Figure 2 shown, the present application provides an intelligent processing system for rail transit AI multimodal data. The system includes: a feature dimension disassembling module 11: using the scenario - data analysis relationship to match the corresponding modal data of the requested task scenario, disassembling the data feature dimension relationship of the matched modal data, and establishing a scenario feature parsing map; a granularity configuration module 12: performing a feature enhancement granularity sensitivity evaluation on the matched modal data according to the scenario feature parsing map, and performing granularity configuration according to the evaluation sensitivity; a task scenario recognition module 13: performing feature enhancement on the matched modal data in the scenario feature parsing map according to the configured granularity, and using the enhanced features to perform recognition of the requested task scenario to obtain the recognition result of the requested task scenario.
[0053] Furthermore, the feature dimension disassembling module 11 is further used to execute the following method:
[0054] Connect to the rail transit data collection platform, obtain all modal data, and construct a full-modal database; according to the entire operation cycle of rail transit, analyze the task scenarios of each modal data in the full-modal database to determine the scenario-data analysis relationship.
[0055] Furthermore, the feature dimension decomposition module 11 is also used to execute the following method:
[0056] Taking the task scenario as the top event, decompose the top event into basic events based on historical samples to obtain the basic events of the task scenario; decompose the accident factors of the basic events of the task scenario according to historical samples to obtain the accident factors of the basic events and the event impact function relationship; based on the fault tree algorithm, construct a causal relationship tree according to the top event-basic event-accident factor and its event impact function relationship; sort out the data relationships of each task scenario based on the causal relationship tree, establish a scenario data association matrix, and obtain the scenario-data analysis relationship.
[0057] Furthermore, the feature dimension decomposition module 11 is also used to execute the following method:
[0058] The full-modal database at least includes: visual data, sensor data, text data, voice data, and geographic information data.
[0059] Furthermore, the feature dimension decomposition module 11 is also used to execute the following method:
[0060] When the matching modal data is single-modal data, decompose the data feature dimensions of the single-modal data to obtain multi-dimensional features; according to the single-modal data and its corresponding multi-dimensional features, perform visual graph conversion to establish the scenario feature analysis map, where the nodes are task scenarios, single-modal data, and multi-dimensional features, and the connection relationship represents the feature quantization relationship between the nodes.
[0061] Furthermore, the feature dimension decomposition module 11 is also used to execute the following method:
[0062] When the matching modal data is multiple modal data, decompose the data feature dimensions of each modal data respectively to obtain the multi-dimensional features of each modal data; according to the multi-dimensional features of each modal data, analyze the fusion relationship of cross-modal data with the task scenario as the target to establish a cross-modal feature relationship; according to the multi-dimensional features of each modal data and the cross-modal feature relationship, perform visual graph conversion to establish the scenario feature analysis map.
[0063] Furthermore, the feature dimension decomposition module 11 is also used to execute the following method:
[0064] Perform temporal and / or spatial alignment on the cross-modal data; obtain co-occurrence data of multi-dimensional features of the cross-modal data according to the abnormal samples of the task scenario, where the co-occurrence data is the data in which the dimensional features of the cross-modal data appear simultaneously in the abnormal samples; calculate the confidence of the cross-modal dimensional features according to the independent occurrence data of the single-modal data features in the cross-modal data in the abnormal samples and the co-occurrence data; determine the fusion relationship of the cross-modal data according to the confidence of the cross-modal dimensional features.
[0065] Further, the feature dimension disassembling module 11 is further configured to perform the following method:
[0066] Obtain single-abnormal evaluation values of each single-modal data feature from the abnormal samples respectively according to the independent occurrence data; obtain co-occurrence abnormal evaluation values according to the co-occurrence data; calculate the difference coefficient between the sum of the single-abnormal evaluation values corresponding to the single-modal data features in the co-occurrence data and the co-occurrence abnormal evaluation values to obtain the confidence of the cross-modal dimensional features.
[0067] Further, the granularity configuration module 12 is further configured to perform the following method:
[0068] Set granularity enhancement parameters corresponding to the granularity levels of feature addition; perform multi-granularity level enhancement processing on the dimensional features by using the granularity enhancement parameters corresponding to the preset granularity levels according to the scenario feature parsing map; perform sensitivity evaluation on each granularity level enhancement processing by using the recognition result of the request task scenario to obtain feature recognition sensitivity; perform granularity determination on the feature recognition sensitivity based on the recognition target and sensitivity requirement threshold of the request task scenario to determine the configured granularity.
[0069] It should be noted that the above sequence of embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above describes specific embodiments of this specification. The processes depicted in the drawings do not necessarily require the specific order and continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.
[0070] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0071] This specification and the drawings are merely exemplary illustrations of the present application and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications therein.
Claims
1. An intelligent processing method for AI multi-modal data in rail transit, characterized in that, Including: Utilize the scenario - data analysis relationship to match the corresponding modal data of the request task scenario, disassemble the relationship of data feature dimensions of the matched modal data, and establish a scenario feature analysis map; Perform a feature - enhancement granularity sensitivity evaluation on the matched modal data according to the scenario feature analysis map, and configure the granularity according to the evaluation sensitivity; Perform feature enhancement on the matched modal data in the scenario feature analysis map according to the configured granularity, and use the enhanced features to identify the request task scenario for the request task scenario, obtaining the recognition result of the request task scenario; Establishing a scenario feature analysis map includes: When the matched modal data is multiple modal data, disassemble the data feature dimensions of each modal data respectively to obtain multi - dimensional features of each modal data; According to the multi - dimensional features of each modal data, perform a fusion relationship analysis of cross - modal data with the task scenario as the target, and establish a cross - modal feature relationship; According to the multi - dimensional features of each modal data and the cross - modal feature relationship, perform a visual graph conversion to establish the scenario feature analysis map; Performing a fusion relationship analysis of cross - modal data with the task scenario as the target includes: Align the cross - modal data in time and / or space; According to the abnormal samples of the task scenario, obtain the co - occurrence data of the multi - dimensional features of the cross - modal data, where the co - occurrence data is the data where the dimensional features of the cross - modal data appear simultaneously in the abnormal samples; According to the independent occurrence data of the single - modal data features in the cross - modal data in the abnormal samples and the co - occurrence data, calculate the confidence of the cross - modal dimensional features; According to the confidence of the cross - modal dimensional features, determine the fusion relationship of the cross - modal data.
2. The intelligent processing method for rail transit AI multi-modal data according to claim 1, wherein, Utilize the scenario - data analysis relationship to match the corresponding modal data of the request task scenario, previously including: Connect to the rail transit data acquisition platform, obtain all modal data, and construct a full - modal database; According to the entire operation cycle of rail transit, perform a task scenario analysis on each modal data in the full - modal database to determine the scenario - data analysis relationship.
3. The intelligent processing method for rail transit AI multi-modal data according to claim 2, characterized in that Performing a task scenario analysis on each modal data in the full - modal database to determine the scenario - data analysis relationship includes: Taking the task scenario as the top event, perform a basic event disassembly on the top event based on historical samples to obtain the basic events of the task scenario; According to historical samples, perform an accident factor decomposition on the basic events of the task scenario to obtain the accident factors of the basic events and the event impact function relationship; Based on the fault tree algorithm, construct a causal relationship tree according to the top event - basic event - accident factor and its event impact function relationship; Based on the causal relationship tree, sort out the data relationships of each task scenario, establish a scenario data association matrix, and obtain the scenario - data analysis relationship.
4. The intelligent processing method for rail transit AI multi-modal data according to claim 2, wherein The full - modal database includes at least: visual data, sensor data, text data, voice data, and geographic information data.
5. The intelligent processing method for rail transit AI multimodal data according to claim 4, wherein Disassembling the relationship of data feature dimensions of the matched modal data and establishing a scenario feature analysis map includes: When the matched modal data is single - modal data, disassemble the data feature dimensions of the single - modal data to obtain multi - dimensional features; Based on the single-modal data and its corresponding multi-dimensional features, perform visual graph transformation to establish the scene feature analysis map, where the nodes are task scenarios, single-modal data, and multi-dimensional features, and the connection relationship represents the feature quantization relationship between the nodes.
6. The intelligent processing method for rail transit AI multi-modal data according to claim 1, wherein Calculate the confidence of cross-modal dimensional features, including: Obtain the single anomaly evaluation value of each single-modal data feature from the abnormal samples respectively according to the independently occurring data; Obtain the co-occurrence anomaly evaluation value according to the co-occurrence data; Calculate the difference coefficient between the sum of the single anomaly evaluation values corresponding to the single-modal data features in the co-occurrence data and the co-occurrence anomaly evaluation value to obtain the confidence of the cross-modal dimensional features.
7. The intelligent processing method for rail transit AI multimodal data according to claim 1, wherein, Perform feature enhancement granularity sensitivity evaluation on the matching modal data according to the scene feature analysis map, and perform granularity configuration according to the evaluation sensitivity, including: Set the granularity enhancement parameters corresponding to the granularity levels of feature increase; According to the scene feature analysis map, perform multi-granularity level enhancement processing on the dimensional features by using the granularity enhancement parameters corresponding to the preset granularity levels; Use the recognition result of the requested task scenario to evaluate the sensitivity of each granularity level enhancement processing to obtain the feature recognition sensitivity; Based on the recognition target and sensitivity requirement threshold of the requested task scenario, perform granularity determination on the feature recognition sensitivity to determine the configured granularity.
8. An intelligent processing system for AI multi-modal data in rail transit, characterized in that, The system is used to execute the intelligent processing method for rail transit AI multi-modal data according to any one of claims 1-7, including: Feature dimension decomposition module: Use the scene-data analysis relationship to match the corresponding modal data of the requested task scenario, decompose the data feature dimension relationship of the matching modal data, and establish a scene feature analysis map; Granularity configuration module: Perform feature enhancement granularity sensitivity evaluation on the matching modal data according to the scene feature analysis map, and perform granularity configuration according to the evaluation sensitivity; Task scenario recognition module: Perform feature enhancement on the matching modal data in the scene feature analysis map according to the configured granularity, and use the enhanced features to perform request task scenario recognition on the requested task scenario to obtain the recognition result of the requested task scenario.
Citation Information
Patent Citations
Coal mine fire and gas disaster multi-modal knowledge base construction and optimization method
CN119322857A
Rail transit intelligent control system and method based on Internet of Things
CN119682818A