Intelligent processing method and system for rail transit AI multi-modal data

By establishing a scene feature analysis map in the rail transit system and evaluating and configuring feature enhancement particle size sensitivity, the problem of low match between scenes and data and insufficient configuration of feature enhancement particle size in multimodal data processing is solved, and more efficient task scenario recognition and multimodal data processing are achieved.

CN119988895AActive Publication Date: 2025-05-13DALIAN METRO TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510473714.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

In the multimodal data processing process in rail transit systems, the low matching degree of scenes and data and insufficient feature enhancement granularity configuration lead to low recognition efficiency.

Method used

Through the scene-data analysis relationship matching the corresponding modal data of the requested task scenario, the data feature dimension relationship is disassembled, the scene feature analysis map is established, and the feature enhancement particle size sensitivity evaluation and granularity configuration are performed based on the map, and the task scene recognition is finally achieved using enhanced features.

Benefits of technology

It improves the accuracy and efficiency of task scenario recognition and improves the efficiency of multimodal data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988895A_ABST
    Figure CN119988895A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and provides an intelligent processing method and system for rail transit AI multi-modal data. The method comprises the following steps: matching corresponding modal data of a task scene by utilizing a scene-data analysis relationship, performing feature dimension disassembly on the matched data, and establishing a scene feature analysis map; performing granularity sensitivity evaluation on the matched modal data according to the scene feature analysis atlas, and performing granularity configuration according to sensitivity; and performing feature enhancement on the matched modal data in the scene feature analysis atlas according to the configuration granularity, and identifying the request task scene by using the enhanced features to obtain an identification result. According to the method, the technical problem of low recognition efficiency caused by low scene and data matching degree and insufficient feature enhancement granularity configuration in a rail transit multi-modal data processing process is solved, and the recognition efficiency is improved through scene feature analysis atlas, granularity sensitivity evaluation and feature enhancement. The task scene recognition precision is improved; and the multi-modal data processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and specifically to an intelligent processing method and system for rail transit AI multimodal data. Background Art

[0002] With the acceleration of urbanization and the rapid development of rail transit, rail transit systems such as subways and light rails have become an important part of urban public transportation. However, the operation and maintenance of rail transit systems face huge challenges, especially in data management and fault diagnosis. Rail transit systems generate massive amounts of multimodal data every day, including visual data, sensor data, text data, voice data, and geographic information data. These data come from diverse sources and have complex structures. Traditional single-modal data processing methods are difficult to effectively integrate and utilize these multimodal data, resulting in low data processing efficiency, low matching between scenes and data, and insufficient feature enhancement granularity configuration, which in turn affects the accuracy and efficiency of task scene recognition.

[0003] In actual operation, the fault diagnosis, resource allocation and operation management of rail transit systems are highly dependent on the intelligent processing and analysis of multimodal data. However, existing data processing technologies are often unable to process multimodal data efficiently, especially when faced with complex mission scenarios, and it is difficult to accurately extract key features and perform effective scene recognition. This not only increases operating costs, but may also lead to potential safety hazards. Therefore, how to efficiently process multimodal data and improve the accuracy and efficiency of mission scene recognition has become a key technical problem in the development of intelligent rail transit. Summary of the invention

[0004] This application provides an intelligent processing method and system for rail transit AI multimodal data, aiming to solve the technical problems of low recognition efficiency caused by poor matching between scenes and data and insufficient feature enhancement granularity configuration during rail transit multimodal data processing.

[0005] The first aspect disclosed in the present application provides an intelligent processing method for rail transit AI multimodal data, the method comprising: utilizing a scene-data analysis relationship to match the corresponding modal data of a request task scenario, performing data feature dimension relationship decomposition on the matching modal data, and establishing a scene feature analysis map; performing feature enhancement granularity sensitivity evaluation on the matching modal data according to the scene feature analysis map, and performing granularity configuration according to the evaluated sensitivity; performing feature enhancement on the matching modal data in the scene feature analysis map according to the configured granularity, performing request task scenario recognition on the request task scenario using the enhanced features, and obtaining a recognition result of the request task scenario.

[0006] Another aspect disclosed in the present application provides an intelligent processing system for rail transit AI multimodal data, the system comprising: a feature dimension decomposition module: utilizing the scene-data analysis relationship to match the corresponding modal data of the requested task scene, performing data feature dimension relationship decomposition on the matched modal data, and establishing a scene feature analysis map; a granularity configuration module: performing feature enhancement granularity sensitivity evaluation on the matched modal data according to the scene feature analysis map, and performing granularity configuration according to the evaluated sensitivity; a task scene recognition module: performing feature enhancement on the matched modal data in the scene feature analysis map according to the configured granularity, performing requested task scene recognition on the requested task scene using the enhanced features, and obtaining a recognition result of the requested task scene.

[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages: The above-mentioned intelligent processing method of rail transit AI multimodal data utilizes the scene-data analysis relationship to match the corresponding modal data of the request task scene, decomposes the data feature dimension relationship of the matching modal data, and establishes a scene feature analysis map; then, according to the scene feature analysis map, the matching modal data is evaluated for feature enhancement granularity sensitivity, and the granularity is configured according to the evaluation sensitivity; finally, according to the configuration granularity, the matching modal data in the scene feature analysis map is feature enhanced, and the request task scene is identified for the request task scene using the enhanced features to obtain the identification result of the request task scene, thereby achieving the technical effect of improving the task scene identification accuracy and improving the multimodal data processing efficiency through scene feature analysis maps, granularity sensitivity evaluation and feature enhancement.

[0008] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0010] Figure 1 The figure is a flow chart of an intelligent processing method of rail transit AI multimodal data in one embodiment.

[0011] Figure 2 This is an architecture diagram of an intelligent processing system for rail transit AI multimodal data in one embodiment.

[0012] Explanation of the reference numerals: feature dimension decomposition module 11, granularity configuration module 12, task scenario recognition module 13. DETAILED DESCRIPTION

[0013] The embodiments of the present application provide an intelligent processing method and system for rail transit AI multimodal data to solve the technical problems of low recognition efficiency caused by low scene and data matching and insufficient feature enhancement granularity configuration during rail transit multimodal data processing.

[0014] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0015] It should be noted that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules that are not explicitly listed or inherent to these processes, methods, products or devices.

[0016] Embodiment 1, as Figure 1 As shown, the present application provides an intelligent processing method for rail transit AI multimodal data, the method comprising: The corresponding modal data of the requested task scenario is matched by using the scene-data analysis relationship, the data feature dimension relationship of the matched modal data is decomposed, and a scene feature analysis graph is established.

[0017] In the embodiment of the present application, firstly, it is necessary to match the multimodal data related to the task scene according to the needs of the task scene, using the relationship between the scene and data analysis. These multimodal data include visual data (such as images captured by the camera), sensor data (such as environmental data such as temperature and humidity), text data (such as operation logs), etc. Subsequently, the matched multimodal data is disassembled into data feature dimensions. Specifically, each modal data is decomposed into features of multiple dimensions. For example, for visual data, it can be disassembled into features such as color and shape; for sensor data, it can be disassembled into time series features, sampling frequency features, etc. Through this disassembly, the inherent structure and information of each modal data can be understood more carefully. After completing the feature dimension disassembly, based on these feature dimension relationships, a scene feature analysis map is established. This map is a visual graphical structure, in which nodes represent task scenes, modal data and their corresponding multidimensional features, and the connection relationship represents the feature quantification relationship between these nodes. For example, a task scene may be related to multiple modal data, and some feature dimensions of these modal data may jointly affect the recognition results of the scene. By constructing such a graph, the complex relationship between task scenarios and multimodal data can be clearly displayed, providing basic support for subsequent feature enhancement and scene recognition.

[0018] For example, assume that the task scenario is track health monitoring. In this scenario, multiple modal data are needed to comprehensively evaluate the health status of the track, including visual data, texture features, vibration data, etc., where visual data refers to images of the track surface captured by a camera. These images have spatial resolution characteristics and texture features, which can help identify physical damage such as cracks and wear on the track surface; vibration data refers to vibration data collected by sensors when a train passes. These data have frequency characteristics and can help identify looseness or deformation of the track structure. In this example, the task scenario of track health monitoring is related to visual data and sensor data. The spatial resolution and texture features of visual data, as well as the frequency characteristics of vibration data of sensor data, jointly affect the identification results of the track health status. By analyzing the feature dimensions of these multimodal data, the health status of the track can be more comprehensively evaluated, thereby improving the accuracy and efficiency of monitoring.

[0019] Further, the present application provides a method for matching the corresponding modal data of the request task scenario using the scenario-data analysis relationship, which includes: Connect to the rail transit data acquisition platform, obtain all modal data, and build a full modal database; according to the full operation cycle of rail transit, perform task scenario analysis on each modal data in the full modal database to determine the scenario-data analysis relationship.

[0020] Preferably, before matching the corresponding modal data of the requested task scenario, it is first necessary to connect to the data collection platform of rail transit to obtain modal data from different sources, including visual data (such as track images, platform cameras), sensor data (such as vibration, temperature, current), text data (such as social media public opinion, maintenance records), voice data (such as alarm broadcasts, fault audio), geographic information data (such as track topology, train location information), etc. Through real-time or regular collection of these data, a full-modal database can be constructed to provide a basis for subsequent task analysis and processing. These data will be cleaned, format converted, and timestamp aligned when stored to ensure data quality and consistency. Subsequently, according to the full operation cycle of rail transit, combined with actual business needs and historical samples, basic events of multiple task scenarios are defined, such as track health monitoring (detection of track cracks, deformation, and foreign body intrusion), passenger flow prediction and scheduling optimization (predicting passenger flow density during peak hours and optimizing train departure intervals), emergency response (rapidly locating the event location and generating emergency plans), equipment fault diagnosis (real-time monitoring of the status of equipment such as train bearings and signal systems), etc. Afterwards, based on these basic events, the scene-data analysis relationship is established through accident factor decomposition and data relationship combing. This scene-data analysis relationship records the data and impact involved in each basic event, providing basic support for subsequent data processing and task identification. For example, in the track health monitoring scenario, visual data (such as track images), sensor data (such as vibration data) and geographic information data (such as track topology) will be associated to detect track cracks or deformation; in the passenger flow prediction and scheduling optimization scenario, visual data (such as camera counting data), sensor data (such as gate pass data) and text data (such as social media public opinion) will be combined to predict passenger flow during peak hours, and train departure intervals will be optimized based on these data; in the emergency response scenario, visual data (such as smoke images) and sensor data (such as temperature and smoke concentration) will be associated and analyzed to quickly locate the location of the event and generate corresponding emergency plans; in the equipment fault diagnosis task, sensor data (such as vibration and current signals) and text data (such as maintenance records) will be associated to monitor the status of train equipment. Through the above process, it is possible to extract multimodal data related to each task scenario from the omnimodal database according to the operational needs of rail transit, and determine the relationship between the scenario and data analysis, providing basic support for subsequent data processing and task identification.

[0021] Furthermore, the present application provides a task scenario analysis of each modality data in the full modality database to determine the scenario-data analysis relationship, including: Taking the task scenario as the top event, the top event is decomposed into basic events based on historical samples to obtain the basic events of the task scenario; the basic events of the task scenario are decomposed into accident factors according to historical samples to obtain the accident factors of the basic events and the event impact function relationship; based on the accident tree algorithm, a causal relationship tree is constructed according to the top event-basic event-accident factor and its event impact function relationship; based on the causal relationship tree, the data relationship of each task scenario is sorted out, and a scenario data association matrix is ​​established to obtain the scenario-data analysis relationship.

[0022] Optionally, first, according to business needs and the entire operation cycle of rail transit, multiple task scenarios are defined, and these task scenarios are used as top events, which means starting the analysis from a high-level task scenario, such as equipment fault diagnosis or track health monitoring. These task scenarios represent important operational goals or problems in the entire rail transit system. Subsequently, based on historical sample data, the top events are decomposed into basic events. This process decomposes the task scenario into multiple more basic events. For example, in the equipment fault diagnosis task, basic events such as sensor abnormalities or mechanical damage are involved. By analyzing the event identifiers of each group of data in the historical sample data, multiple basic events of each task scenario can be identified, such as track surface crack events and track structure loosening events in the track health monitoring scenario. Afterwards, these basic events are decomposed into accident factors based on historical sample data. Accident factors refer to various factors that may cause a basic event to occur, which may include environmental factors, operating errors or equipment failures. By extracting the data corresponding to each basic event, and using the Pearson correlation coefficient, Spearman rank correlation coefficient, etc. to quantify the correlation between each factor in the data involved in each basic event and the basic event, for example, when the basic event is the size of the track crack, obtain the data related to the event in the historical sample data, such as temperature, humidity, wheel-rail friction, material aging, etc., and quantify the correlation between each factor and the size of the track crack by using the Pearson correlation coefficient, Spearman rank correlation coefficient, etc. Then, check the correlation of each factor with the correlation lower limit value, eliminate the factors less than the correlation lower limit value, and use the remaining factors as accident factors, and then use these accident factors as indexes to analyze the extracted data and calculate the probability of each accident factor each time a basic event occurs (the number of occurrences of the accident factor is calculated by the ratio of the number of occurrences of the basic event). By comparing the probability of each accident factor with the lower limit of probability, the logical relationship between each accident factor and the basic event is determined (greater than means AND, less than means OR). Finally, each accident factor is associated with the basic event through correlation and logical relationship to form the event impact function relationship of each basic event. The event impact function relationship represents the degree of influence of each accident factor on the occurrence of the basic event.

[0023] After obtaining the accident factors and event impact function relationships of each basic event, the accident tree algorithm is applied to build a causal relationship tree. The accident tree algorithm is a logic-based tool for analyzing the relationship between events and factors. The initial causal relationship tree is constructed by taking the top event as the top node of the tree, i.e., the root node, connecting the basic event as the intermediate node to the top event, and connecting the accident factor as the leaf node to the corresponding basic event. Subsequently, the logical relationship between the nodes in the initial causal relationship tree is defined according to the logical relationship in the event impact function relationship to indicate how the accident factors jointly cause the basic events, and how the basic events jointly cause the top events. In general, the relationship between the basic event and the top event is or, for example, basic event 1 (track surface cracks) is caused by accident factor 1 (wheel-rail friction) or accident factor 2 (material aging); basic event 2 (track structure loosening) is caused by accident factor 1 (bolt loosening) and accident factor 2 (foundation settlement); the top event (track health monitoring) is caused by basic event 1 (track surface cracks) or basic event 2 (track structure loosening). After the logical relationship configuration is completed, a causal relationship tree is obtained. Through this causal relationship tree, the dependencies and mutual influences between different events and factors can be clearly understood. Afterwards, based on the causal relationship tree, the data relationships in each task scenario are sorted out to establish a scenario data association matrix. The data association matrix is ​​a two-dimensional matrix used to describe the relationship between different modal data and events and factors in the task scenario. In this matrix, rows represent different data modalities (such as sensor data, visual data, text data, etc.), columns represent different events or accident factors, and the values ​​in the matrix represent the correlation and logical relationship between data and events or factors. This matrix helps to identify which data is most important for events or factors in a specific task scenario, thereby further optimizing the use of data and the analysis of task scenarios. Through the above process, a complete scenario-data analysis relationship is finally obtained, and the analysis relationship between each task scenario and multimodal data is clarified, providing structured support for subsequent data processing and task identification.

[0024] Furthermore, the present application provides that the omnimodal database includes at least: visual data, sensor data, text data, voice data, and geographic information data.

[0025] Optionally, a full-modal database refers to a database containing multiple types of data. In the rail transit system, these data come from different monitoring devices and information systems for comprehensive analysis and decision-making. Specifically, the full-modal database includes at least visual data, sensor data, text data, voice data, and geographic information data. Among them, the visual data mainly comes from cameras, usually including track images, car monitoring videos, platform monitoring screens, etc. These data are used to detect track conditions, identify passenger behavior, monitor equipment operation status, etc. For example, image analysis can be used to detect whether there are cracks or other obstacles on the track; sensor data comes from various sensors installed on equipment and tracks, such as temperature sensors, vibration sensors, pressure sensors, etc. These data can provide real-time physical status monitoring, which is used to evaluate equipment health, track stability, smoothness of train operation, etc. For example, sensors can help monitor the vibration of the train or temperature changes of the track; text data comes from different documents, reports, and transmissions. Sensor logs, maintenance records, operation manuals, social media, etc. These data are usually structured or unstructured, and contain information such as system operation status, fault reports, and operation instructions. By analyzing text data, potential fault information, maintenance suggestions, or user feedback can be extracted; voice data includes voice instructions from station broadcasts, passengers or staff, alarm signals, communication records, etc. Through voice recognition and analysis, these data can help monitor and analyze passenger behavior, station safety, equipment failures, or emergency events, etc. For example, fault alarm audio or emergency broadcasts can be used to determine whether an emergency has occurred; geographic information data includes geographic location information, station location, train location, road topology, etc. of the rail transit network. This type of data is used for map display, route planning, traffic flow monitoring, fault location, etc. Through the geographic information system (GIS), accurate positioning and efficient management of the rail transit system can be achieved. These different types of data form a multimodal data system that can provide a comprehensive view of all aspects of rail transit, thereby helping real-time monitoring, fault diagnosis, resource optimization, and intelligent decision-making. By integrating these modal data, the intelligent operation and maintenance of the system can be better realized, operational efficiency can be improved, and potential risks can be reduced.

[0026] Furthermore, the present application provides a method for decomposing the data feature dimension relationship of the matching modal data and establishing a scene feature analysis graph, including: When the matching modal data is single modal data, the data feature dimension of the single modal data is decomposed to obtain multi-dimensional features; based on the single modal data and its corresponding multi-dimensional features, a visual graphic conversion is performed to establish the scene feature analysis graph, wherein the nodes are task scenes, single modal data, and multi-dimensional features, and the connection relationship represents the feature quantification relationship between the nodes.

[0027] Optionally, after matching the modal data of the current task scene through the scene-data analysis relationship, it will be determined whether the matched modal data is single modal data, that is, whether the current modal data comes from a single data source and contains only one type of information. If the matched modal data is single modal data, the single modal data will be feature-decomposed and converted into multi-dimensional features. These features are used to quantify different attributes of the data and provide support for subsequent analysis and visualization. For example, for visual data, feature dimensions may include resolution, pixels, texture, edge strength, color, etc.; for vibration data, feature dimensions may include sampling frequency, time domain statistics, frequency domain energy, multi-axis correlation, envelope features, etc.; for text data, feature dimensions include word frequency distribution, entity density, syntactic complexity, sentiment polarity, topic distribution, etc.; and for location data, its feature dimensions may be coordinate accuracy, timestamp synchronization, spatial topology, elevation features, motion trajectory, etc. After the feature dimension decomposition of the single modal data is completed, these features are converted into visual graphics, that is, the relationship between the features is represented in a graphical way, which is convenient for analysis and understanding. The scene feature analysis map generated by this process will help to identify and analyze the subsequent task scenes. In the scene feature analysis map, three types of nodes are defined, namely, task scene nodes, single modal data nodes and multi-dimensional feature nodes. Connection relationships are also defined between nodes. The connection relationship represents the feature quantitative relationship between nodes, that is, the correlation between different data features and task scenes. For example, connecting the task scene node with the single modal data node represents the application of the modal data in the task scene; connecting the single modal data node with the multi-dimensional feature node represents the multi-dimensional features of the modal data; connecting the quantitative relationship between the multi-dimensional feature nodes represents the mutual influence and dependency between the features. These nodes and connection relationships together constitute the scene feature analysis map. Through the above process, the data feature dimension decomposition of the single modal data can be carried out, and the scene feature analysis map can be established through visual graphic conversion to clarify the relationship between the task scene, single modal data and multi-dimensional features, providing basic support for subsequent feature enhancement and task scene recognition.

[0028] Furthermore, the present application provides a method for establishing a scene feature analysis graph, including: When the matching modal data is multiple modal data, the data feature dimensions of each modal data are decomposed respectively to obtain the multi-dimensional features of each modal data; according to the multi-dimensional features of each modal data, the fusion relationship of cross-modal data is analyzed with the task scenario as the target, and a cross-modal feature relationship is established; according to the multi-dimensional features of each modal data and the cross-modal feature relationship, a visual graphic conversion is performed to establish the scene feature analysis map.

[0029] Optionally, when the matching modal data is multiple modal data, the aforementioned data feature dimension decomposition will be performed on each modal data first. After the data feature decomposition of each modality is completed, the data of each modality will be converted into corresponding multi-dimensional features. Subsequently, with the task scenario as the target, the fusion relationship of cross-modal data is analyzed. The purpose of cross-modal fusion is to combine data from different modalities (such as visual data and vibration data) to more fully understand the task scenario. To achieve this, different modal data will be aligned in time and space. For example, the timestamps of visual data, vibration data, and position data will be synchronized to ensure the corresponding relationship between the data at the same time point and spatial position. Then, according to the abnormal samples of the task scene, the co-occurrence of the multi-dimensional features of the cross-modal data in the abnormal samples is analyzed. For example, in the track crack abnormal samples, the edge features of the visual data (edge ​​strength is 0.7) and the frequency domain energy of the vibration data (low-frequency energy accounts for 60%) appear at the same time; in the track structure loose abnormal samples, the envelope features of the vibration data (the envelope peak appears at 50Hz) and the spatial topology of the position data (the track topology structure is a straight line segment) appear at the same time. After that, the confidence of the cross-modal dimensional features is calculated, and then the fusion relationship between the cross-modal data is defined according to the confidence calculation results, and the defined fusion relationship is stored to obtain the cross-modal feature relationship as the basis for subsequent task scene recognition. Then, the multi-dimensional features are fused according to the cross-modal feature relationship, and then the visual graphics are converted according to the fused characteristics to establish a scene feature analysis map. This map includes task scene nodes, single-modal data nodes, multi-dimensional feature nodes, fusion feature nodes, and the connection relationship between nodes. Through the above process, it is possible to decompose the feature dimensions of multimodal data, analyze the fusion relationship of cross-modal data, and establish a scene feature analysis map through visual graphic conversion to clarify the relationship between task scenarios, multimodal data and their feature dimensions, providing basic support for subsequent feature enhancement and task scenario recognition.

[0030] Furthermore, the present application provides a method for analyzing the fusion relationship of cross-modal data with a task scenario as the target, including: The cross-modal data is aligned in time and / or space; based on the abnormal samples of the task scenario, the co-occurrence data of the multi-dimensional features of the cross-modal data is obtained, and the co-occurrence data is the data in which the dimensional features of the cross-modal data appear simultaneously in the abnormal samples; based on the independent occurrence data of the single-modal data features in the cross-modal data in the abnormal samples and the co-occurrence data, the confidence of the cross-modal dimensional features is calculated; based on the confidence of the cross-modal dimensional features, the fusion relationship of the cross-modal data is determined.

[0031] Optionally, when processing data of multiple modalities, it is necessary to ensure that the data are aligned in time and space. For example, visual data (such as track images) and sensor data (such as vibration signals) may be collected at different times and spatial locations, so synchronization and alignment are required. In this process, the different modal data will be synchronized according to the timestamp to ensure the alignment of the timestamps of the different modal data. If geographic location data (such as sensor location and track image location) is involved, it is also necessary to ensure the consistency of different data sources in space. This is to map data from different sources into the same coordinate system through coordinate transformation, thereby ensuring the consistency of data in physical space. Once the cross-modal data is aligned in time and / or space, the abnormal samples of the current mission scenario are retrieved from the abnormal sample library, such as track cracks, loose track structure, etc., and then the abnormal samples are compared with the multi-dimensional features of the cross-modal data for keywords to identify the co-occurrence data between the abnormal samples and the multi-dimensional features. This co-occurrence data indicates that the multi-dimensional features of different modal data appear in the abnormal samples at the same time. For example, a frequency domain feature in the vibration signal may appear in the abnormal sample at the same time as a feature in the visual data (such as the crack edge of the track image). Subsequently, based on the co-occurrence data obtained above, the confidence of the cross-modal dimension features is calculated. Confidence is an important indicator to measure the strength and reliability of the relationship between different modal data features. In this process, it is analyzed whether the data features of each modality appear independently in the abnormal samples. For example, if the low-frequency energy of the vibration signal appears repeatedly in multiple abnormal samples, but is unrelated to other modal features (such as visual data), it can be considered that the independence of this feature is strong. Then, it is calculated based on the co-occurrence data. The co-occurrence data indicates the co-occurrence of multiple modal features in the same abnormal sample. By calculating the frequency of co-occurrence data, etc., the understanding of the cross-modal feature relationship is further enhanced. Afterwards, the confidence of the cross-modal dimension features is calculated by combining the independent occurrence data and the co-occurrence data. This confidence value reflects the strength of the association between cross-modal features and helps to determine whether there is an effective fusion relationship. Finally, according to the confidence of the calculated cross-modal dimension features, the fusion relationship between different modal data can be further determined. If the confidence is high, it means that different modal data have a strong correlation in the task scenario. These data can be fused together to more accurately analyze the task scenario and make decisions. For example, in equipment fault diagnosis, if the visual data and vibration data have a high confidence in abnormal conditions, it means that there is a strong correlation between these features and they are suitable for closer fusion. At this time, the data of the two can be fused to form a more accurate fault diagnosis result. Through the above process, the system can complete the alignment of cross-modal data, co-occurrence data analysis, confidence calculation and fusion relationship determination, provide basic support for the subsequent construction of scene feature analysis maps and task scenario identification, and improve the accuracy and efficiency of task scenario analysis.

[0032] Furthermore, the present application provides a method for calculating the confidence of cross-modal dimensional features, including: According to the independent occurrence data, the single anomaly evaluation value of each unimodal data feature is obtained from the abnormal samples respectively; according to the co-occurrence data, the co-occurrence anomaly evaluation value is obtained; according to the sum of the single anomaly evaluation values ​​corresponding to the unimodal data features in the co-occurrence data, the difference coefficient is calculated with the co-occurrence anomaly evaluation value to obtain the confidence of the cross-modal dimensional feature.

[0033] Optionally, in the process of cross-modal data processing, the abnormal evaluation of single-modal data features and the abnormal evaluation of co-occurrence data are very important steps for calculating the confidence of cross-modal dimension features. Specifically, the absolute difference between the independent data in each single-modal data feature and the mean of the normal sample of the data is first calculated, and then the calculation result is compared with the standard deviation of the data to obtain the single abnormal evaluation value of each single-modal data feature. At the same time, in the abnormal sample, the number of times the different modal data features appear abnormally at the same time is counted, and the statistical co-occurrence number is compared with the total number of feature occurrences to obtain the co-occurrence abnormal evaluation value of the co-occurrence data. Subsequently, the single abnormal evaluation values ​​of all independent data are accumulated, and the accumulated result is compared with the co-occurrence abnormal evaluation value of the co-occurrence data to determine the difference coefficient. The difference coefficient helps measure the difference between the individual abnormality of each modal feature and its abnormal intensity when it co-occurs in multimodal data. Finally, by using the difference coefficient as the confidence of the cross-modal dimension feature, it provides a basis for the subsequent determination of the cross-modal data fusion relationship.

[0034] The feature enhancement granularity sensitivity of the matching modal data is evaluated according to the scene feature analysis map, and the granularity is configured according to the evaluated sensitivity.

[0035] In one embodiment, when evaluating the sensitivity of feature enhancement granularity, it is first necessary to analyze the impact of the features of different modal data on the task scene at different granularities. Granularity refers to the level and resolution of data features in details. Fine-grained features usually contain more detailed information and can provide higher accuracy, while coarse-grained features may be more concise and cover a wider range, but may not be accurate enough in details. In this process, multiple granularity levels are defined, each level corresponds to different feature enhancement parameters, and then according to the scene feature analysis map, the feature enhancement parameters corresponding to the preset granularity level are used to perform multi-granularity level enhancement processing on the dimensional features. Subsequently, the recognition result of the requested task scene is used to perform sensitivity evaluation on the enhancement processing of each granularity level, obtain the feature recognition sensitivity, and judge the feature recognition sensitivity according to the sensitivity requirement threshold, and select the granularity level that meets the sensitivity requirement threshold as the configuration granularity to ensure that the task scene can obtain appropriate information. Through the above process, it is possible to perform feature enhancement granularity sensitivity evaluation on the matching modal data based on the scene feature analysis map, and perform granularity configuration according to the evaluation sensitivity, thereby improving the accuracy and efficiency of task scene recognition.

[0036] Furthermore, the present application provides a method for evaluating the sensitivity of feature enhancement granularity of matching modal data according to the scene feature analysis graph, and configuring the granularity according to the evaluation sensitivity, including: Set the granularity enhancement parameters corresponding to the granularity level of feature addition; perform multi-granularity enhancement processing on the dimensional features using the granularity enhancement parameters corresponding to the preset granularity level according to the scene feature analysis map; perform sensitivity evaluation on each granularity level enhancement processing using the recognition result of the requested task scene to obtain the feature recognition sensitivity; perform granularity judgment on the feature recognition sensitivity based on the recognition target and sensitivity requirement threshold of the requested task scene to determine the configuration granularity.

[0037] Optionally, according to historical data or domain knowledge, multiple granularity levels are preset, and each level corresponds to different feature enhancement parameters. For example, low-granularity enhancement, global distribution adjustment, and the feature enhancement parameter is 0.1; medium-granularity enhancement: segmented processing, and the feature enhancement parameter is 0.5; high-granularity enhancement: local feature repair, and the feature enhancement parameter is 0.9. Subsequently, according to the scene feature parsing graph, the dimension features are enhanced at multiple granularity levels using the granularity enhancement parameters corresponding to the preset granularity levels. For example, in the track health monitoring scenario, high-granularity enhancement will be selected to sharpen the local crack area (such as Vision Transformer attention focusing); in the passenger flow prediction and scheduling optimization scenario, low-granularity enhancement will be selected to simulate the global illumination (generate platform images at different time periods). Then, using the recognition result of the request task scenario, the sensitivity evaluation of each granularity level enhancement process is carried out to obtain the feature recognition sensitivity S. This process is obtained by calculating the proportion of correctly recognized samples in the total samples (i.e., the accuracy rate). Then, the sensitivity is compared with the scene threshold, and the scene threshold is determined based on the sensitivity requirement threshold θ, such as high-sensitivity scenario (S > θ), medium-sensitivity scenario (θ / 2 < S ≤ θ), low-sensitivity scenario (S ≤ θ / 2), and then the corresponding configured granularity is matched according to the comparison result. Through the above process, the system can perform feature enhancement granularity sensitivity evaluation on the matched modal data based on the scene feature parsing graph and perform granularity configuration according to the evaluated sensitivity, thereby improving the accuracy and efficiency of task scenario recognition.

[0038] Perform feature enhancement on the matched modal data in the scene feature parsing graph according to the configured granularity, and use the enhanced features to perform request task scenario recognition on the request task scenario to obtain the recognition result of the request task scenario.

[0039] In one embodiment, according to the configured granularity, feature enhancement processing will be performed on the matched modal data in the scene feature parsing graph. The process of feature enhancement is to adjust the detail level of the features according to different granularity requirements and improve the expression ability of the data. Specifically, according to the feature requirements of high, medium, and low granularity, the corresponding data features will be enhanced so that they can provide richer information in task scenario recognition. For example, local crack area sharpening is performed through Vision Transformer attention focusing, and high-frequency noise is processed according to the time window to perform segmented noise reduction on vibration signals. These enhanced features are input into the task scenario recognition model to identify a specific task scenario. The task scenario recognition model analyzes the enhanced data features and outputs the recognition result of the task scenario, which indicates whether the task scenario is successfully recognized and the accuracy of the recognition. In this way, effective task scenario recognition can be performed based on task requirements and feature granularity configuration, improving the accuracy and reliability of recognition.

[0040] The task scene recognition model is built based on a convolutional neural network, and the training data includes sample modal features and sample task scenes. During the training process, iterative training is performed through steps such as forward propagation, loss calculation, back propagation, and parameter optimization until the maximum number of iterations is met or the model converges. After the training is completed, the model will be verified. If the accuracy meets the expected expectations, the current task scene recognition model will be output. Otherwise, hyperparameters such as the learning rate and the number of training batches will be adjusted to further improve the recognition effect of the model.

[0041] In summary, the embodiments of the present application have at least the following technical effects: The embodiment of the present application utilizes the scene-data analysis relationship to match the corresponding modal data of the request task scene, decomposes the data feature dimension relationship of the matching modal data, and establishes a scene feature analysis map; performs feature enhancement granularity sensitivity evaluation on the matching modal data according to the scene feature analysis map, and performs granularity configuration according to the evaluation sensitivity; performs feature enhancement on the matching modal data in the scene feature analysis map according to the configured granularity, and utilizes the enhanced features to perform request task scene recognition on the request task scene to obtain the recognition result of the request task scene. These technical effects jointly solve the technical problems of low scene and data matching and insufficient feature enhancement granularity configuration in the process of multimodal data processing of rail transit, resulting in low recognition efficiency, and achieve the technical effect of improving the task scene recognition accuracy and improving the multimodal data processing efficiency through scene feature analysis maps, granularity sensitivity evaluation and feature enhancement.

[0042] Embodiment 2 is based on the same inventive concept as the intelligent processing method of rail transit AI multimodal data in the aforementioned embodiment. Figure 2 As shown, the present application provides an intelligent processing system for rail transit AI multimodal data, and the system includes: a feature dimension decomposition module 11: using the scene-data analysis relationship to match the corresponding modal data of the request task scene, performing data feature dimension relationship decomposition on the matching modal data, and establishing a scene feature analysis map; a granularity configuration module 12: performing feature enhancement granularity sensitivity evaluation on the matching modal data according to the scene feature analysis map, and performing granularity configuration according to the evaluated sensitivity; a task scene recognition module 13: performing feature enhancement on the matching modal data in the scene feature analysis map according to the configured granularity, performing request task scene recognition on the request task scene using the enhanced features, and obtaining a recognition result of the request task scene.

[0043] Furthermore, the feature dimension decomposition module 11 is also used to execute the following method: Connect to the rail transit data acquisition platform, obtain all modal data, and build a full modal database; according to the full operation cycle of rail transit, perform task scenario analysis on each modal data in the full modal database to determine the scenario-data analysis relationship.

[0044] Furthermore, the feature dimension decomposition module 11 is also used to execute the following method: Taking the task scenario as the top event, the top event is decomposed into basic events based on historical samples to obtain the basic events of the task scenario; the basic events of the task scenario are decomposed into accident factors according to historical samples to obtain the accident factors of the basic events and the event impact function relationship; based on the accident tree algorithm, a causal relationship tree is constructed according to the top event-basic event-accident factor and its event impact function relationship; based on the causal relationship tree, the data relationship of each task scenario is sorted out, and a scenario data association matrix is ​​established to obtain the scenario-data analysis relationship.

[0045] Furthermore, the feature dimension decomposition module 11 is also used to execute the following method: The omnimodal database includes at least: visual data, sensor data, text data, voice data, and geographic information data.

[0046] Furthermore, the feature dimension decomposition module 11 is also used to execute the following method: When the matching modal data is single modal data, the data feature dimension of the single modal data is decomposed to obtain multi-dimensional features; based on the single modal data and its corresponding multi-dimensional features, a visual graphic conversion is performed to establish the scene feature analysis graph, wherein the nodes are task scenes, single modal data, and multi-dimensional features, and the connection relationship represents the feature quantification relationship between the nodes.

[0047] Furthermore, the feature dimension decomposition module 11 is also used to execute the following method: When the matching modal data is multiple modal data, the data feature dimensions of each modal data are decomposed respectively to obtain the multi-dimensional features of each modal data; according to the multi-dimensional features of each modal data, the fusion relationship of cross-modal data is analyzed with the task scenario as the target, and a cross-modal feature relationship is established; according to the multi-dimensional features of each modal data and the cross-modal feature relationship, a visual graphic conversion is performed to establish the scene feature analysis map.

[0048] Furthermore, the feature dimension decomposition module 11 is also used to execute the following method: The cross-modal data is aligned in time and / or space; based on the abnormal samples of the task scenario, the co-occurrence data of the multi-dimensional features of the cross-modal data is obtained, and the co-occurrence data is the data in which the dimensional features of the cross-modal data appear simultaneously in the abnormal samples; based on the independent occurrence data of the single-modal data features in the cross-modal data in the abnormal samples and the co-occurrence data, the confidence of the cross-modal dimensional features is calculated; based on the confidence of the cross-modal dimensional features, the fusion relationship of the cross-modal data is determined.

[0049] Furthermore, the feature dimension decomposition module 11 is also used to execute the following method: According to the independent occurrence data, the single anomaly evaluation value of each unimodal data feature is obtained from the abnormal samples respectively; according to the co-occurrence data, the co-occurrence anomaly evaluation value is obtained; according to the sum of the single anomaly evaluation values ​​corresponding to the unimodal data features in the co-occurrence data, the difference coefficient is calculated with the co-occurrence anomaly evaluation value to obtain the confidence of the cross-modal dimensional feature.

[0050] Furthermore, the granularity configuration module 12 is also used to execute the following method: Set the granularity enhancement parameters corresponding to the granularity level of feature addition; perform multi-granularity enhancement processing on the dimensional features using the granularity enhancement parameters corresponding to the preset granularity level according to the scene feature analysis map; perform sensitivity evaluation on each granularity level enhancement processing using the recognition result of the requested task scene to obtain the feature recognition sensitivity; perform granularity judgment on the feature recognition sensitivity based on the recognition target and sensitivity requirement threshold of the requested task scene to determine the configuration granularity.

[0051] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description and does not represent the advantages and disadvantages of the embodiments. And the above-mentioned specific embodiments of this specification are described. The processes depicted in the accompanying drawings do not necessarily require the specific order and continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0052] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

[0053] This specification and drawings are merely exemplary illustrations of the present application and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, a person skilled in the art may make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalents, the present application intends to include these modifications and variations.

Claims

1. An intelligent processing method for rail transit AI multimodal data, characterized in that: include: Use the scene-data analysis relationship to match the corresponding modal data of the request task scene, decompose the data feature dimension relationship of the matching modal data, and establish a scene feature analysis map; Performing feature enhancement granularity sensitivity evaluation on the matching modality data according to the scene feature analysis graph, and performing granularity configuration according to the evaluated sensitivity; According to the configuration granularity, feature enhancement is performed on the matching modal data in the scene feature analysis map, and the request task scene is identified by using the enhanced features to obtain an identification result of the request task scene.

2. The intelligent processing method of rail transit AI multimodal data according to claim 1 is characterized in that: Use the scene-data analysis relationship to match the corresponding modal data of the request task scene, which previously included: Connect to the rail transit data collection platform to obtain all modal data and build a full modal database; According to the full operation cycle of rail transit, each modal data in the full modal database is analyzed for a task scenario to determine a scenario-data analysis relationship.

3. The intelligent processing method of rail transit AI multimodal data according to claim 2 is characterized in that: Performing task scenario analysis on each modality data in the full modality database to determine the scenario-data analysis relationship includes: Taking the task scenario as the top event, decomposing the top event into basic events based on historical samples to obtain the basic events of the task scenario; Decomposing the basic events of the mission scenario by accident factors according to historical samples, and obtaining the accident factors of the basic events and the event impact function relationship; Based on the accident tree algorithm, a causal relationship tree is constructed according to the relationship between top events, basic events, accident factors and their event impact functions; Based on the causal relationship tree, the data relationship of each task scenario is sorted out, a scenario data association matrix is ​​established, and the scenario-data analysis relationship is obtained.

4. The intelligent processing method of rail transit AI multimodal data according to claim 2 is characterized in that: The omnimodal database includes at least: visual data, sensor data, text data, voice data, and geographic information data.

5. The intelligent processing method of rail transit AI multimodal data according to claim 4 is characterized in that: Decompose the data feature dimension relationship of the matching modal data and establish a scene feature analysis map, including: When the matching modal data is single modal data, the single modal data is decomposed into data feature dimensions to obtain multi-dimensional features; According to the unimodal data and its corresponding multidimensional features, a visual graphic conversion is performed to establish the scene feature analysis graph, wherein the nodes are task scenes, unimodal data, and multidimensional features, and the connection relationship represents the feature quantification relationship between the nodes.

6. The intelligent processing method for rail transit AI multimodal data according to claim 5 is characterized in that: Establish scene feature analysis graph, including: When the matching modal data is multiple modal data, data feature dimension decomposition is performed on each modal data to obtain multi-dimensional features of each modal data; According to the multi-dimensional characteristics of each modal data, the fusion relationship of cross-modal data is analyzed with the task scenario as the target, and the cross-modal feature relationship is established; According to the multi-dimensional features of each modal data and the cross-modal feature relationship, a visual graphic conversion is performed to establish the scene feature analysis map.

7. The intelligent processing method of rail transit AI multimodal data according to claim 6 is characterized in that: The fusion relationship analysis of cross-modal data is carried out based on the task scenario, including: Temporally and / or spatially align cross-modal data; According to the abnormal samples of the task scenario, co-occurrence data of the multi-dimensional features of the cross-modal data are obtained, where the co-occurrence data is data in which the dimensional features of the cross-modal data appear simultaneously in the abnormal samples; Calculating the confidence of the cross-modal dimension feature according to the independent occurrence data of the single-modal data feature in the cross-modal data in the abnormal sample and the co-occurrence data; According to the confidence of the cross-modal dimensional features, a fusion relationship of the cross-modal data is determined.

8. The intelligent processing method for rail transit AI multimodal data according to claim 7 is characterized in that: Calculate the confidence of cross-modal dimension features, including: Obtaining a single anomaly evaluation value of each single modal data feature from the abnormal samples according to the independent occurrence data; Obtaining a co-occurrence anomaly evaluation value according to the co-occurrence data; The confidence of the cross-modal dimensional feature is obtained by calculating the difference coefficient between the sum of the single anomaly evaluation values ​​corresponding to the single modal data features in the co-occurrence data and the co-occurrence anomaly evaluation value.

9. The intelligent processing method of rail transit AI multimodal data according to claim 6 is characterized in that: According to the scene feature analysis graph, feature enhancement granularity sensitivity evaluation is performed on the matching modality data, and granularity configuration is performed according to the evaluation sensitivity, including: Set the particle size enhancement parameters corresponding to the particle size level of feature increase; According to the scene feature analysis graph, the dimensional features are enhanced at multiple granularity levels using granularity enhancement parameters corresponding to preset granularity levels; Using the recognition results of the request task scenario, the sensitivity of each granularity level enhancement process is evaluated to obtain the feature recognition sensitivity; Based on the recognition target and sensitivity requirement threshold of the requested task scenario, the feature recognition sensitivity is determined in granularity to determine the configuration granularity.

10. Intelligent processing system for AI multimodal data of rail transit, characterized by: The system is used to execute the intelligent processing method of rail transit AI multimodal data according to any one of claims 1 to 9, comprising: Feature dimension decomposition module: Use the scene-data analysis relationship to match the corresponding modal data of the request task scene, decompose the data feature dimension relationship of the matching modal data, and establish a scene feature analysis map; Granularity configuration module: performs feature enhancement granularity sensitivity evaluation on matching modal data according to the scene feature analysis map, and performs granularity configuration according to the evaluation sensitivity; Task scene recognition module: according to the configuration granularity, feature enhancement is performed on the matching modal data in the scene feature analysis map, and the enhanced features are used to perform request task scene recognition on the request task scene to obtain the recognition result of the request task scene.

Citation Information

Patent Citations

  • Multi-modal analysis method, system and equipment for industrial inspection scene and medium

    CN119128810A

  • Coal mine fire and gas disaster multi-modal knowledge base construction and optimization method

    CN119322857A

  • Rail transit intelligent control system and method based on Internet of Things

    CN119682818A

  • LLM-based client intention identification and response system, method and device, and medium

    CN119808789A

  • Dynamic data pipeline construction method based on artificial intelligence and multi-modal data processing

    CN119830200A