Data storage method of rail transit command center network big data
By adopting a combination of MongoDB and Redis in the big data storage of rail transit command center network, structured and unstructured data processing problems are solved, efficient and flexible data storage and management are achieved, and system performance and reliability are improved.
Patent Information
- Application Number
- CN202510405256.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-08
AI Technical Summary
The existing big data storage method of rail transit command center network cannot efficiently process structured and unstructured data synchronously, resulting in increased system complexity, high management difficulty and operation and maintenance costs, and obvious performance bottlenecks in high concurrency scenarios.
MongoDB database is used to store structured and unstructured data, and the data format is uniformly converted through JSON format, the Redis cache layer is configured to improve data access speed and system response performance, and MongoDB's dynamic field design and sharding technology are used to establish indexes to support efficient query and analysis.
It realizes unified management and storage of structured and unstructured data, improves data processing efficiency, reduces format conversion complexity, enhances system flexibility and reliability, and optimizes data access speed and system response performance.
Smart Images

Figure CN120277071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data storage, and specifically to a data storage method for the big data of the rail transit command center network. Background Art
[0002] The existing data of the rail transit command center network big data is generally stored fixedly. These stored data are important bases to ensure the safe, efficient and reliable operation of trains. In the existing data management system, structured data is usually stored in a relational database. The relational database has significant advantages in processing structured data, including data consistency and integrity, efficient query and data operation, security guarantee, data integrity constraint, support for multi-user concurrent access, and mature ecosystem and tool support. These advantages make the relational database an important infrastructure for enterprise data management and business systems, and are widely used in various scenarios with data-intensive and transaction processing requirements.
[0003] However, in the existing data storage method of the rail transit command center network big data, it is unable to cope well with unstructured data; unstructured data includes text data, multimedia data, document data, etc. The formats and contents of these data are diverse and do not conform to the fixed mode requirements of the relational database; due to the limitations of the relational database, unstructured data often needs to be stored in a file system or object storage; however, these storages lack powerful query and analysis functions and require additional tools for data processing; this not only increases the complexity of the system, but also introduces more integration and maintenance work; data synchronization and consistency management between different storage systems are also a major challenge, increasing the management difficulty and operation and maintenance cost of the system;
[0004] Secondly, the relational database needs to pre-define the data schema (Schema), which means that before data storage, the structure and field types of the data must be clearly defined; when the data structure changes, the schema must be modified and the data reorganized; this schema-driven design is not flexible enough to cope with rapidly changing and diverse data requirements; moreover, in the rail transit command center, data access is frequent and highly concurrent, especially during peak periods, and the system needs to process a large number of read and write operations; the performance of the relational database in a high-concurrency scenario may become a bottleneck; although the relational database can ensure data consistency through lock mechanisms and transaction processing, in a high-concurrency environment, lock contention and transaction conflicts will significantly affect the response time and processing ability of the database. A large number of read and write operations not only increase the load on the database, but may also cause the database connection pool to be exhausted, resulting in slower system response and even system crashes;
[0005] In view of the above problems, the present invention proposes a solution. Summary of the Invention
[0006] In view of the deficiencies of the prior art, the present invention provides a data storage method for the big data of the rail transit command center network, which solves the problem that it cannot synchronously and quickly process structured data and unstructured data.
[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: A data storage method for the big data of the rail transit command center network, including the following steps:
[0008] S1. Data collection and conversion: Collect structured data and unstructured data in the rail transit system, and convert the collected structured data and unstructured data into JSON data format;
[0009] S2. Data storage: Use the MongoDB database to store the converted data, allowing structured and unstructured data to be stored simultaneously;
[0010] S3. Establish an index in the MongoDB database to support efficient querying and analysis of structured and unstructured data, and configure a Redis cache layer to store frequently accessed data, improving data access speed and system response performance.
[0011] Preferably, in step S1, data collection is specifically carried out through multiple methods such as sensors, data interfaces, web crawlers, and manual input;
[0012] The structured data and unstructured data in the rail transit system specifically include:
[0013] Structured data:
[0014] a. Train operation schedule, including train number, starting station, terminal station, departure time, arrival time;
[0015] b. Line information, including line number, line name, line length, number of stations;
[0016] c. Station information, including station number, station name, station location;
[0017] d. Train status information, including current train speed, location, operation status;
[0018] Unstructured data:
[0019] a. Text data:
[0020] Passenger feedback: Written opinions and suggestions provided by passengers through text messages, emails, and social media;
[0021] Fault log: Error log files generated by the system, recording abnormal and fault information during train operation;
[0022] b. Multimedia data:
[0023] Video surveillance data: including real-time videos and recordings collected by each surveillance camera in the rail transit system;
[0024] Voice recordings: including recordings of conversations between passengers and the command center;
[0025] c. Document data:
[0026] Reports and documents: including technical reports, analysis reports, and meeting minutes in the form of PDF and Word documents.
[0027] Preferably, in step S1, during the process of converting data into JSON format, data cleaning and preprocessing are included to remove redundant and incorrect data;
[0028] The expression for converting the structured data into JSON data format is:
[0029]
[0030] where S represents the set of collected structured data toJSON(s i ) is a function that converts a single structured data s i into JSON format, and n represents the number of elements in the structured data set S;
[0031] The expression for converting unstructured data into JSON data format is:
[0032]
[0033] where U represents the set of collected unstructured data Tokenize(u j ) is a function that performs word segmentation on a single unstructured data u j , and m represents the number of elements in the unstructured data set U.
[0034] Preferably, in step S2, the MongoDB database can store documents with different structures in the same collection by dynamically adding and deleting fields, adapting to the coexistence of structured data and unstructured data.
[0035] Preferably, in step S2, the MongoDB database adopts sharding technology and replica set architecture.
[0036] Preferably, in step S3, the MongoDB database indexes include single-field indexes and composite indexes, and the Redis cache layer is configured to be automatically refreshed to ensure the timeliness and consistency of cached data.
[0037] Preferably, in the step S3, the Redis cache layer includes a master-slave replication mechanism and a cache eviction policy based on the LRU algorithm.
[0038] Preferably, the expression of the LRU algorithm is:
[0039]
[0040] where t i is the current time, indicating that it is the most recently used data, C is the cache capacity, X is the set of data elements stored in the cache, T is the usage time of the cache, and x i is the accessed data, and x k is the least recently used data in the cache, and t k is the smallest timestamp in the cache.
[0041] Preferably, in the step S3, the data with high access frequency is determined according to historical access records and predefined rules;
[0042] The predefined rules are as follows:
[0043] a. Access frequency threshold: Set an access frequency threshold, and only the data items whose access frequency exceeds this threshold will be considered as data with high access frequency;
[0044] b. Data priority: Set different priorities for different data items according to the importance of the data items and business requirement factors;
[0045] c. Time window: Set the time window length for calculating the access frequency, which can be specifically the most recent week and the most recent month.
[0046] Preferably, in the step S3, a data synchronization mechanism is configured between the Redis cache layer and the MongoDB database.
[0047] The present invention provides a data storage method for the big data of the rail transit command center network. It has the following beneficial effects:
[0048] 1. By uniformly converting and managing structured data and unstructured data in JSON format, the present invention makes data processing and transmission more convenient, reduces the complexity of data parsing and format conversion. At the same time, through the MongoDB database, fields can be dynamically added and deleted, allowing documents with different structures to be stored in the same collection. This schema-free design enables MongoDB to flexibly adapt to different types of data, avoiding the limitation of the traditional relational database that requires a predefined fixed schema. By converting all data into JSON format and storing it in MongoDB, unified management and storage of data structures are achieved.
[0049] 2. By establishing single-field indexes and composite indexes in MongoDB, the present invention significantly improves the query and analysis efficiency of structured and unstructured data; and it configures a Redis cache layer for storing frequently accessed data. The Redis cache layer adopts a master-slave replication mechanism to ensure the high availability of cached data, and performs cache eviction based on the LRU algorithm. At the same time, frequently accessed data is determined according to historical access records and predefined rules to ensure that the data stored in the cache layer is the most frequently accessed data. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The technical solutions of the present invention will be clearly and completely described below with reference to the drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] Embodiment:
[0053] Please refer to the attached Figure 1 , the embodiment of the present invention provides a data storage method for the big data of the rail transit command center network, including the following steps:
[0054] S1. Data collection and conversion: Collect structured data and unstructured data in the rail transit system, and convert the collected structured data and unstructured data into the JSON data format;
[0055] S2. Data storage: Use the MongoDB database to store the converted data, allowing the simultaneous storage of structured and unstructured data;
[0056] S3. Establish indexes in the MongoDB database to support the efficient query and analysis of structured and unstructured data, and configure a Redis cache layer for storing frequently accessed data to improve data access speed and system responsiveness.
[0057] Specifically, by converting data into JSON format, the uniformity of data format is achieved, simplifying the subsequent data processing and transmission processes and improving data processing efficiency. The schema-free design and sharding technology of the MongoDB database enable flexible and efficient data storage, meeting the requirements of large-scale data storage and high-performance queries while ensuring high availability and reliability of data. Indexes are established in the MongoDB database and a Redis cache layer is configured. By caching frequently accessed data, the data access speed and system response performance are significantly improved. The LRU algorithm ensures that the data in the cache is always the most recently used, enhancing the utilization efficiency of the cache. Frequently accessed data is determined through historical access records and predefined rules, further optimizing system performance. The data synchronization mechanism guarantees the consistency and integrity of cached data and database data, enhancing the reliability of the system.
[0058] In step S1, data collection is specifically carried out through multiple methods such as sensors, data interfaces, web crawlers, and manual input.
[0059] The structured and unstructured data in the rail transit system specifically includes:
[0060] Structured data:
[0061] a. Train operation timetables, including train numbers, starting stations, terminal stations, departure times, and arrival times;
[0062] b. Line information, including line numbers, line names, line lengths, and the number of stations;
[0063] c. Station information, including station numbers, station names, and station locations;
[0064] d. Train status information, including the current speed, location, and operation status of the train;
[0065] Unstructured data:
[0066] a. Text data:
[0067] Passenger feedback: Written opinions and suggestions provided by passengers via text messages, emails, and social media;
[0068] Fault logs: Error log files generated by the system, recording abnormal and fault information during train operation;
[0069] b. Multimedia data:
[0070] Video surveillance data: Including real-time videos and recordings collected by various surveillance cameras in the rail transit system;
[0071] Voice recordings: Including recordings of conversations between passengers and the command center;
[0072] c. Document data:
[0073] Reports and documents: including technical reports, analysis reports, and meeting minutes in PDF and Word document forms;
[0074] In step S1, the process of converting data into JSON format includes data cleaning and preprocessing to remove redundant and incorrect data;
[0075] The expression for converting structured data into JSON data format is:
[0076]
[0077] where S represents the set of structured data collected, toJSON(s i ) is a function that converts a single structured data s i into JSON format, and n represents the number of elements in the set of structured data S;
[0078] The expression for converting unstructured data into JSON data format is:
[0079]
[0080] where U represents the set of unstructured data collected, Tokenize(u j ) is a function that performs word segmentation on a single unstructured data u j and m represents the number of elements in the set of unstructured data U.
[0081] Specifically, the data of the rail transit system is collected through sensors, data interfaces, web crawlers, and manual input. Sensors are used to monitor the running status of trains in real-time. Data interfaces are used to obtain train operation schedules and line information. Web crawlers are used to capture passenger feedback information. Manual input is used to record the observations and reports of on-site staff. The collected data includes structured data and unstructured data. Structured data such as train operation schedules and line information, and unstructured data such as passenger feedback and video surveillance data. To facilitate storage and processing, all these data are converted into JSON format. The expression for converting structured data into JSON data format can convert all structured data into JSON format, which is convenient for storage and management in MongoDB. Moreover, the unified data format of structured data simplifies the data processing and transmission process, improves the data processing efficiency, and the JSON format has good compatibility, facilitating data interaction with other systems or applications. From the expression for converting unstructured data into JSON data format, word segmentation processing enables unstructured data to be parsed and understood, facilitating subsequent data processing and analysis. Through word segmentation processing, useful information can be extracted from unstructured data, improving the utilization value of the data, and standardizing unstructured data into a processable format, facilitating storage and query in the database.
[0082] In step S2, the MongoDB database can store documents with different structures in the same collection by dynamically adding and deleting fields, adapting to the coexistence of structured data and unstructured data. In step S2, the MongoDB database adopts a sharding technology and a replica set architecture.
[0083] Specifically, by using the MongoDB database, large-scale structured and unstructured data can be stored and managed in one system, achieving efficient storage and management of data, while ensuring high availability and reliability of the data.
[0084] In step S3, the MongoDB database indexes include single-field indexes and composite indexes. The Redis cache layer is configured to be automatically refreshed to ensure the timeliness and consistency of the cached data.
[0085] In step S3, the Redis cache layer includes a master-slave replication mechanism and a cache eviction policy based on the LRU algorithm.
[0086] The LRU algorithm expression is:
[0087]
[0088] where t iis the current time, indicating that it is the most recently used data, C is the cache capacity, X is the set of data elements stored in the cache, T is the usage time of the cache, and x i is the accessed data, x k is the least recently used data in the cache, t k is the smallest timestamp in the cache.
[0089] Specifically, create single-field indexes and composite indexes in the MongoDB database. Single-field indexes are suitable for simple queries, and composite indexes are suitable for complex queries, which can significantly improve the query efficiency of data; configure the Redis cache layer to store frequently accessed data; Redis adopts a master-slave replication mechanism and a cache eviction policy based on the LRU algorithm. The master-slave replication mechanism can ensure the high availability of cached data, and the LRU algorithm can effectively manage the cache space to ensure that the data stored in the cache is the most recently used data, improving the cache hit rate. In addition, through the LRU algorithm expression, we can free up space to store new data by evicting data that has not been used for a long time, optimize system performance, and reduce data access latency; the LRU algorithm can also automatically adjust the cache content according to the data access pattern to adapt to dynamic access requirements.
[0090] In step S3, the frequently accessed data is determined according to historical access records and predefined rules;
[0091] The predefined rules are as follows:
[0092] a. Access frequency threshold: Set an access frequency threshold, and only data items whose access frequency exceeds this threshold will be considered frequently accessed data;
[0093] b. Data priority: Set different priorities for different data items according to factors such as the importance of the data item and business requirements;
[0094] c. Time window: Set the length of the time window used to calculate the access frequency, which can be specifically the most recent week and the most recent month;
[0095] In step S3, a data synchronization mechanism is configured between the Redis cache layer and the MongoDB database.
[0096] Specifically, determining the frequently accessed data through historical access records and predefined rules can ensure that the data stored in the cache layer is the most frequently accessed data, improving the access speed and response performance of the system; the data synchronization mechanism ensures the data consistency and integrity between the Redis cache layer and the MongoDB database, enhancing the reliability of the system.
[0097] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A data storage method for the big data of the rail transit command center network, characterized in that, It includes the following steps: S1. Data collection and conversion: Collect structured and unstructured data in the rail transit system, and convert the collected structured and unstructured data into JSON data format; S2. Data storage: Use the MongoDB database to store the converted data, allowing the simultaneous storage of structured and unstructured data; S3. Establish indexes in the MongoDB database to support efficient querying and analysis of structured and unstructured data, and configure a Redis cache layer to store frequently accessed data, improving data access speed and system response performance.
2. The data storage method of the rail transit command center line network big data according to claim 1, characterized in that, In step S1, data collection is specifically carried out through multiple methods such as sensors, data interfaces, web crawlers, and manual input; The structured and unstructured data in the rail transit system specifically includes: Structured data: a. Train operation timetables, including train numbers, starting stations, terminal stations, departure times, and arrival times; b. Line information, including line numbers, line names, line lengths, and the number of stations; c. Station information, including station numbers, station names, and station locations; d. Train status information, including the current speed, location, and operation status of the train; Unstructured data: a. Text data: Passenger feedback: Written opinions and suggestions provided by passengers through text messages, emails, and social media; Fault logs: Error log files generated by the system, recording abnormal and fault information during train operation; b. Multimedia data: Video surveillance data: Including real-time videos and recordings collected by various surveillance cameras in the rail transit system; Voice recordings: Including recordings of conversations between passengers and the command center; c. Document data: Reports and documents: Including technical reports, analysis reports, and meeting minutes in the form of PDF and Word documents.
3. The data storage method of the rail transit command center line network big data according to claim 1, characterized in that, In step S1, during the process of converting data into JSON format, data cleaning and preprocessing are included to remove redundant and error data; The expression for converting structured data into JSON data format is: Among them, S represents the set of collected structured data toJSON(s i ) is a function that converts a single piece of structured data s i into the JSON format, and n represents the number of elements in the set of structured data S; The expression for converting unstructured data into JSON data format is: Among them, U represents the set of unstructured data collected Tokenize(u j ) is a function for tokenizing a single unstructured data u j The function, m represents the number of elements in the unstructured data set U.
4. The data storage method of the rail transit command center line network big data according to claim 1, characterized in that In step S2, the MongoDB database can store documents with different structures in the same collection by dynamically adding and deleting fields, adapting to the coexistence of structured and unstructured data.
5. The data storage method of the rail transit command center line network big data according to claim 1, characterized in that In step S2, the MongoDB database adopts a sharding technology and replica set architecture.
6. The data storage method of the rail transit command center line network big data according to claim 1, characterized in that In step S3, the MongoDB database indexes include single-field indexes and composite indexes, and the Redis cache layer is configured to be automatically refreshed to ensure the timeliness and consistency of cached data.
7. The data storage method of the rail transit command center line network big data according to claim 1, characterized in that, In step S3, the Redis cache layer includes a master-slave replication mechanism and a cache eviction policy based on the LRU algorithm.
8. The data storage method of the rail transit command center line network big data according to claim 1, characterized in that The expression of the LRU algorithm is: where t i is the current time, indicating that it is the most recently used data, C is the cache capacity, X is the set of data elements stored in the cache, T is the usage time of the cache, and x i is the accessed data, and x k is the least recently used data in the cache, and t k is the smallest timestamp in the cache.
9. The data storage method of the rail transit command center line network big data according to claim 1, characterized in that In step S3, frequently accessed data is determined based on historical access records and predefined rules; The predefined rules are as follows: a. Access frequency threshold: Set an access frequency threshold, and only data items with an access frequency exceeding this threshold will be considered frequently accessed data; b. Data priority: Different priorities are set for different data items according to the importance of the data items and business requirement factors; c. Time window: Set the length of the time window used to calculate the access frequency, which can specifically be the most recent week and the most recent month.
10. The data storage method of the rail transit command center line network big data according to claim 1, wherein, In the step S3, a data synchronization mechanism is configured between the Redis cache layer and the MongoDB database.
Citation Information
Cited By
SPEC term structured storage method and system
CN120743908A