Method and device for managing railway passenger data

By constructing a full lifecycle management framework for railway passenger transport data, the problems of data silos and low productization have been solved, enabling efficient development of data products and enhancement of commercial value. Big data and artificial intelligence technologies are used for data management and analysis.

CN122132457APending Publication Date: 2026-06-02CHINA ACADEMY OF RAILWAY SCI CORP LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA ACADEMY OF RAILWAY SCI CORP LTD
Filing Date
2026-02-27
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

The management of railway passenger transport data suffers from problems such as data silos, low levels of data productization, lack of full lifecycle management and unsystematic technical means, resulting in long development cycles, high operating costs, uncontrollable quality, and difficulty in measuring value.

Method used

We will construct a systematic management framework covering the entire lifecycle of railway passenger transport data. Through preprocessing, statistical analysis, model prediction, and joint analysis, we will realize data asset ownership confirmation, differentiated storage, and updates. Combining big data, artificial intelligence, and privacy computing technologies, we will provide a full lifecycle data product management method.

Benefits of technology

It has improved the development efficiency, service quality, and commercial value of data products, solved the problem of data silos, and enabled rapid iteration and large-scale operation of data products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132457A_ABST
    Figure CN122132457A_ABST
Patent Text Reader

Abstract

This invention provides a method for managing railway passenger transport data, comprising: preprocessing diverse heterogeneous data to form railway passenger transport data, which is divided into statistical analysis data, model prediction data, and joint analysis data; for statistical analysis data, performing real-time or offline indicator calculations to generate real-time or offline indicators for storage and sharing; for model prediction data, constructing a prediction model for training and calculation to generate feature indicators for storage and sharing; and for joint analysis data, analyzing it through a predetermined joint analysis model to generate score indicators for storage and sharing. This invention also provides a device, storage medium, and electronic device for the full lifecycle management of railway passenger transport data. Thus, by constructing a systematic data management framework covering the entire lifecycle of railway passenger transport data, this invention improves the development efficiency, service quality, and commercial value of data products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data management technology, and in particular to a method, apparatus, storage medium, and electronic device for managing railway passenger data. Background Technology

[0002] With the rapid development of my country's railway passenger transport system and the continuous deepening of its digitalization, railway passenger transport systems (such as the 12306 system) generate massive amounts of multi-source heterogeneous data during daily production and operation. This business data includes, but is not limited to, details of ticket sales / reservations / waitlists, train schedules, passenger information characteristics, station passenger flow input / output, real-time train locations, equipment operating status, catering orders, and air-rail / rail-water intermodal transport information. As a crucial production factor, railway passenger transport data holds immense value and is of great significance for improving railway transport efficiency, optimizing passenger travel services, achieving precise marketing, supporting scientific decision-making, and promoting cross-industry collaborative development. Since its inception, the passenger transport system has continuously enriched and improved its functions, and the types of data have expanded. Based on this, the railway department has designed and developed a large amount of railway passenger transport data, which is showing a development trend of multi-domain integration, multi-dimensional expansion, intelligent operation, and ecological services.

[0003] With the increasing diversity of railway passenger transport data formats and the rapid iteration of data value mining technologies, the challenges faced by the development and management of railway passenger transport data products are becoming more severe. Currently, the railway sector faces the following main pain points in railway passenger transport data management and application:

[0004] I. The problem of data silos has not been fully resolved. Data is scattered across different business systems (such as ticketing systems, scheduling systems, and passenger service systems), with varying formats and standards. It lacks effective collection, integration, and governance, making it difficult to form a unified data asset.

[0005] Second, the level of data productization is still not high. The application of railway passenger transport data is mostly limited to simple report statistics and ad-hoc queries. There is a lack of systematic methods to process, model, and encapsulate raw data into reusable, operable, and high-value data products (such as "station / city passenger flow forecast", "holiday popular route analysis", "passenger value rating assessment", "frequent flyer precise recommendation service").

[0006] Third, there is a lack of data product lifecycle management. For the entire lifecycle of railway passenger data—from demand verification, system design, product development, deployment, operation monitoring, value assessment, version iteration, to decommissioning—there is a lack of standardized, automated, and visualized management processes and platforms. This results in long development cycles, high operating costs, uncontrollable quality, and difficulty in measuring the value of railway passenger data.

[0007] Fourth, the technical means for developing data products have not yet formed a systematic framework. Existing management models rely heavily on manual processes and scattered tools, which require a high level of expertise from data product developers and lack integrated solutions that deeply integrate with modern technologies such as big data, artificial intelligence, and privacy computing. This makes it impossible to meet the needs of rapid iteration and large-scale operation of data products.

[0008] Therefore, the railway passenger transport sector urgently needs a systematic data product management method and platform that covers the entire lifecycle to solve the above problems and fully explore and release the value of railway passenger transport data products.

[0009] In conclusion, the existing technology obviously has inconveniences and defects in practical use, so it is necessary to improve it. Summary of the Invention

[0010] To address the aforementioned shortcomings, the present invention aims to provide a method, apparatus, storage medium, and electronic device for managing railway passenger transport data. By constructing a systematic data management framework covering the entire lifecycle of railway passenger transport data, it improves the development efficiency, service quality, and commercial value of data products.

[0011] To solve the above-mentioned technical problems, the present invention is implemented as follows:

[0012] In a first aspect, embodiments of the present invention provide a method for managing railway passenger transport data, including:

[0013] The preprocessing step involves preprocessing the diverse and heterogeneous data to form railway passenger transport data, which is divided into statistical analysis data, model prediction data, and joint analysis data.

[0014] The statistical analysis step involves calculating real-time or offline indicators for the statistical analysis data, generating corresponding real-time or offline indicators for storage and sharing.

[0015] The model prediction step involves, for the model prediction data, constructing a prediction model for training and computation, generating corresponding feature indicators for storage and sharing;

[0016] The joint analysis step involves analyzing the joint analysis data using a predetermined joint analysis model to generate corresponding score indicators for storage and sharing.

[0017] According to the railway passenger data management method of the present invention, the preprocessing step further includes:

[0018] The diverse and heterogeneous data is cleaned, transformed, and standardized to form the railway passenger transport data.

[0019] The railway passenger transport data is recorded as a data asset in the table, and data ownership is confirmed for the data asset.

[0020] According to the railway passenger data management method of the present invention, the statistical analysis step further includes:

[0021] For statistical analysis data with real-time requirements, streaming extraction is performed, and real-time indicators are calculated using a streaming computing platform and a pre-defined data element calculation logic library to generate corresponding real-time indicators for storage and sharing.

[0022] For statistical analysis data that does not require real-time processing, the data is loaded into the data warehouse. The distributed storage and computing resources of the big data cluster, as well as custom functions, are used to calculate offline indicators and generate corresponding offline indicators for storage and sharing.

[0023] According to the railway passenger data management method of the present invention, the model prediction step further includes:

[0024] The prediction model is designed and constructed using the extracted basic feature labels and the model Y value;

[0025] The prediction model quantifies the basic feature labels, optimizes the model parameters by performing multiple rounds of model training on the training dataset, and generates corresponding feature indicators for storage and sharing.

[0026] According to the railway passenger data management method of the present invention, the joint analysis step further includes:

[0027] The railway passenger transport indicators and non-railway industry indicators are numerically processed and then used together with the model Y value as input to the joint analysis model.

[0028] The joint analysis model performs multiple rounds of model training and gradient updates through federated learning to generate corresponding score indicators for storage and sharing.

[0029] According to the railway passenger data management method of the present invention, the method further includes the following steps after the preprocessing step:

[0030] The differentiated storage step involves classifying and separating the railway passenger data into hot and cold categories based on its source, importance, timeliness, data volume, mining method, frequency of use, scope of use, and / or whether there is a need for joint queries.

[0031] The railway passenger data management method according to the present invention further includes:

[0032] The data update step involves formulating different data update strategies based on factors such as the access frequency, data granularity, data range, computational cost, and / or result set size of the railway passenger data during external service processes, and updating the railway passenger data according to the data update strategies.

[0033] Secondly, embodiments of the present invention provide a full lifecycle management device for railway passenger transport data constructed based on any one of the methods described above, the device comprising:

[0034] The preprocessing module is used to preprocess diverse and heterogeneous data to form railway passenger transport data, which is divided into statistical analysis data, model prediction data, and joint analysis data.

[0035] The statistical analysis module is used to perform real-time or offline indicator calculations on the statistical analysis data, and generate corresponding real-time or offline indicators for storage and sharing.

[0036] The model prediction module is used to train and calculate a prediction model for the model prediction data, and generate corresponding feature indicators for storage and sharing.

[0037] The joint analysis module is used to analyze the joint analysis data using a predetermined joint analysis model, and generate corresponding score indicators for storage and sharing.

[0038] Thirdly, embodiments of the present invention provide a storage medium for storing a computer program for performing any of the methods described herein.

[0039] Fourthly, embodiments of the present invention provide an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement any of the methods described above.

[0040] This invention provides a data development and sharing platform covering the entire lifecycle of railway passenger transport data. It employs differentiated development processes designed for three types of railway passenger transport data characteristics: for statistical analysis data, real-time or offline indicator calculations are performed to generate corresponding real-time or offline indicators; for model prediction data, a prediction model is built, trained, and calculated to generate corresponding feature indicators; for joint analysis data, a predetermined joint analysis model is used for analysis to generate corresponding score indicators. Therefore, this invention improves the development efficiency, service quality, and commercial value of data products by constructing a systematic data management framework covering the entire lifecycle of railway passenger transport data. Attached Figure Description

[0041] Figure 1 This is a flowchart illustrating the railway passenger data management method provided in Embodiment 1 of the present invention;

[0042] Figure 2 This is a flowchart illustrating the railway passenger data management method provided in Embodiment 2 of the present invention;

[0043] Figure 3 This is a framework diagram of the railway passenger data product development and sharing platform provided in Embodiment 3 of the present invention;

[0044] Figure 4 This is a flowchart of the development process for statistical analysis data provided in Embodiment 4 of the present invention;

[0045] Figure 5 This is a flowchart of the development process for model predictive data provided in Embodiment 5 of the present invention;

[0046] Figure 6 This is a flowchart of the development process for synthetic analytical data provided in Embodiment Six of the present invention;

[0047] Figure 7 This is a schematic diagram of the structure of the railway passenger data management device provided in Embodiment 7 of the present invention;

[0048] Figure 8 This is a schematic diagram of the structure of the railway passenger data management device provided in Embodiment 8 of the present invention;

[0049] Figure 9 This is a schematic diagram of the structure of the electronic device provided in Embodiment 9 of the present invention. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0051] It should be noted that references to "an embodiment," "embodiment," "example embodiment," etc., in this specification refer to the described embodiment including specific features, structures, or characteristics, but not every embodiment must include these specific features, structures, or characteristics. Furthermore, such expressions do not refer to the same embodiment. Moreover, when describing specific features, structures, or characteristics in conjunction with embodiments, whether or not explicitly described, it is indicated that incorporating such features, structures, or characteristics into other embodiments is within the knowledge of those skilled in the art.

[0052] Furthermore, certain terms are used in the specification and subsequent claims to refer to specific components or parts. Those skilled in the art will understand that manufacturers may use different names or terms to refer to the same component or part. This specification and subsequent claims do not distinguish components or parts by differences in name, but rather by differences in function. The terms "comprising" and "including" used throughout the specification and subsequent claims are open-ended and should be interpreted as "including but not limited to." Additionally, the term "connection" here includes any direct and indirect electrical connection means. Indirect electrical connection means include connections made through other means.

[0053] The following description, in conjunction with the accompanying drawings, details the railway passenger data management method provided by the embodiments of the present invention through specific implementations and application scenarios.

[0054] This invention aims to overcome the current shortcomings in the development and sharing of railway passenger transport data products, and to provide a full lifecycle management method. The core objective is to construct an integrated management framework covering the entire process of railway passenger transport data from "birth" to "retirement," achieving standardized production, automated operation and maintenance, intelligent operation, and visualized control of railway passenger transport data, thereby improving the development efficiency, service quality, and commercial value of data products.

[0055] Figure 1 This is a flowchart illustrating a railway passenger data management method provided in Embodiment 1 of the present invention, the method comprising:

[0056] Step S101, preprocessing step, preprocessing the diverse heterogeneous data to form railway passenger transport data, which is divided into statistical analysis data, model prediction data and joint analysis data.

[0057] Preferably, the diverse heterogeneous data comes from heterogeneous data from different railway systems, platforms, and even external units, including but not limited to ticketing data such as passenger ticket purchases, refunds, changes, and waitlists; passenger behavior data such as passenger travel preferences, ticket purchase frequency, and meal ordering information; train operation data such as train operation plans, train timetables, and real-time train locations; equipment status data such as the opening status of turnstiles and ticket windows, and the number of handheld ticket checking devices activated; publicly available data such as real-time / future weather and traffic control information; and embedded log data from railway passenger transport systems and platforms.

[0058] Preferably, the preprocessing step further includes:

[0059] The diverse and heterogeneous data is cleaned, transformed, and standardized to form railway passenger transport data.

[0060] Railway passenger transport data will be recorded as a data asset and data ownership will be confirmed.

[0061] Better yet, the rights to be confirmed include ownership, processing, use, and management rights.

[0062] Step S102, statistical analysis step: For statistical analysis data, perform real-time indicator calculation or offline indicator calculation, generate corresponding real-time indicators or offline indicators for storage and sharing.

[0063] Preferably, the statistical analysis data includes train schedules, ticket purchase and refund information, station passenger flow input / output, popular stations, etc.

[0064] Preferably, the statistical analysis step further includes:

[0065] For statistical analysis data with real-time requirements, streaming extraction is performed, and real-time indicators are calculated using a streaming computing platform and a pre-defined data element calculation logic library to generate corresponding real-time indicators for storage and sharing.

[0066] For statistical analysis data that does not require real-time processing, the data warehouse can be loaded with distributed storage and computing resources of the big data cluster, as well as custom functions to perform offline indicator calculations, generate corresponding offline indicators for storage and sharing.

[0067] Step S103, Model Prediction Step: For model prediction data, a prediction model is built, trained, and calculated to generate corresponding feature indicators for storage and sharing.

[0068] Preferably, the model predictive data includes passenger value ratings, station passenger flow and distribution predictions, and passenger residence locations.

[0069] Preferably, the model prediction step further includes:

[0070] A prediction model is designed and constructed using the extracted basic feature labels and the model Y value.

[0071] The basic feature labels are numerically processed by the prediction model, and the model parameters are optimized by training the training dataset in multiple rounds. The corresponding feature indicators are then generated for storage and sharing.

[0072] Step S104, Joint Analysis Step: For joint analysis data, analysis is performed using a predetermined joint analysis model to generate corresponding score indicators for storage and sharing.

[0073] Preferably, the joint analytical data includes identification of air-rail intermodal passengers and analysis of potential customers for air-rail intermodal transport.

[0074] Preferably, the joint analysis step further includes:

[0075] Railway passenger transport indicators and non-railway industry indicators are quantified and then used together with the model Y value as input to the joint analysis model.

[0076] The joint analysis model uses federated learning to perform multiple rounds of model training and gradient updates, generating corresponding score indicators for storage and sharing.

[0077] It is worth noting that steps S102 to S103 are parallel and have no sequential relationship. In other words, the appropriate step is selected to be executed based on the type of railway passenger data. Step S203 is executed for statistical analysis data, step S204 is executed for model prediction data, and step S205 is executed for joint analysis data.

[0078] This invention relates to a railway passenger transport data development and sharing platform covering the entire lifecycle, and a development process designed with differentiated features for three types of railway passenger transport data. The railway passenger transport data development and sharing platform is the foundation for the internal and external circulation and sharing of railway passenger transport data elements. Through key processing steps such as unified coding, dictionary extraction, invalid data filtering, missing data completion, anonymization, and privacy encryption, it achieves effective integration of data from various systems and platforms, greatly solving the problems of insufficient available data and "data silos." By confirming data asset ownership, the usable data catalogs and scopes for each department are determined, thereby achieving access control. The platform integrates mainstream data product development technology stacks and supports horizontal expansion. It also includes a large number of railway passenger transport-specific data product algorithm models, meeting the design and development needs of various data products. Furthermore, the platform provides multiple data product sharing methods, enabling the rapid release of railway passenger transport data value while maintaining security and controllability.

[0079] This invention designs different development processes for railway passenger transport data requirements, namely statistical analysis, model prediction, and joint analysis. These three data product development processes integrate cutting-edge technologies such as multidimensional data analysis, big data distributed computing, artificial intelligence models, and privacy computing technologies (such as federated learning). They cover data product requirements in scenarios involving the flow and sharing of internal and external data elements in railway passenger transport, and greatly improve the development efficiency of railway passenger transport data by leveraging a self-developed reusable indicator calculation logic library and pre-trained artificial intelligence models.

[0080] Figure 2 This is a flowchart illustrating the railway passenger data management method provided in Embodiment 2 of the present invention, the method comprising:

[0081] Step S201, preprocessing step, preprocesses the diverse heterogeneous data to form railway passenger transport data, which is divided into statistical analysis data, model prediction data and joint analysis data.

[0082] Preferably, the preprocessing step further includes:

[0083] The diverse and heterogeneous data is cleaned, transformed, and standardized to form railway passenger transport data.

[0084] Railway passenger transport data will be recorded as a data asset and data ownership will be confirmed.

[0085] Better yet, the rights to be confirmed include ownership, processing, use, and management rights.

[0086] Step S202, Differentiated storage step, involves classifying and storing railway passenger data in a hierarchical and cold / hot format based on the source, importance, timeliness, data volume, mining method, frequency of use, scope of use, and / or whether there is a need for joint queries.

[0087] Step S203, statistical analysis step: For statistical analysis data, perform real-time indicator calculation or offline indicator calculation, generate corresponding real-time indicators or offline indicators for storage and sharing.

[0088] Preferably, the statistical analysis step further includes:

[0089] For statistical analysis data with real-time requirements, streaming extraction is performed, and real-time indicators are calculated using a streaming computing platform and a pre-defined data element calculation logic library to generate corresponding real-time indicators for storage and sharing.

[0090] For statistical analysis data that does not require real-time processing, the data warehouse can be loaded with distributed storage and computing resources of the big data cluster, as well as custom functions to perform offline indicator calculations, generate corresponding offline indicators for storage and sharing.

[0091] Step S204, Model Prediction Step: For model prediction data, a prediction model is built, trained, and calculated to generate corresponding feature indicators for storage and sharing.

[0092] Preferably, the model prediction step further includes:

[0093] A prediction model is designed and constructed using the extracted basic feature labels and the model Y value.

[0094] The basic feature labels are numerically processed by the prediction model, and the model parameters are optimized by training the training dataset in multiple rounds. The corresponding feature indicators are then generated for storage and sharing.

[0095] Step S205, Joint Analysis Step: For joint analysis data, analysis is performed using a predetermined joint analysis model to generate corresponding score indicators for storage and sharing.

[0096] Preferably, the joint analysis step further includes:

[0097] Railway passenger transport indicators and non-railway industry indicators are quantified and then used together with the model Y value as input to the joint analysis model.

[0098] The joint analysis model uses federated learning to perform multiple rounds of model training and gradient updates, generating corresponding score indicators for storage and sharing.

[0099] It is worth noting that steps S203 through S205 are parallel and have no sequential relationship. In other words, the appropriate step is selected based on the type of railway passenger data. Step S203 is executed for statistical analysis data, step S204 is executed for model prediction data, and step S205 is executed for joint analysis data.

[0100] Step S206, data update step: Based on the access frequency, data granularity, data range, computing cost and / or result set size of railway passenger data in the external service process, formulate different data update strategies, and update the railway passenger data according to the data update strategies.

[0101] To better understand the present invention, the technical solution of the present invention will be described below with reference to specific embodiments.

[0102] Part One: Railway Passenger Transport Data Product Development and Sharing Platform

[0103] Figure 3 This is a framework diagram of the railway passenger transport data product development and sharing platform provided in Embodiment 3 of the present invention. Its key processes include: collection and assetization of multi-source heterogeneous data; hierarchical classification and differentiated processing and storage of data assets; management of data analysis and computing engine; design and development of railway passenger transport data system; circulation and sharing of railway passenger transport data; and updating and destruction of railway data products.

[0104] I. Collection and Assetization of Multi-Source Heterogeneous Data: This is a prerequisite for the construction of the railway passenger transport data system and the internal and external circulation and sharing of railway data elements. This stage involves the intra-network / cross-network transmission, cleaning, transformation, and data entry and ownership confirmation (including ownership, processing rights, usage rights, and management rights) of heterogeneous data from different railway systems, platforms, and even external units (including but not limited to ticketing data such as passenger ticket purchases, refunds, changes, and waitlists; passenger behavior data such as travel preferences, ticket purchase frequency, and meal ordering information; train operation data such as train schedules, timetables, and real-time train locations; equipment status data such as gate and window opening status and the number of handheld ticket checking devices activated; publicly available data such as real-time / future weather and traffic control information; and embedded log data from the railway passenger transport system and platform). Its core objective is to solve problems such as incomplete available data, low data quality, and difficulties in joint data queries across railway bureaus.

[0105] Taking the catering data of various railway bureaus as an example, train meal ordering information can be uniformly collected through channels such as the Railway 12306 APP, website, and mini-program. Data cleaning, transformation, and loading into the database are performed using log parsing. Currently, this process relies on manual code development by data ETL personnel, processing different types of data separately. However, through the railway passenger data development and sharing platform proposed in this invention, standardized processing can be achieved, eliminating problems such as inconsistent data formats and poor data quality. Another type of data is collected from the catering databases of various railway bureaus. This type of data suffers from problems such as missing field values, low data quality, and inconsistent format standards. For example, railway bureau A's catering data is stored in a Greenplum database, while railway bureau B's catering data is stored in a Sybase database. The data types of these two databases differ, and compatibility between these data types needs to be ensured during the unified extraction and collection process. In the catering system of Railway Bureau A, a product with ID (field name sp_id) of 00001 (field name sp_name) represents 500ml mineral water of brand A. However, in the catering system of Railway Bureau B, the ID (field name commodity_id) of the same product is 00002, and the product name (field name commodity_name) is mineral water of brand A (500ml). The platform described in this invention will use a uniformly defined product dictionary table to format the data details collected from different railway bureaus as goods_id=S000001, goods_type=mineral water, goods_brand=brand A, goods_vol=500ml. This allows for cross-bureau joint queries when analyzing popular products and revenue of various products across the entire railway network. Issues such as missing field values ​​and outlier interference (dirty data) can be addressed by associating with the platform's dictionary table to complete the data and correct interference values. For example, if Railway Bureau A's product sales details table records the product ID and sales amount but not the sales quantity, this can be completed by associating with the product dictionary. The commodity sales details table of Railway Bureau B records commodity_id=00002, but commodity_name=A mineral water 1L. This can be corrected through this platform to goods_type=mineral water, goods_brand=A brand, and goods_vol=500ml.

[0106] II. Data Asset Classification and Differentiated Storage: Railway passenger transport data will be classified and stored separately for hot and cold data based on its source, importance, timeliness, volume, mining method, frequency of use, scope of use, and whether there is a need for joint queries. For example, the massive daily log data from various platforms and systems will be loaded into a distributed file system, and key indicators will be extracted using big data analytics. Production data involving real-time transactions or frequently changing statuses will be stored in a real-time database or real-time data warehouse, using SSDs or other storage media for easy and fast querying. Business data with lower timeliness but higher mining value will be stored in an analytical database, and daily business reports will be developed using a multi-dimensional analysis engine. Data from data middleware, real-time interface transmissions, etc., can be parsed in real time and loaded into a data lake for storage through streaming computing platforms (such as Flink CDC). Based on the data asset ownership confirmation results, different access and usage permissions will be assigned to various users / departments of the railway data product development and sharing platform to meet permission isolation requirements and avoid disputes over the division of operating revenue during the subsequent data product service period.

[0107] III. Data Analysis and Computation Method Library Management: To meet the needs of developing different railway passenger transport data, it should possess data analysis tools for different purposes. These include tools such as Hive / Gbase for offline label calculation, Doris / Clickhouse for real-time multidimensional analysis, artificial intelligence models suitable for analysis objectives such as passenger flow prediction and passenger value assessment, and privacy-preserving computation methods such as differential privacy, secure multi-party computation, and federated learning that may be used when conducting joint analysis with external partners.

[0108] The railway passenger transport data development and sharing platform proposed in this invention incorporates the aforementioned data analysis methods, core algorithms, and models into a unified knowledge base for management. It supports operations such as adding, modifying, updating, and deleting within the same management page. When developing railway passenger transport data, users can select the required calculation methods or models by clicking through the dropdown menu on the page, eliminating the need to repeatedly develop these reusable algorithms or models, thereby greatly improving the efficiency of railway passenger transport data development.

[0109] IV. Design and Development of Railway Passenger Transport Data System: This invention conducts feasibility analysis on the internal and external data product requirements collected through industry surveys and in-depth interviews. The needs of each data product service recipient are refined and categorized to construct a complete and feasible data product system. Differentiated data product development routes are also studied. This invention classifies railway passenger transport data into three categories: statistical analysis data (such as train schedules, ticket purchase and refund information, station passenger flow input / output, and popular stations), model prediction data (such as passenger value ratings, station passenger flow and distribution predictions, and passenger residence locations), and joint analysis data (such as identification of air-rail intermodal passengers and analysis of potential air-rail intermodal customers). Different analysis methods are employed for each of these three types of railway passenger transport data. For statistical analysis data, offline / real-time computing and multi-dimensional analysis engines can be used. For model prediction data, an artificial intelligence model can be designed and constructed by combining basic feature labels and Y-values, and score indicators can be obtained through multiple rounds of model training. For joint analysis data, privacy-preserving computing technologies (such as federated learning) can be used.

[0110] The railway passenger transport data development and sharing platform integrates the computing engines, algorithms, models, and development processes required for the development of the aforementioned data products. It provides a unified data product design and development page, allowing business personnel to select key parameters such as data product type, required data source type and tables, data product computing engine, and data product development process through clicks. This method greatly simplifies the data product development process and improves development efficiency. Developers do not need to deeply understand the business logic of each data product, significantly lowering the barrier to entry for using the platform.

[0111] V. Railway Passenger Data Circulation and Sharing: For different internal and external data product service recipients, such as railway departments, cultural and tourism departments, financial service institutions, integrated transportation hubs, urban planning departments, and aviation departments, differentiated methods for railway passenger data circulation and sharing are provided based on their actual needs. These methods include: visual dashboards, real-time data streams, periodically refreshed materialized views, data interface calls, profiling system tags, and joint analysis model feature inputs and model parameters. For example, for the real-time passenger flow needs of railway stations in integrated transportation hubs, since the required calculation result set is accessed approximately every 15-30 seconds, real-time data interface calls are considered for data product circulation and sharing when providing railway passenger data services to them.

[0112] VI. Railway Passenger Data Update and Destruction: Different data product update strategies will be formulated based on factors such as the access frequency (e.g., once per hour), data granularity (e.g., passenger trips / stations / day), data scope (e.g., passenger travel trajectories over the past year), computational costs (including storage, bandwidth, memory, and CPU cores), and result set size (e.g., passenger characteristic tags with travel records in the past year) during the external service process. For some data products with high computational frequency, multiple versions will be retained for key indicator trend analysis. For data products with low access frequency, 1-3 historical versions will be retained. Taking financial service institutions' data products as an example, credit products require value assessment of users when granting credit limits. These user value level tags can usually be obtained offline through historical travel data and historical consumption data (e.g., tickets, catering, hotel accommodation, air-rail travel, etc.) of railway passengers over the past few years. Since short-term changes in user profile tags are not significant, the refresh frequency can be set to approximately one month, and historical tag data from versions 1-3 prior can be destroyed. This refresh and destruction strategy can be set on the platform page.

[0113] Part Two: Differentiated Development Process for Railway Passenger Data

[0114] In the context of internal and external circulation and sharing of railway data elements, this study collects, organizes, and analyzes the passenger transport data product needs of internal and external data product service recipients, including other railway departments, financial service institutions, cultural and tourism departments, integrated transportation hubs, and urban planning departments. These needs are categorized into three main types: statistical analysis data (such as train schedules, ticket purchase and refund information, station passenger flow input / output, and popular stations), model prediction data (such as passenger value ratings, future passenger flow and distribution predictions at stations, and passenger residence locations), and joint analysis data (such as identification of shared air-rail passengers and analysis of potential customers for air-rail intermodal transport). Differentiated designs are implemented for these three types of railway passenger transport data needs, proposing three different development and sharing processes. These three different railway passenger transport data development processes are all integrated into the data product development module of the railway passenger transport data development and sharing platform. Furthermore, all data sources, computing resources, execution engines, and algorithm models in the aforementioned railway passenger transport data development and sharing platform can be easily accessed. After analyzing and classifying the internal and external data product requirements for railway passenger transport, business personnel can conveniently and quickly complete the development and sharing of these three different types of railway passenger transport data on the same page through drop-down selection and simple parameter configuration. This is something that existing data product development models cannot achieve.

[0115] (1) Statistical analysis data.

[0116] Figure 4This is a flowchart of the development process for statistical analysis data provided in Embodiment 4 of the present invention. The development process includes two main parts: offline indicator calculation and real-time indicator calculation. First, based on the capabilities of the railway passenger transport data development and sharing platform, heterogeneous data from multiple sources, such as data files in the file system, embedded logs in the log collection system, data tables / Binlog operation logs in the real-time production database, and manually created Excel tables, undergoes standardization processing (dictionary table extraction, semantic unification, etc.). Simultaneously, noisy data is removed, and missing information is supplemented to form usable data assets for table entry. Subsequently, on the one hand, the real-time data extracted via streaming is processed through Flink CDC, and relying on the self-developed data element calculation logic library, it achieves simultaneous reading and calculation to obtain real-time production indicators, such as real-time passenger flow data for each waiting room in the station. On the other hand, data assets with lower real-time requirements are batch-loaded into Gbase or Hive data warehouses, and offline indicators are developed using the distributed storage and computing resources of the big data cluster, as well as self-developed UDF / UDAF functions, such as historical passenger flow input / output volumes for stations / cities. Finally, these offline / real-time metrics are written to the SSD storage medium of the Doris real-time cache database cluster and connected to the railway passenger data development and sharing platform in the form of data tables or materialized views for external data product services.

[0117] Taking the daily passenger flow input / output needs of various railway stations in Beijing as an example, the data product requirements can first be categorized as "statistical analysis type - offline calculation type". Then, the data source is selected as Doris / Gbase data warehouse, the data granularity is railway station / day level, the province and city is Beijing, and the date range is today. The platform will automatically match the calculation logic of station passenger flow input / output, analyze and count the number of people arriving / boarding at each railway station in Beijing today (in this process, the train ticket status will be automatically filtered, and invalid tickets such as those that have been refunded or changed will not be counted), and save the calculation results to the data product data table / materialized view defined on the platform page.

[0118] (2) Model predictive data:

[0119] Figure 5This is a flowchart of the development process for model predictive data provided in Embodiment 5 of the present invention. The development process of this data product relies on basic feature index data (such as passenger identity tags, passenger consumption preferences, historical passenger flow data, etc.) obtained through mathematical statistics. At the same time, it conducts extensive literature research and industry exchanges to build predictive models for different business needs. In this process, the extracted feature data needs to be numerically processed, and the model parameters are optimized by performing multiple rounds of model training on the training dataset. Finally, the Doris-generated feature indexes are used as model inputs to predict future passenger flow, passenger credit scores, passenger value levels, and other score indicators. The predicted indicators output by the model are loaded into the HBase distributed column-family database, and data services are provided through the railway passenger transport data development and sharing platform.

[0120] Taking the passenger value rating assessment needs of financial service institutions as an example, this data product requirement can be categorized as model prediction. First, a portion of passenger travel and consumption characteristic labels are extracted from the Doris data warehouse, numerically processed, and used as measurable model input. This is combined with the results of technical research and industry exchanges, such as the RFM model (a model that uses the time of the most recent purchase, purchase frequency, and purchase amount to measure user activity, user stickiness / loyalty, and user spending power), the CLV model (Customer Lifetime Value model, used to predict the future profit contribution of a customer group), and finally, the model Y value is set (for example, in the RFM model, the output R / F / M are divided into three levels 0, 1, 2, 3, and 4 to quantify user activity, loyalty, and spending power). A prediction model is built, a target prediction accuracy is set, and the model is trained using a test dataset, updating the model parameters until the accuracy target is achieved. Then, using the numerical feature labels of the production dataset as input, the model's RFM score is obtained. For example, (0,2,4) represents someone who has not recently made a purchase, has a moderate frequency of purchases in the past but strong purchasing power, and needs to be retained through targeted marketing. In the financial services industry (such as credit card limit granting), a higher initial limit can be considered.

[0121] (3) Joint analysis type:

[0122] Figure 6This is a flowchart illustrating the development process of jointly analyzed data provided in Embodiment Six of the present invention. The core of this data product development process lies in the use of privacy-preserving computation technology. Taking the joint analysis scenario of air-rail intermodal transport as an example, firstly, the railway department and the airline de-identify and encrypt their respective travel details data, and then select a portion of the encrypted data as the training set for the joint analysis model. In the stage of identifying co-passengers (referring to passengers who have purchased both train and plane tickets), multiple rounds of model training and gradient updates are performed through longitudinal federated learning, continuously adjusting and optimizing the local model parameters until the final co-passenger prediction value meets the accuracy requirements (generally, the success rate of co-passenger identification at this stage is required to be infinitely close to 100%). After identifying air-rail intermodal passengers, railway passenger transport indicators (i.e., railway passenger characteristic tags) and other industry indicators (such as airline passenger characteristic tags) are quantified. These, along with model Y values ​​obtained through technical surveys or expert discussions, are used as input to a joint analysis model. The trained local model is then used for analysis and prediction to obtain score outputs (e.g., 0 - not a potential air-rail intermodal customer, 1 - has not purchased an air-rail intermodal ticket but can be considered a potential customer for air-rail intermodal development, 2 - has already purchased an air-rail intermodal ticket and is a high-quality customer for whom promotion efforts can be intensified). This allows for the determination of whether these air-rail intermodal passengers are potential customers for air-rail intermodal tickets. Finally, these score data are loaded into the lake warehouse for airline subscription use.

[0123] It should be noted that the railway passenger data management method provided in this embodiment of the invention can be executed by an electronic device, a apparatus, or a control module within that apparatus for executing the method. This embodiment of the invention uses an apparatus executing the method as an example to illustrate the railway passenger data management apparatus provided in this embodiment of the invention.

[0124] Figure 7 This is a schematic diagram of the railway passenger data management device provided in Embodiment 7 of the present invention. The device 100 includes a preprocessing module 10, a statistical analysis module 20, a model prediction module 30, and a joint analysis module 40, wherein:

[0125] The preprocessing module 10 is used to preprocess diverse heterogeneous data to form railway passenger transport data, which is divided into statistical analysis data, model prediction data, and joint analysis data.

[0126] The statistical analysis module 20 is used to perform real-time or offline indicator calculations on statistical analysis data, and generate corresponding real-time or offline indicators for storage and sharing.

[0127] The model prediction module 30 is used to train and calculate a prediction model for model prediction data, and generate corresponding feature indicators for storage and sharing.

[0128] The joint analysis module 40 is used to analyze joint analysis data through a predetermined joint analysis model, generate corresponding score indicators for storage and sharing.

[0129] Figure 8 This is a schematic diagram of the railway passenger data management device provided in Embodiment 8 of the present invention. The device 100 includes a preprocessing module 10, a statistical analysis module 20, a model prediction module 30, a joint analysis module 40, a differential storage module 50, and / or a data update module 60, wherein:

[0130] The preprocessing module 10 is used to preprocess diverse heterogeneous data to form railway passenger transport data, which is divided into statistical analysis data, model prediction data, and joint analysis data.

[0131] Preferably, the preprocessing module 10 further performs the following actions:

[0132] The diverse and heterogeneous data is cleaned, transformed, and standardized to form railway passenger transport data.

[0133] Railway passenger transport data will be recorded as a data asset and data ownership will be confirmed.

[0134] The differentiated storage module 50 is used to classify and store railway passenger data separately as hot and cold based on the source, importance, timeliness, data volume, mining method, frequency of use, scope of use, and / or whether there is a joint query requirement.

[0135] The statistical analysis module 20 is used to perform real-time or offline indicator calculations on statistical analysis data, and generate corresponding real-time or offline indicators for storage and sharing.

[0136] Preferably, the statistical analysis module 20 further performs the following actions:

[0137] For statistical analysis data with real-time requirements, streaming extraction is performed, and real-time indicators are calculated using a streaming computing platform and a pre-defined data element calculation logic library to generate corresponding real-time indicators for storage and sharing.

[0138] For statistical analysis data that does not require real-time processing, the data warehouse can be loaded with distributed storage and computing resources of the big data cluster, as well as custom functions to perform offline indicator calculations, generate corresponding offline indicators for storage and sharing.

[0139] The model prediction module 30 is used to train and calculate a prediction model for model prediction data, and generate corresponding feature indicators for storage and sharing.

[0140] Preferably, the model prediction module 30 further performs the following actions:

[0141] A prediction model is designed and constructed using the extracted basic feature labels and the model Y value.

[0142] The basic feature labels are numerically processed by the prediction model, and the model parameters are optimized by training the training dataset in multiple rounds. The corresponding feature indicators are then generated for storage and sharing.

[0143] The joint analysis module 40 is used to analyze joint analysis data through a predetermined joint analysis model, generate corresponding score indicators for storage and sharing.

[0144] Preferably, the joint analysis module 40 further performs the following actions:

[0145] Railway passenger transport indicators and non-railway industry indicators are quantified and then used together with the model Y value as input to the joint analysis model.

[0146] The joint analysis model uses federated learning to perform multiple rounds of model training and gradient updates, generating corresponding score indicators for storage and sharing.

[0147] The data update module 60 is used to formulate different data update strategies based on the access frequency, data granularity, data range, computing cost and / or result set size of railway passenger data in the external service process, and update the railway passenger data according to the data update strategies.

[0148] The railway passenger data management device provided in this embodiment of the invention can achieve Figures 1-2 The various processes implemented in the railway passenger transport data management method embodiment shown are not described in detail here to avoid repetition.

[0149] The railway passenger transport data management device provided in this invention employs a differentiated development process designed for three types of railway passenger transport data characteristics: for statistical analysis data, real-time or offline indicator calculations are performed to generate corresponding real-time or offline indicators; for model prediction data, a prediction model is constructed, trained, and calculated to generate corresponding feature indicators; for joint analysis data, a predetermined joint analysis model is used for analysis to generate corresponding score indicators. Therefore, this invention improves the development efficiency, service quality, and commercial value of data products by constructing a systematic data management framework covering the entire lifecycle of railway passenger transport data.

[0150] The present invention also provides a storage medium for storing, for example, Figures 1-6A computer program for any of the railway passenger data management methods. For example, computer program instructions, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the present invention through the operation of the computer, achieving the same technical effect. To avoid repetition, these will not be elaborated further here. The program instructions for invoking the methods of the present invention may be stored in a fixed or removable storage medium, and / or transmitted via data streams in broadcast or other signal carrying media, and / or stored in the storage medium of a computer device operating according to the program instructions.

[0151] According to one embodiment of the present invention, the present invention also provides such a Figure 9 The illustrated electronic device 400 may optionally include a storage medium 200 for storing computer programs and a processor 300 for executing the computer programs. When the computer program is executed by the processor 300, it implements any of the aforementioned railway passenger data management methods, triggering the electronic device 400 to execute methods and / or technical solutions based on the foregoing embodiments, achieving the same technical effects. To avoid repetition, these will not be elaborated upon here. It should be noted that the electronic devices in this embodiment include mobile electronic devices and non-mobile electronic devices. For example, mobile electronic devices may be mobile phones, tablets, laptops, handheld computers, in-vehicle electronic devices, wearable devices, super mobile personal computers, netbooks, or personal digital assistants, etc., while non-mobile electronic devices may be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This embodiment does not specifically limit the scope of the invention.

[0152] It should be noted that the present invention can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of the present invention can be executed by a processor to implement the steps or functions described above. Similarly, the software program of the present invention (including associated data structures) can be stored in a computer-readable recording medium, such as RAM memory, a magnetic or optical drive, a floppy disk, or similar devices. Furthermore, some steps or functions of the present invention can be implemented in hardware, for example, as circuitry that works with a processor to perform the various steps or functions.

[0153] This invention can be implemented on a computer as a computer-based method, or in dedicated hardware, or a combination of both. Executable code or portions thereof for the method according to the invention can be stored on a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Optionally, the computer program product includes non-transitory program code components stored on a computer-readable medium so as to execute the method according to the invention when the program product is executed on a computer.

[0154] In an optional embodiment, the computer program includes computer program code components adapted to perform all the steps of the method according to the invention when the computer program is run on a computer. Optionally, the computer program is embodied on a computer-readable medium.

[0155] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0156] Of course, the present invention may have other various embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.

Claims

1. A method for managing railway passenger transport data, characterized in that, include: The preprocessing step involves preprocessing the diverse and heterogeneous data to form railway passenger transport data, which is divided into statistical analysis data, model prediction data, and joint analysis data. The statistical analysis step involves calculating real-time or offline indicators for the statistical analysis data, generating corresponding real-time or offline indicators for storage and sharing. The model prediction step involves, for the model prediction data, constructing a prediction model for training and computation, generating corresponding feature indicators for storage and sharing; The joint analysis step involves analyzing the joint analysis data using a predetermined joint analysis model to generate corresponding score indicators for storage and sharing.

2. The method for managing railway passenger transport data according to claim 1, characterized in that, The preprocessing step further includes: The diverse and heterogeneous data is cleaned, transformed, and standardized to form the railway passenger transport data. The railway passenger transport data is recorded as a data asset in the table, and data ownership is confirmed for the data asset.

3. The method for managing railway passenger data according to claim 1, characterized in that, The statistical analysis steps further include: For statistical analysis data with real-time requirements, streaming extraction is performed, and real-time indicators are calculated using a streaming computing platform and a pre-defined data element calculation logic library to generate corresponding real-time indicators for storage and sharing. For statistical analysis data that does not require real-time processing, the data is loaded into the data warehouse. The distributed storage and computing resources of the big data cluster, as well as custom functions, are used to calculate offline indicators and generate corresponding offline indicators for storage and sharing.

4. The method for managing railway passenger data according to claim 1, characterized in that, The model prediction step further includes: The prediction model is designed and constructed using the extracted basic feature labels and the model Y value; The prediction model quantifies the basic feature labels, optimizes the model parameters by performing multiple rounds of model training on the training dataset, and generates corresponding feature indicators for storage and sharing.

5. The method for managing railway passenger data according to claim 1, characterized in that, The joint analysis step further includes: The railway passenger transport indicators and non-railway industry indicators are numerically processed and then used together with the model Y value as input to the joint analysis model. The joint analysis model performs multiple rounds of model training and gradient updates through federated learning to generate corresponding score indicators for storage and sharing.

6. The method for managing railway passenger transport data according to claim 1, characterized in that, Following the preprocessing step, the following also includes: The differentiated storage step involves classifying and separating the railway passenger data into hot and cold categories based on its source, importance, timeliness, data volume, mining method, frequency of use, scope of use, and / or whether there is a need for joint queries.

7. The method for managing railway passenger transport data according to claim 1, characterized in that, The method further includes: The data update step involves formulating different data update strategies based on factors such as the access frequency, data granularity, data range, computational cost, and / or result set size of the railway passenger data during external service processes, and updating the railway passenger data according to the data update strategies.

8. A railway passenger transport data lifecycle management device constructed based on the method described in any one of claims 1 to 7, characterized in that, The device includes: The preprocessing module is used to preprocess diverse and heterogeneous data to form railway passenger transport data, which is divided into statistical analysis data, model prediction data, and joint analysis data. The statistical analysis module is used to perform real-time or offline indicator calculations on the statistical analysis data, and generate corresponding real-time or offline indicators for storage and sharing. The model prediction module is used to train and calculate a prediction model for the model prediction data, and generate corresponding feature indicators for storage and sharing. The joint analysis module is used to analyze the joint analysis data using a predetermined joint analysis model, and generate corresponding score indicators for storage and sharing.

9. A storage medium, characterized in that, Used to store a computer program for performing the method according to any one of claims 1 to 7.

10. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.