Rail transit data asset service directory construction method and system
Patent Information
- Application Number
- CN202510844448.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-06-23
AI Technical Summary
[0004]1、大多数轨道交通数据资产标签体系是通用的,缺乏针对轨道交通行业的专门适配,导致轨道交通数据资产分类粗糙,无法满足轨道交通行业特定的数据管理需求
[0067] The method and system for constructing a data asset service catalog for rail transit of the present invention can achieve the following:
Smart Images

Figure CN120821721B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and in particular to a method and system for constructing a data asset service catalog for rail transit. Background Technology
[0002] Rail transit data assets refer to data resources, recorded in physical or electronic form, owned or managed by enterprises or institutions that can bring them future economic benefits. In the rail transit industry, with the continuous improvement of data collection, storage, and analysis capabilities, massive amounts of data have been generated across multiple scenarios such as operation management, equipment monitoring, and passenger flow analysis. This data covers multiple dimensions, including train operation status, passenger behavior, dispatching information, equipment health status, and ticketing transactions.
[0003] However, inconsistencies in the data structure, naming conventions, and classification standards of rail transit data lead to fragmented data management, making it difficult to provide unified data service sharing and establish a unified data asset service catalog. Specifically:
[0004] 1. Most rail transit data asset labeling systems are generic and lack specific adaptations for the rail transit industry, resulting in crude classification of rail transit data assets that cannot meet the specific data management needs of the rail transit industry.
[0005] 2. The use of the labeling system relies on human experience and lacks unified rules, which can easily lead to inconsistencies or errors in labeling during intelligent labeling.
[0006] 3. Most existing intelligent labeling methods only assign labels and cannot record key information such as sensitive information, row and column range, and data volume of rail transit data assets, resulting in insufficient subsequent data management and security control.
[0007] 4. In the process of data sharing, existing methods usually rely on manual control of desensitization, encryption or approval, which is inefficient, error-prone and unable to achieve fine control at the row and column level.
[0008] 5. Existing data queries rely on SQL or keyword searches, requiring users to understand the database structure, making it difficult to perform semantic-level intelligent retrieval.
[0009] In the prior art, such as Chinese invention patent with publication number CN119622556A, an intelligent method and system for managing traffic asset data is disclosed. This includes: performing multi-dimensional analysis to determine data classification target labels, constructing a machine learning model to obtain an intelligent data classification model, identifying and classifying asset data, generating data labels to annotate data in the traffic asset database; establishing a response correlation between traffic asset data and data application scenarios, and performing data label dependency and transparency analysis to obtain data call constraints; searching and filtering data in the traffic asset database based on the data labels in the data call constraints, and outputting scenario application data. This invention solves the technical problem in the prior art where the failure to pre-define the correlation between data and application scenarios during data management leads to complex search condition settings, inaccurate data matching, and omissions in data querying and application, thus improving data management efficiency. However, the aforementioned prior art does not disclose how to accurately locate the corresponding data asset service catalog through natural language queries. Summary of the Invention
[0010] To address the technical problems existing in the prior art, the present invention aims to provide a method and system for constructing a data asset service catalog for rail transit, which can realize intelligent labeling of rail transit data assets and construct a data asset service catalog based on rail transit data.
[0011] To achieve the above-mentioned objectives, this invention provides a method for constructing a data asset service catalog for rail transit, comprising the following steps:
[0012] Based on rail transit data rules, perform corresponding rule matching on rail transit data sources;
[0013] The rail transit data rules and corresponding rail transit data tags are pre-stored in the rail transit data rule library;
[0014] When the rail transit data source matches the rail transit data rules, the rail transit data table corresponding to the rail transit data source is assigned a corresponding rail transit data asset tag.
[0015] Based on the preset data asset service catalog template, the rail transit data table and the corresponding rail transit data asset tags are integrated to construct the data asset service catalog.
[0016] According to one technical solution of the present invention, the rail transit data source includes at least one level of rail transit database, rail transit data pattern, rail transit data table, rail transit data field and rail transit data field value;
[0017] The rail transit data rules include regular expression rules, keyword rules, dimension combination rules, and condition triggering rules;
[0018] The expression rules include: rules for pattern matching of rail transit data field names, rail transit data table names, and rail transit data field values;
[0019] The keyword rules include: applying matching rules to remarks, field comments, database names, and table names in rail transit data tables;
[0020] The dimension combination rules include: rules that match the field names of rail transit data, the field types in the rail transit data table, and the remarks.
[0021] The conditional triggering rules include: rules that use the value of a rail transit data field conforming to a sensitive information pattern as the triggering condition;
[0022] The specific steps for performing rule matching on the rail transit data source are as follows:
[0023] Select the required rail transit data tags and corresponding rail transit data rules from the rail transit data rule library;
[0024] The rail transit data rules are combined using regular expression symbols to obtain combined rail transit data rules;
[0025] Based on the combined rail transit data rules, rule matching is performed on one or more levels of rail transit database, rail transit data pattern, rail transit data table, rail transit data field, and rail transit data field value.
[0026] According to one technical solution of the present invention, it further includes:
[0027] When the rail transit data source matches the rail transit data rules, the sensitive information, the range of rows and columns of the sensitive information, and the amount of sensitive information in the rail transit data are recorded based on the rail transit data tags corresponding to the sensitive information.
[0028] Different security policies are configured based on the different levels of sensitive information in the data asset service catalog; the different levels of sensitive information correspond to different rail transit data tags.
[0029] The security policy includes approval, data anonymization, data encryption, and access control for the range of sensitive information rows and / or the amount of sensitive information at this level.
[0030] According to one technical solution of the present invention, it further includes:
[0031] Based on the general large language model, the corresponding rail transit data assets are obtained by querying the data asset service catalog.
[0032] According to a technical solution of the present invention, based on a general large language model, the corresponding rail transit data assets are retrieved from the data asset service catalog, and the steps are as follows:
[0033] Vectorize the rail transit data assets in the data asset service catalog;
[0034] Extract the semantics of the input query text;
[0035] Based on the semantics, the most similar rail transit data assets are recalled from the vectorized rail transit data assets as the initial recall result;
[0036] Based on the semantic matching of the rail transit data asset tags in the initial recall results, the rail transit data assets corresponding to the rail transit data asset tags that meet the matching rules are retained as candidate rail transit data assets.
[0037] Based on the candidate rail transit data assets, prompt words are constructed, and corresponding term explanations are obtained based on a preset term database;
[0038] The candidate rail transit data assets, the prompt words, and the terminology explanations are input into a general large language model, and a candidate rail transit data asset is output as the final recommended rail transit data asset.
[0039] The present invention also provides a data asset service catalog construction system for rail transit, including a rule matching module for performing corresponding rule matching on rail transit data sources based on rail transit data rules;
[0040] A rail transit data rule library is used to pre-store the rail transit data rules and corresponding rail transit data tags;
[0041] The tagging module is used to assign corresponding rail transit data asset tags to the rail transit data table corresponding to the rail transit data source when the rail transit data source matches the rail transit data rules.
[0042] The data integration module is used to integrate the rail transit data table and the corresponding rail transit data asset tags to construct a data asset service catalog based on a preset data asset service catalog template.
[0043] According to one technical solution of the present invention, the rail transit data source includes at least one level of rail transit database, rail transit data pattern, rail transit data table, rail transit data field and rail transit data field value;
[0044] The rail transit data rules include regular expression rules, keyword rules, dimension combination rules, and condition triggering rules;
[0045] The expression rules include: rules for pattern matching of rail transit data field names, rail transit data table names, and rail transit data field values;
[0046] The keyword rules include: applying matching rules to remarks, field comments, database names, and table names in rail transit data tables;
[0047] The dimension combination rules include: rules that match the field names of rail transit data, the field types in the rail transit data table, and the remarks.
[0048] The conditional triggering rules include: rules that use the value of a rail transit data field conforming to a sensitive information pattern as the triggering condition;
[0049] The rule matching module includes:
[0050] The rule selection module is used to select the required rail transit data tags and corresponding rail transit data rules from the rail transit data rule library;
[0051] The rule combination module is used to combine the rail transit data rules using regular expression symbols to obtain combined rail transit data rules;
[0052] The hierarchical matching module is used to perform rule matching on one or more levels of the rail transit database, rail transit data pattern, rail transit data table, rail transit data field, and rail transit data field value based on the combined rail transit data rules.
[0053] According to one technical solution of the present invention, the labeling module is further configured to, when the rail transit data source matches the rail transit data rules, record the sensitive information, the range of rows and columns of the sensitive information, and the amount of sensitive information in the rail transit data based on the rail transit data labels corresponding to the sensitive information;
[0054] The construction system also includes:
[0055] The permission configuration module is used to configure different security policies based on the different levels of sensitive information in the data asset service catalog; the different levels of sensitive information correspond to different rail transit data tags.
[0056] The security policy includes approval, data anonymization, data encryption, and access control for the range of sensitive information rows and / or the amount of sensitive information at this level.
[0057] According to one technical solution of the present invention, it further includes:
[0058] The asset query module is used to retrieve the corresponding rail transit data assets from the data asset service catalog based on the general large language model.
[0059] According to one technical solution of the present invention, the asset query module includes:
[0060] The asset vectorization module is used to vectorize rail transit data assets in the data asset service catalog;
[0061] The query text vectorization module is used to extract the semantics of the input query text;
[0062] The initial recall module is used to recall multiple rail transit data assets with the highest similarity from the vectorized rail transit data assets based on the semantics, as the initial recall result;
[0063] The asset screening module is used to match the rail transit data asset tags in the initial recall results based on the semantics, and retain the rail transit data assets corresponding to the rail transit data asset tags that meet the matching rules as candidate rail transit data assets.
[0064] An enhanced prompt word construction module is used to construct prompt words based on the candidate rail transit data assets and obtain corresponding term explanations based on a preset term database;
[0065] A generalized large language model is used to input the candidate rail transit data assets, the prompt words, and the terminology explanations into the generalized large language model, and output a candidate rail transit data asset as the final recommended rail transit data asset.
[0066] Compared with the prior art, the present invention has the following advantages:
[0067] The method and system for constructing a data asset service catalog for rail transit of the present invention can achieve the following:
[0068] 1. The standardization of the tagging system provides a foundation for intelligent tagging.
[0069] 2. Intelligent labeling of rail transit data assets, and the intelligent labeling is not limited to just labeling, but also records more key information of rail transit data assets.
[0070] 3. Intelligent labeling and hierarchical classification of rail transit data assets, and construction of a data asset service catalog based on rail transit data.
[0071] 4. A closed loop for hierarchical data security sharing: When sharing data services, the system will automatically perform desensitization, encryption, approval control, and row / column permission restrictions based on the information recorded by intelligent annotation.
[0072] 5. The combination of understanding large models and tag systems enables accurate location of the corresponding data asset service catalog through natural language queries. Attached Figure Description
[0073] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0074] Figure 1 The diagram illustrates the principle of a method for constructing a data asset service catalog for rail transit according to an embodiment of the present invention. Detailed Implementation
[0075] The description of the embodiments in this specification should be taken in conjunction with the accompanying drawings, which should form part of the complete specification. In the drawings, the shape or thickness of the embodiments may be exaggerated and may be indicated in a simplified or convenient manner. Furthermore, parts of the various structures in the drawings will be described separately; it is worth noting that elements not shown in the figures or not described in words are in a form known to those skilled in the art.
[0076] The descriptions of the embodiments herein, including any references to directions and orientations, are for ease of description only and should not be construed as limiting the scope of the invention. The following description of preferred embodiments involves combinations of features, which may exist independently or in combination; the invention is not particularly limited to the preferred embodiments. The scope of the invention is defined by the claims.
[0077] like Figure 1 As shown, the present invention provides a method for constructing a data asset service catalog for rail transit, comprising the following steps:
[0078] Based on rail transit data rules, perform corresponding rule matching on rail transit data sources;
[0079] Rail transit data rules and corresponding rail transit data tags are pre-stored in the rail transit data rule database;
[0080] When the rail transit data source matches the rail transit data rules, the corresponding rail transit data asset label is assigned to the rail transit data table corresponding to the rail transit data source.
[0081] Based on the preset data asset service catalog template, the data asset service catalog is constructed by integrating the rail transit data table and the corresponding rail transit data asset tags.
[0082] In this embodiment, a rail transit data asset labeling system is defined to construct a rail transit data asset labeling system based on data hierarchical management. According to the requirements of rail transit data classification and hierarchical management and data governance, combined with the business logic and technical characteristics of the data, an urban rail transit data asset labeling system is constructed.
[0083] The rail transit data labeling system includes three main categories: data business labels, data technology labels, and data management labels.
[0084] (1) Data service labels are divided into data hierarchical labels and data classification labels according to the characteristics of rail transit services.
[0085] Data classification labels mainly describe the categories and characteristics of rail transit data;
[0086] Data classification labels mainly describe the security level and external accessibility of rail transit data;
[0087] (2) Data technology tags, which describe the specific types or attributes used in each stage of the entire data lifecycle based on the characteristics of rail transit data technology, are divided into three categories: data collection, data storage, and data processing.
[0088] Data collection-related tags mainly include the frequency of rail transit data collection, acquisition method, data timeliness, and data source;
[0089] Data storage tags mainly include rail transit data storage cycle, data upload time, and data type;
[0090] Data processing tags mainly include the processing type of rail transit data, time granularity, spatial dimension, ticket type dimension, whether to distinguish statistical caliber, and indicator type.
[0091] (3) The data management label mainly describes the service targets and approval levels of rail transit data.
[0092] Each tag has its own attributes defined based on business needs and technical characteristics. These attributes mainly include basic information, business attributes, and technical attributes.
[0093] The metadata attributes of the tag are as follows:
[0094] (1) Basic information mainly describes the label’s naming, encoding, and attribution, including label name, label encoding, category, and category encoding;
[0095] 1) Tag Name: This mainly describes the specific name of the tag;
[0096] 2) Label Encoding: Primarily describes the standard encoding of the label;
[0097] 3) Category: Primarily describes the name of the parent category to which the tag belongs;
[0098] 4) Category Code: Primarily describes the code of the parent category name to which the tag belongs;
[0099] (2) Business attributes are mainly used to clarify the specific purpose and significance of the label in the business scenario, including label definition, label value range, business description, label application scope, and label annotation method;
[0100] 1) Tag Definition: This section mainly describes and explains the specific concepts behind the tag.
[0101] 2) Tag Values: Primarily describes the collection of specific tag value names;
[0102] 3) Business Description: This section mainly describes the business logic and specifications of the tags and their values.
[0103] 4) Label Basis: This section mainly describes the source of the label and the basis for setting the label;
[0104] 5) Scope of application of the label: This mainly describes the scope of the object that the label is primarily used to annotate. For example, the data collection label is suitable for annotating raw data.
[0105] (3) Technical attributes refer to the various technical characteristics used to describe and define labels, including label labeling methods, label labeling rules, and label labeling cycles;
[0106] 1) Labeling method: This mainly describes the specific labeling method, including intelligent labeling and manual labeling;
[0107] 2) Tag Labeling Rules: This section mainly describes the specific technical rules for tag labeling, including but not limited to keywords, regular expressions, etc.
[0108] 3) Labeling cycle: This mainly describes the labeling cycle, i.e. how often the label is labeled.
[0109] The above-mentioned rail transit data asset labels are shown in Tables 1 and 2 below, which are examples of data classification labels and data grading labels in data business labels (the corresponding rule formulas are omitted, but based on existing technology, the corresponding rule formulas can be constructed according to the required label categories).
[0110] Table 1 Examples of Data Category Tags in Data Service Tags
[0111] Tag metadata attributes:
[0112] Tag Name: Subject Perspective Tag Code: YW_FL_SJZT
[0113] Category: Data Classification; Category Code: YW_FL
[0114] Business attributes:
[0115]
[0116]
[0117] Technical attributes:
[0118]
[0119] Table 2. Examples of Data Classification Tags in Data Service Tags. Tag Metadata Attributes: Tag Name: Security Level; Tag Code: YW_FJ_AQDJ; Category: Data Classification; Category Code: YW_FJ; Service Attributes:
[0120]
[0121] Technical attributes:
[0122]
[0123]
[0124] 2. Intelligent Labeling Based on Technical Attributes: Intelligent labeling tasks are constructed by combining labeling methods, rules, and cycles based on the technical attributes of the labels. This allows for the free combination of keywords and regular expressions to label data assets from multiple dimensions, including data sources, databases, tables, fields, data, and remarks. Simultaneously, the labeling process extracts and records hierarchical feature information of the rail transit data assets, such as sensitive information, row and column ranges, and data volume. Upon completion, the basis for labeling the rail transit data assets is automatically generated, ensuring the labeling process is traceable and verifiable, providing a reliable basis for data sharing and application (e.g., in table A, the number of data rows matching label 1 rule is 100).
[0125] This implementation defines a rail transit data asset tagging system, which is designed specifically for the rail transit industry and can adapt to specific data types such as train operation, passenger flow, equipment status, and dispatch records.
[0126] Establish standardized labeling specifications to improve the consistency and reusability of rail transit data asset management.
[0127] Based on the tags constructed above, intelligent annotation is performed as follows:
[0128] (1) Based on the constructed tag system, construct regular expression matching rules;
[0129] (2) Construct scheduling tasks, configure the scanning range of the annotations and the scheduling period;
[0130] (3) Configure specific labeling rules under the scheduling task;
[0131] (4) Collect data metadata / data samples;
[0132] (5) Perform regular expression matching or keyword search on the aforementioned collected structural information and field values;
[0133] (6) Tag hit calculation;
[0134] (7) Assign a label to the hit table and record information.
[0135] The rail transit data source includes at least one level of rail transit database, rail transit data schema, rail transit data table, rail transit data field, and rail transit data field value;
[0136] Rail transit data rules include regular expression rules, keyword rules, dimension combination rules, and condition triggering rules;
[0137] The expression rules include: rules for pattern matching of rail transit data field names, rail transit data table names, and rail transit data field values;
[0138] Keyword rules include: applying rules that include matching to remarks in rail transit data tables, field comments in rail transit data tables, rail transit database names, and rail transit data table names;
[0139] The dimension combination rules include: rules that match the field names of rail transit data, the field types in the rail transit data table, and the remarks.
[0140] Conditional triggering rules include: rules that use the value of a rail transit data field conforming to a sensitive information pattern as a triggering condition;
[0141] In the method for constructing the data asset service catalog for rail transit, corresponding rule matching is performed on the rail transit data sources. The specific steps are as follows:
[0142] Select the required rail transit data labels and corresponding rail transit data rules from the rail transit data rule library;
[0143] Regular expression symbols are used to combine rail transit data rules to obtain combined rail transit data rules;
[0144] Based on combined rail transit data rules, rule matching is performed on one or more levels of rail transit database, rail transit data pattern, rail transit data table, rail transit data field, and rail transit data field value.
[0145] This embodiment mainly involves the detailed process of intelligent annotation, as follows:
[0146] (1) Based on the constructed tag system, construct regular expression matching rules:
[0147] It supports the construction of various types of matching rules, which can be combined in the form of tasks;
[0148] Regular expression rules: perform pattern matching on field names, table names, and field values;
[0149] Keyword rules: Perform containment matching for comments, field comments, library / table names, etc.
[0150] Dimension combination rules: For example, only fields whose names, types, and comments all match are marked.
[0151] Conditional triggering rules: such as data volume exceeding a threshold, field values conforming to sensitive information patterns.
[0152] (2) Construct the scheduling task, configure the scanning range and scheduling period for the annotations:
[0153] It can be configured with various structured data sources (such as MySQL, PostgreSQL, Hive, GaussDB and other mainstream databases), and can customize the scan level and scan cycle according to four levels: data source -> database -> schema (PostgreSQL / GaussDB) -> data table.
[0154] (3) Configure specific annotation rules under the scheduling task:
[0155] Rule selection: Select the tag to be applied and its regular expression from the rule base;
[0156] Rule combination: Supports matching multiple rule combinations; Example: It needs to satisfy the condition that the table name is prefixed with A and contains field B, thus matching tag C;
[0157] Sampling strategy (field value matching): Select whether to enable field value sampling and the sampling range (full sample / sampling).
[0158] (4) Collect data metadata / data samples:
[0159] Connect to the data source selected by the scheduling task, and extract structured metadata and data samples, including:
[0160] Database, schema, tables, field structure and their comments, types, indexes, etc.;
[0161] Table data: N records (full or sampled) are used for field value rule matching.
[0162] (5) Perform regular expression matching or keyword search on the aforementioned collected structural information and field values:
[0163] 1) Database name / schema / table name / table name / field name matching: whether it contains or fully matches a certain tag rule;
[0164] 2) Field value matching: If no field is specified, all fields in the row record will be matched for the tag.
[0165] Precise field matching: specify field values to complete tag matching;
[0166] 3) Note matching: Whether the tag's key terms appear in the notes;
[0167] (6) Tag hit calculation:
[0168] The hit rate is calculated according to the configuration rules. This step is used to calculate the completeness of the matching rules for subsequent improvement and modification, but it is not a necessary condition for intelligent labeling.
[0169] (7) Assign labels to the hit table and record information:
[0170] 1) Assign labels to the hit table;
[0171] 2) Record information: the source of the final tag hit for each field / table (field name, value, remarks), matching details, rule information, and the basis for the hit; at the same time, the hierarchical feature information of rail transit data assets is extracted during the annotation process, such as the sensitive information of rail transit data assets, row and column range, data volume, sample row data, etc.
[0172] The aforementioned intelligent annotation, through a rule-based labeling system, provides executable standards for automated annotation, improving annotation accuracy and consistency. Furthermore, through rule definition, intelligent annotation enables efficient, batch data processing, reducing manual intervention.
[0173] Based on the intelligently labeled rail transit data assets mentioned above, a data asset service catalog is constructed. The data asset service catalog mainly includes basic information, a data asset service catalog tree, and data service APIs.
[0174] (1) The basic information of the data service catalog includes: asset name, asset type, asset description, asset number, responsible department, regulatory department, asset-related data table information, rail transit data asset tag information, etc.
[0175] (2) The data asset service catalog tree is mainly constructed by referring to the hierarchical classification labels to build a dual catalog tree, which facilitates quick location of the data service catalog. The hierarchical catalog tree of the data asset service is constructed according to the hierarchical labels, and the classification catalog tree of the data asset service is constructed according to the classification labels.
[0176] (3) Data service API information: API information exposed to the outside world based on the data service catalog, mainly including HTTP request method, HTTP request path, request input parameters, request returned data content, etc.
[0177] The method for constructing a data asset service catalog for rail transit also includes:
[0178] When the rail transit data source matches the rail transit data rules, the sensitive information, the range of rows and columns of the sensitive information, and the amount of sensitive information in the rail transit data are recorded based on the rail transit data labels corresponding to the sensitive information.
[0179] Configure different security policies based on the different levels of sensitive information in the data asset service catalog; different levels of sensitive information correspond to different rail transit data tags;
[0180] The security strategy includes approval, data anonymization, data encryption, and access control for the range of sensitive information and / or the amount of sensitive information at this level.
[0181] In this implementation, a hierarchical system is used to achieve secure control over data sharing.
[0182] When sharing and utilizing rail transit data assets, a hierarchical system based on the security level, approval level, and openness of rail transit data assets is used. Combined with intelligent annotation to record key information such as sensitive information, row and column range, and data volume of rail transit data assets, approval control, desensitization, encryption, and row and column permission control are automatically completed.
[0183] No longer relying on manual identification of sensitive fields, the content is automatically classified through a labeling mechanism.
[0184] Security policies (approval, desensitization, encryption) are automatically configured based on tags and levels, requiring no manual maintenance.
[0185] Furthermore, the combination of row-level and column-level permissions enables field-level control rather than just table-level control. This achieves a balance between security and data availability, improving sharing efficiency while ensuring audit compliance.
[0186] The key information of the rail transit data assets recorded synchronously during the annotation process is as follows:
[0187] Does the data contain sensitive information (name, ID number, GPS coordinates, etc.)?
[0188] Specifically, the fields involved, the number of data rows, and the amount of data;
[0189] Access rights and compliance requirements for rail transit data assets;
[0190] The above data provides a basis for subsequent data security management and access control, and enhances data governance capabilities.
[0191] Based on intelligent annotation information, security policies are automatically executed as follows:
[0192] Desensitization: Automatically masking or replacing sensitive fields (such as hiding parts of ID card numbers or mobile phone numbers);
[0193] Encryption: Encrypt highly sensitive data during storage to ensure secure transmission;
[0194] Approval control: Automatically triggers approval processes for high-privilege data;
[0195] Row-level access control: Set access permissions based on user roles to prevent unauthorized access to data;
[0196] By establishing a closed-loop mechanism for data security management, the compliance and security of data sharing can be ensured, thereby enhancing the company's ability to protect its rail transit data assets.
[0197] The method for constructing a data asset service catalog for rail transit also includes:
[0198] Based on the general large language model, the corresponding rail transit data assets are retrieved from the data asset service catalog.
[0199] Based on a general large language model, the corresponding rail transit data assets are retrieved from the data asset service catalog. The steps are as follows:
[0200] Vectorize the rail transit data assets in the data asset service catalog;
[0201] Extract the semantics of the input query text;
[0202] Based on semantics, the most similar rail transit data assets are recalled from vectorized rail transit data assets as the initial recall result;
[0203] Based on the initial recall results of semantic matching, the rail transit data asset tags that meet the matching rules are retained as candidate rail transit data assets.
[0204] Prompt words are constructed based on candidate rail transit data assets, and corresponding term explanations are obtained based on a pre-set term database;
[0205] The candidate rail transit data assets, prompts, and terminology explanations are input into a general large language model, and a candidate rail transit data asset is output as the final recommended rail transit data asset.
[0206] In this implementation, a locally deployed open-source large-scale model is used for fine-tuning training with key information extracted during labeling and annotation, enabling users to quickly query rail transit data assets using natural language. PromptTuning is employed to ensure the large-scale model accurately understands and applies the rail transit data asset labeling system. This approach abandons the existing method of constructing multi-condition queries, instead using natural language input. By having the large-scale model parse the query semantics, it converts natural language into SQL (SQL, or SQL query language) for efficient querying. Large-scale model query process:
[0207] (1) Constructing a knowledge graph / structure library for rail transit data assets:
[0208] Each rail transit data asset is defined by the following dimensions:
[0209] 1) The table information structure includes table name, field name, data type, field description, sample value, Chinese name, field description, and sensitivity level;
[0210] 2) The tag information system includes topic tags (such as "user behavior"), function tags (such as "data timeliness"), and sensitive tags (such as "user information" and "train information");
[0211] 3) Asset attribute metadata includes data source, update time, responsible person, whether it can be shared, and data volume;
[0212] (2) Vectorize the information and store it in a vector database:
[0213] 1) Construct an embedding vector from the asset's summary information;
[0214] 2) Store in a vector database
[0215] (3) User query vectorization and tag pre-screening:
[0216] 1) The system transforms the user's question into a fixed-dimensional vector using an embedding model, representing the semantic content of the question;
[0217] Returning to the Top-K assets, only the top K assets (e.g., the top 5 to 10 assets) that are closest to the semantics of the question are retained. Subsequent analysis and recommendations are only performed on these candidate assets. This avoids the model going astray by analyzing a large number of irrelevant assets and provides highly relevant contextual material for the generation of prompt words.
[0218] 2) Match the returned asset's tag list (e.g., "Line 7", "Entry / Exit" tags).
[0219] 3) Quickly filter out irrelevant rail transit data assets to improve accuracy.
[0220] (4) Construct an enhanced Prompt+ terminology completion tool.
[0221] Based on the recall results, construct a Prompt:
[0222] User question: How to query asset information for Line 7 stations?
[0223] Candidate Assets:
[0224] 1. Table: seven_line_exit_data, Tags: Line 7, Entry / Exit, Table name: seven_line, Fields: line_no, exit_num...
[0225] 2. Table: all_line_exit_data, Tags: All Lines, Statistics, Inbound / Outbound Stations, Fields: line_no, exit_num....
[0226] Terminology explanation: "Entry and exit" refers to the number of subway lines entering and exiting stations on a given day;
[0227] Problem analysis: The "Line 7" should aggregate the table name "seven_line", which contains an "exit_num" field;
[0228] Recommendation reason (generated from natural language by a large model, such as "the exit_num field of this table can be used to count the number of trains entering and leaving the station");
[0229] The large model combines table structure, labels, and field matching degree to perform inference and ranking, and recommends the asset results associated with seven_line_exit_data;
[0230] (5) Output recommended assets + matching explanation;
[0231] The output includes asset metadata, table names and field lists, and tag matching information.
[0232] The present invention provides a data asset service catalog construction system for rail transit, comprising:
[0233] The rule matching module is used to perform corresponding rule matching on rail transit data sources based on rail transit data rules;
[0234] The rail transit data rule library is used to pre-store rail transit data rules and corresponding rail transit data tags;
[0235] The labeling module is used to assign corresponding rail transit data asset labels to the rail transit data table corresponding to the rail transit data source when the rail transit data source matches the rail transit data rules.
[0236] The data integration module is used to integrate rail transit data tables and corresponding rail transit data asset tags to construct a data asset service catalog based on a preset data asset service catalog template.
[0237] In the data asset service catalog construction system for rail transit, the rail transit data source includes at least one level of rail transit database, rail transit data schema, rail transit data table, rail transit data field, and rail transit data field value;
[0238] Rail transit data rules include regular expression rules, keyword rules, dimension combination rules, and condition triggering rules;
[0239] The expression rules include: rules for pattern matching of rail transit data field names, rail transit data table names, and rail transit data field values;
[0240] Keyword rules include: applying rules that include matching to remarks in rail transit data tables, field comments in rail transit data tables, rail transit database names, and rail transit data table names;
[0241] The dimension combination rules include: rules that match the field names of rail transit data, the field types in the rail transit data table, and the remarks.
[0242] Conditional triggering rules include: rules that use the value of a rail transit data field conforming to a sensitive information pattern as a triggering condition;
[0243] The rule matching module includes:
[0244] The rule selection module is used to select the required rail transit data labels and corresponding rail transit data rules from the rail transit data rule library;
[0245] The rule combination module is used to combine rail transit data rules using regular expression symbols to obtain combined rail transit data rules;
[0246] The hierarchical matching module is used to perform rule matching on one or more levels of rail transit databases, rail transit data patterns, rail transit data tables, rail transit data fields, and rail transit data field values based on combined rail transit data rules.
[0247] In the data asset service catalog construction system for rail transit, the labeling module is also used to record sensitive information, the range of rows and columns of sensitive information, and the amount of sensitive information in rail transit data based on the rail transit data labels corresponding to sensitive information, when the rail transit data source and rail transit data rules are matched.
[0248] The system also includes:
[0249] The permission configuration module is used to configure different security policies based on the different levels of sensitive information in the data asset service catalog; different levels of sensitive information correspond to different rail transit data tags.
[0250] The security strategy includes approval, data anonymization, data encryption, and access control for the range of sensitive information and / or the amount of sensitive information at this level.
[0251] The data asset service catalog construction system for rail transit also includes:
[0252] The asset query module is used to retrieve the corresponding rail transit data assets from the data asset service catalog based on a general large language model.
[0253] The asset query module in the rail transit data asset service catalog construction system includes:
[0254] The asset vectorization module is used to vectorize rail transit data assets in the data asset service catalog;
[0255] The query text vectorization module is used to extract the semantics of the input query text;
[0256] The initial recall module is used to recall multiple rail transit data assets with the highest similarity from vectorized rail transit data assets based on semantics, as the initial recall result;
[0257] The asset screening module is used to select rail transit data assets based on the initial recall results of semantic matching, and retain the rail transit data assets corresponding to the rail transit data asset tags that meet the matching rules as candidate rail transit data assets.
[0258] The enhanced prompt word construction module is used to construct prompt words based on candidate rail transit data assets and obtain corresponding term explanations based on a preset term database;
[0259] A generalized large language model is used to input candidate rail transit data assets, prompt words, and terminology explanations into the generalized large language model, and output a candidate rail transit data asset as the final recommended rail transit data asset.
[0260] The present invention discloses a method and system for constructing a data asset service catalog for rail transit, wherein the method comprises the following steps: based on rail transit data rules, performing corresponding rule matching on rail transit data sources; both rail transit data rules and corresponding rail transit data tags are pre-stored in a rail transit data rule library; when the rail transit data source matches the rail transit data rules, assigning corresponding rail transit data asset tags to the rail transit data table corresponding to the rail transit data source; and constructing a data asset service catalog by integrating the rail transit data table and the corresponding rail transit data asset tags according to a preset data asset service catalog template.
[0261] Furthermore, it should be noted that the present invention can be provided as a method, apparatus, or computer program product. Therefore, embodiments of the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0262] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0263] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0264] It should also be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0265] Finally, it should be noted that the above description represents a preferred embodiment of the present invention. It should be pointed out that although preferred embodiments have been described, those skilled in the art, once they understand the basic inventive concept of the present invention, can make various improvements and modifications without departing from the principles described herein. These improvements and modifications should also be considered within the scope of protection of the present invention. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
Claims
1. A method for constructing a data asset service catalog for rail transit, characterized in that, The steps are as follows: Based on rail transit data rules, perform corresponding rule matching on rail transit data sources; The rail transit data source includes at least one level of rail transit database, rail transit data schema, rail transit data table, rail transit data field, and rail transit data field value; The specific steps for performing rule matching on the rail transit data source are as follows: Select the required rail transit data labels and corresponding rail transit data rules from the rail transit data rule library; The rail transit data rules are combined using regular expression symbols to obtain combined rail transit data rules; Based on the combined rail transit data rules, rule matching is performed on one or more levels of rail transit database, rail transit data pattern, rail transit data table, rail transit data field and rail transit data field value; The rail transit data rules and corresponding rail transit data tags are pre-stored in the rail transit data rule library; When the rail transit data source matches the rail transit data rules, the rail transit data table corresponding to the rail transit data source is assigned a corresponding rail transit data asset tag. Based on the preset data asset service catalog template, the rail transit data table and the corresponding rail transit data asset tags are integrated to construct the data asset service catalog.
2. The method for constructing a data asset service catalog for rail transit according to claim 1, characterized in that, The rail transit data rules include regular expression rules, keyword rules, dimension combination rules, and condition triggering rules; The regular expression rules include: rules for pattern matching of rail transit data field names, rail transit data table names, and rail transit data field values; The keyword rules include: applying matching rules to the remarks in the rail transit data table, the field comments in the rail transit data table, the rail transit database name, and the rail transit data table name; The dimension combination rules include: rules that match the field names of rail transit data, the field types in the rail transit data table, and the remarks. The conditional triggering rules include: rules that use the value of a rail transit data field conforming to a sensitive information pattern as the triggering condition.
3. The method for constructing a data asset service catalog for rail transit according to claim 1 or 2, characterized in that, Also includes: When the rail transit data source matches the rail transit data rules, the sensitive information, the range of rows and columns of the sensitive information, and the amount of sensitive information in the rail transit data are recorded based on the rail transit data tags corresponding to the sensitive information. Configure different security policies based on the different levels of sensitivity of information in the data asset service catalog; The different levels of the sensitive information correspond to different rail transit data tags; The security policy includes approval, data anonymization, data encryption, and access control for the range of sensitive information rows and / or the amount of sensitive information at this level.
4. The method for constructing a data asset service catalog for rail transit according to claim 3, characterized in that, Also includes: Based on the general large language model, the corresponding rail transit data assets are obtained by querying the data asset service catalog.
5. The method for constructing a data asset service catalog for rail transit according to claim 4, characterized in that, Based on the general large language model, the corresponding rail transit data assets are retrieved from the data asset service catalog. The steps are as follows: Vectorize the rail transit data assets in the data asset service catalog; Extract the semantics of the input query text; Based on the semantics, the most similar rail transit data assets are recalled from the vectorized rail transit data assets as the initial recall result; Based on the semantic matching of the rail transit data asset tags in the initial recall results, the rail transit data assets corresponding to the rail transit data asset tags that meet the matching rules are retained as candidate rail transit data assets. Based on the candidate rail transit data assets, prompt words are constructed, and corresponding term explanations are obtained based on a preset term database; The candidate rail transit data assets, the prompt words, and the terminology explanations are input into a general large language model, and a candidate rail transit data asset is output as the final recommended rail transit data asset.
6. A data asset service catalog construction system for rail transit, characterized in that, include: The rule matching module is used to perform corresponding rule matching on rail transit data sources based on rail transit data rules; The rail transit data source includes at least one level of rail transit database, rail transit data schema, rail transit data table, rail transit data field, and rail transit data field value; The rule matching module further includes: The rule selection module is used to select the required rail transit data labels and corresponding rail transit data rules from the rail transit data rule library; The rule combination module is used to combine the rail transit data rules using regular expression symbols to obtain combined rail transit data rules. The hierarchical matching module is used to perform rule matching on one or more levels of rail transit database, rail transit data pattern, rail transit data table, rail transit data field and rail transit data field value based on the combined rail transit data rules. A rail transit data rule library is used to pre-store the rail transit data rules and corresponding rail transit data tags; The tagging module is used to assign corresponding rail transit data asset tags to the rail transit data table corresponding to the rail transit data source when the rail transit data source matches the rail transit data rules. The data integration module is used to integrate the rail transit data table and the corresponding rail transit data asset tags to construct a data asset service catalog based on a preset data asset service catalog template.
7. The data asset service catalog construction system for rail transit according to claim 6, characterized in that, The rail transit data rules include regular expression rules, keyword rules, dimension combination rules, and condition triggering rules; The regular expression rules include: rules for pattern matching of rail transit data field names, rail transit data table names, and rail transit data field values; The keyword rules include: applying matching rules to the remarks in the rail transit data table, the field comments in the rail transit data table, the rail transit database name, and the rail transit data table name; The dimension combination rules include: rules that match the field names of rail transit data, the field types in the rail transit data table, and the remarks. The conditional triggering rules include: rules that use the value of a rail transit data field conforming to a sensitive information pattern as the triggering condition.
8. The data asset service catalog construction system for rail transit according to claim 6 or 7, characterized in that, The labeling module is also used to record the sensitive information, the range of rows and columns of sensitive information, and the amount of sensitive information in the rail transit data based on the rail transit data labels corresponding to the sensitive information, when the rail transit data source matches the rail transit data rules. The construction system also includes: The permission configuration module is used to configure different security policies based on the different levels of sensitive information in the data asset service directory; The different levels of the sensitive information correspond to different rail transit data tags; The security policy includes approval, data anonymization, data encryption, and access control for the range of sensitive information rows and / or the amount of sensitive information at this level.
9. The data asset service catalog construction system for rail transit according to claim 8, characterized in that, Also includes: The asset query module is used to retrieve the corresponding rail transit data assets from the data asset service catalog based on the general large language model.
10. The rail transit data asset service catalog construction system of claim 9, wherein, The asset query module includes: The asset vectorization module is used to vectorize rail transit data assets in the data asset service catalog; The query text vectorization module is used to extract the semantics of the input query text; The initial recall module is used to recall multiple rail transit data assets with the highest similarity from the vectorized rail transit data assets based on the semantics, as the initial recall result; The asset screening module is used to match the rail transit data asset tags in the initial recall results based on the semantics, and retain the rail transit data assets corresponding to the rail transit data asset tags that meet the matching rules as candidate rail transit data assets. An enhanced prompt word construction module is used to construct prompt words based on the candidate rail transit data assets and obtain corresponding term explanations based on a preset term database; A general large language model is used to receive the candidate rail transit data assets, the prompt words, and the terminology explanations, and output a candidate rail transit data asset as the final recommended rail transit data asset.
Citation Information
Patent Citations
Intelligent traffic asset data management method and system
CN119622556A
Intelligent data checking method, data asset searching and using method based on intelligent data checking and computer program product
CN118673002A
Method and system for determining a directory entry's class of service based on the value of a specifier in the entry
US20030078995A1