Railway data valuation method and related equipment

By employing a multi-level weighted calculation method, combined with the basic evaluation indicators and characteristics of railway datasets, the problem of railway data valuation in existing technologies has been solved, achieving scientific and accurate railway data valuation.

CN120952891APending Publication Date: 2025-11-14CHINA ACADEMY OF RAILWAY SCI CORP LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510990439.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing data evaluation methods are not applicable to the railway transportation industry and cannot effectively value railway data.

Method used

A multi-level weighted calculation method is adopted to obtain basic evaluation indicators of railway datasets, including data quality, data scale, data activity, demand, construction cost, operation and maintenance cost, management cost and model complexity. Combined with data cost value, application value-added coefficient and data model value-added coefficient, a scientific and accurate valuation is made.

Benefits of technology

It enables scientific and accurate valuation of railway data, taking into account the multi-dimensional characteristics and industry features of railway data, and provides a scientific quantitative evaluation method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952891A_ABST
    Figure CN120952891A_ABST
Patent Text Reader

Abstract

The invention relates to the field of railway data estimation, in particular to a railway data estimation method and related equipment. The method comprises the following steps: acquiring a railway data set to be estimated; obtaining a plurality of basic evaluation indexes based on the to-be-evaluated railway data set; performing multi-level weighting calculation on the plurality of basic evaluation indexes to obtain a plurality of second-level evaluation indexes; the secondary evaluation indexes comprise data quality, data scale, data activity, demand degree, construction cost, operation and maintenance cost, management cost and model complexity; obtaining a plurality of first-level evaluation indexes according to the plurality of second-level evaluation indexes; the first-level evaluation indexes comprise a data cost value, an application increment coefficient and a data model increment coefficient; and obtaining an estimation result of the to-be-estimated railway data set according to the first-level evaluation index.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of railway data valuation, and in particular to a method and related equipment for valuing railway data. Background Technology

[0002] As data gradually becomes a fundamental strategic resource for society, its importance is increasing daily. Currently, data circulation is mainly achieved through three channels: sharing, openness, and trading, with data trading being the most important. In terms of data valuation, most methods focus on general data asset valuation, as well as methods for specific fields such as clinical diseases, marine environment, highways, power control, and cities. However, the valuation of railway-specific data assets is not yet addressed. Existing data valuation methods are not applicable to the railway transportation industry. Therefore, how to value railway data is a pressing issue that needs to be addressed. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a method and related equipment for valuing railway data, for realizing the valuation of railway data.

[0004] In a first aspect, embodiments of the present invention provide a method for estimating railway data, including: Obtain the dataset of railways to be valued; Based on the aforementioned dataset of railways to be valued, several basic evaluation indicators are obtained; The multiple basic evaluation indicators are weighted at multiple levels to obtain multiple secondary evaluation indicators; the secondary evaluation indicators include: data quality, data scale, data activity, demand, construction cost, operation and maintenance cost, management cost, and model complexity. Multiple primary evaluation indicators are derived from the multiple secondary evaluation indicators; the primary evaluation indicators include: data cost value, application value-added coefficient, and data model value-added coefficient. The valuation results of the railway dataset to be valued are obtained based on the primary evaluation indicators.

[0005] In one possible implementation, the multi-level weighted calculation of the multiple basic evaluation indicators yields multiple secondary evaluation indicators, including: The multiple basic evaluation indicators are weighted and calculated to obtain multiple tertiary evaluation indicators; the tertiary evaluation indicators include: standardization, completeness, accuracy, consistency, timeliness, accessibility, data volume, growth, access popularity, update rate, in-path application, out-of-path application, training cost, dataset size, and model parameters. The weighted calculation of the multiple tertiary evaluation indicators yields the multiple secondary evaluation indicators.

[0006] In one possible implementation, the weighted calculation of the plurality of tertiary evaluation indicators to obtain the plurality of secondary evaluation indicators includes: The data quality is obtained by weighting the standardization, completeness, accuracy, consistency, timeliness, and accessibility.

[0007] In one possible implementation, the weighted calculation of the plurality of tertiary evaluation indicators to obtain the plurality of secondary evaluation indicators includes: The data size is obtained by weighting the data volume and the growth rate.

[0008] In one possible implementation, the weighted calculation of the plurality of tertiary evaluation indicators to obtain the plurality of secondary evaluation indicators includes: The data activity is obtained by weighting the access popularity and the update rate.

[0009] In one possible implementation, the weighted calculation of the plurality of tertiary evaluation indicators to obtain the plurality of secondary evaluation indicators includes: The demand degree is obtained by weighting the in-road application and the out-of-road application.

[0010] In one possible implementation, the weighted calculation of the plurality of tertiary evaluation indicators to obtain the plurality of secondary evaluation indicators includes: The model complexity is obtained by weighting the training cost, the dataset size, and the model parameters.

[0011] In one possible implementation, the multiple basic evaluation metrics include: Data standards, data models, metadata, business rules, authoritative reference sources, security specifications, data element integrity, data record integrity, content correctness, format compliance, data duplication rate, data uniqueness, dirty data occurrence rate, consistency of identical data, consistency of related data, time period correctness, time point time timeliness, time sequence, accessibility, availability, storage space, total number of records, total number of fields, record growth rate, field growth rate, query data volume, number of users accessing the site, access duration, return visit rate, proportion of cold data, data modification rate, data update rate, application integration degree, scenario coverage, sharing popularity, social popularity, number of off-the-road inquiries, data universality, number of data citations, number of market competitors, graphics processor computing power, central processing unit computing power, video memory capacity, storage scale, cluster scale, data scale, total number of model parameters, model accuracy, number of model layers, and computational complexity.

[0012] In one possible implementation, the weighted calculation of the multiple basic evaluation indicators to obtain multiple tertiary evaluation indicators includes: The standardization is obtained by weighting the data standard, the data model, the metadata, the business rules, the authoritative reference source, and the security specification. The integrity is obtained by weighting the integrity of the data elements and the integrity of the data records; The accuracy is obtained by weighting the correctness of the content, the compliance of the format, the data duplication rate, the data uniqueness, and the occurrence rate of dirty data. The consistency is obtained by weighting the consistency of the identical data and the consistency of the related data; The timeliness is obtained by weighting the correctness of the time period, the timeliness of the time point, and the temporality. The accessibility is obtained by weighting the accessibility and the availability. The data volume is obtained by weighting the storage space, the total number of records, and the total number of fields. The growth rate is obtained by weighting the growth rate of the record and the growth rate of the field. The access popularity is obtained by weighting the query data volume, the number of users accessing the service, the access duration, the return visit rate, and the proportion of cold data. The update rate is obtained by weighting the data modification rate and the data update rate. The in-road application is obtained by weighting the application integration degree, the scene coverage degree, and the sharing popularity. The off-road applications are obtained by weighting the social popularity, the number of off-road inquiries, the data universality, the number of data citations, and the number of market competitions. The training cost is obtained by weighting the computing power of the graphics processor, the computing power of the central processing unit, the video memory capacity, the storage scale, and the cluster scale. The dataset size is obtained by performing a weighted calculation on the data size. The model parameters are obtained by weighting the total number of model parameters, the model accuracy, the number of model layers, and the computational complexity.

[0013] In a second aspect, embodiments of the present invention provide a railway data estimation device, comprising: The acquisition module is used to acquire the dataset of railways to be valued. The evaluation module is used to obtain multiple basic evaluation indicators based on the railway dataset to be valued; The processing module is used to perform multi-level weighted calculations on the multiple basic evaluation indicators to obtain multiple secondary evaluation indicators; the secondary evaluation indicators include: data quality, data scale, data activity, demand, construction cost, operation and maintenance cost, management cost, and model complexity; The processing module is also used to obtain multiple primary evaluation indicators based on the multiple secondary evaluation indicators; the primary evaluation indicators include: data cost value, application value-added coefficient, and data model value-added coefficient; The valuation module is used to obtain the valuation results of the railway dataset to be valued based on the primary evaluation indicators.

[0014] Thirdly, embodiments of the present invention provide an electronic device, comprising: At least one processor; and At least one memory communicatively connected to the processor, wherein: The memory stores program instructions that can be executed by the processor, and the processor can execute the method described in the first aspect by calling the program instructions.

[0015] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer instructions that cause the computer to perform the method described in the first aspect.

[0016] In this embodiment of the invention, a multi-dimensional quantitative evaluation of the railway dataset is achieved through basic evaluation indicators. The final valuation result is obtained by performing multi-level weighted calculations on the basic evaluation indicators, thus realizing a scientific and accurate valuation of the railway dataset. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating a method for estimating railway data provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a multi-level weighted calculation provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a railway data device provided in an embodiment of the present invention; Figure 4This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0019] To better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0020] It should be understood that the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0021] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0022] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0023] It should be understood that although terms such as first, second, third, etc., may be used to describe numbers in embodiments of the present invention, these numbers should not be limited to these terms. These terms are only used to distinguish numbers from each other. For example, without departing from the scope of embodiments of the present invention, a first number may also be referred to as a second number, and similarly, a second number may also be referred to as a first number.

[0024] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0025] To achieve the valuation of railway data, embodiments of the present invention provide a method for valuing railway data. Figure 1 A flowchart illustrating a railway data estimation method provided in an embodiment of the present invention. Figure 1 As shown, the method includes: Step 101: Obtain the dataset of railways to be valued.

[0026] In this context, "railway dataset" refers to the collection of data generated, transmitted, and stored in the railway industry during business applications such as transportation scheduling, operation and maintenance management, and production construction. Railway datasets often exist in various forms, including images, text, video, and audio. Compared to other types of data, railway datasets are characterized by rapid updates, diverse types, large data volumes, and high confidentiality. In this embodiment of the invention, the railway dataset refers to the dataset obtained by performing preliminary processing steps such as cleaning, classification, and structuring on the original dataset.

[0027] Step 102: Based on the dataset of railways to be valued, obtain several basic evaluation indicators.

[0028] In this embodiment of the invention, multiple basic evaluation indicators are used to achieve quantitative evaluation of railway data valuation. Specifically, the basic evaluation indicators may include data standards, data models, metadata, business rules, authoritative reference sources, security specifications, data element integrity, data record integrity, content correctness, format compliance, data duplication rate, data uniqueness, dirty data occurrence rate, consistency of identical data, consistency of related data, correctness of time periods, timeliness of time points, time sequence, accessibility, availability, storage space, total number of records, total number of fields, record growth rate, field growth rate, query data volume, number of users accessing the service, access duration, return visit rate, proportion of cold data, data modification rate, data update rate, application integration degree, scenario coverage, sharing popularity, social popularity, number of external inquiries, data universality, number of data citations, number of market competitors, graphics processor computing power, central processing unit computing power, video memory capacity, storage scale, cluster scale, data scale, total number of model parameters, model accuracy, number of model layers, and computational complexity.

[0029] Because different basic evaluation indicators have different value ranges, to facilitate data alignment, after each basic evaluation indicator is calculated according to its corresponding method, it is mapped to an integer score of 1 to 5 points using four thresholds. A higher score indicates better performance of the railway data in that basic evaluation indicator. The threshold format is {t4; t3; t2; t1}, where t4 > t3 > t2 > t1. When the calculated value > t4, it is mapped to 5 points. When t4 ≥ the calculated value > t3, it is mapped to 4 points. When t3 ≥ the calculated value > t2, it is mapped to 3 points. When t2 ≥ the calculated value > t1, it is mapped to 2 points. When t1 ≥ the calculated value, it is mapped to 1 point.

[0030] Specifically, the calculation methods and grade levels for each basic evaluation indicator are as follows: Data standards. The calculation method is: number of elements conforming to the data standards / total number of elements in the dataset. The threshold values ​​are {0.9; 0.8; 0.7; 0.6}.

[0031] Data model. The calculation method is: number of elements meeting the model requirements / total number of elements in the dataset. The threshold values ​​are {0.99; 0.95; 0.90; 0.85}.

[0032] Metadata. The calculation method is: number of elements conforming to the metadata definition / total number of elements in the dataset. The threshold values ​​are {0.99; 0.95; 0.90; 0.85}.

[0033] Business rules. The calculation method is: number of elements that meet the business rule requirements / total number of elements in the dataset. The threshold values ​​are {0.99; 0.95; 0.90; 0.85}.

[0034] Authoritative reference source. Calculated as: number of elements from authoritative reference sources / total number of elements in the dataset. The threshold values ​​are {0.9; 0.8; 0.7; 0.6}.

[0035] Safety standards. The calculation method is: number of elements conforming to safety management standards / total number of elements in the dataset. The threshold values ​​are {0.99; 0.95; 0.90; 0.85}.

[0036] Data element integrity. Calculated as: number of elements meeting the integrity assignment criteria / total number of elements in the dataset. Threshold values ​​are {0.99; 0.95; 0.90; 0.85}.

[0037] Data record integrity. The calculation method is: number of records meeting the integrity assignment criteria / total number of records in the dataset. The threshold values ​​are {0.99; 0.95; 0.90; 0.85}.

[0038] Content accuracy. The calculation method is: number of records with reasonable and expected values ​​ / total number of records in the dataset. The threshold values ​​are {0.99; 0.98; 0.97; 0.96}.

[0039] Format compliance. Calculated as: number of elements meeting the expected format / total number of elements in the dataset. The threshold values ​​are {0.99; 0.98; 0.97; 0.96}.

[0040] Data duplication rate. Calculated as: number of unexpectedly duplicated records / total number of records in the dataset. Thresholds are {0.01; 0.02; 0.03; 0.04}.

[0041] Data uniqueness. The calculation method is: number of elements meeting the uniqueness standard / total number of master data elements in the dataset. The threshold values ​​are {0.99; 0.97; 0.95; 0.9}.

[0042] Dirty data occurrence rate. Calculated as: number of records containing invalid data / total number of records in the dataset. Thresholds are {0.01; 0.03; 0.06; 0.1}.

[0043] Data consistency is measured by the following formula: number of records meeting the synchronous modification query requirements / total number of records in the dataset. The threshold values ​​are {0.99; 0.98; 0.97; 0.96}.

[0044] Data consistency across related records. The calculation method is: number of records meeting the association constraint rules / total number of records in the dataset. The threshold values ​​are {0.99; 0.95; 0.90; 0.85}.

[0045] Time period accuracy. The calculation method is: number of records within the required date range / total number of records in the dataset associated with a date. The threshold values ​​are {0.99; 0.95; 0.90; 0.85}.

[0046] Timeliness of timestamps. Calculated as: number of records with timestamps recorded as needed / total number of records associated with timestamps in the dataset. Threshold values ​​are {0.99; 0.95; 0.90; 0.85}.

[0047] Temporal sequence. The calculation method is: number of correctly recorded records with a temporal sequence relationship / number of records in the dataset that should have a temporal sequence relationship. The threshold values ​​are {0.99; 0.95; 0.90; 0.85}.

[0048] Accessible. Calculated as: number of records accessible when needed / total number of records in the dataset. Thresholds are {0.99; 0.95; 0.90; 0.75}.

[0049] Availability. Calculated as: number of records that can be accessed on demand by the data interface / total number of records in the dataset. Thresholds are {0.99; 0.95; 0.90; 0.75}.

[0050] Storage space. The value represents the total storage space occupied by the dataset. The price tiers are {1PB; 1TB; 1GB; 1MB}.

[0051] Total number of records. The value indicates how many records the dataset contains. The thresholds are {200M; 2M; 20K; 200}.

[0052] Total number of fields. The value represents the total number of data elements (fields) contained in the dataset. The thresholds are {500; 100; 20; 5}.

[0053] Record the growth rate. The value represents the average number of new records added to the dataset per day. The thresholds are {10M; 100K; 1K; 10}.

[0054] Field growth rate. The value represents the average number of data elements (fields) added to the dataset each month. The increments are {7; 5; 3; 1}.

[0055] Data volume queried. The value represents the average number of records returned per query on the dataset. The thresholds are {100; 60; 30; 5}.

[0056] Number of users accessing the dataset. The value represents the number of times the dataset is accessed by users per hour. The thresholds are {10K; 1K; 100; 10}.

[0057] Access duration. The value represents the average number of minutes per user access to the dataset. The thresholds are {30; 10; 3; 1}.

[0058] Return visit rate. Calculated as: number of users who visited this dataset two or more times / total number of users who visited this dataset. The thresholds are {0.4; 0.3; 0.2; 0.1}.

[0059] Cold data ratio. Calculated as: number of records in the dataset that meet the cold data standard / total number of records in the dataset. The threshold values ​​are {0.8; 0.4; 0.2; 0.1}.

[0060] Data modification rate. Calculated as: average number of records that changed per day / total number of records in the dataset. Thresholds are {0.3; 0.2; 0.1; 0.01}.

[0061] Data update rate. The value represents the time interval between dataset updates. The increments are {1ms; 1s; 1h; 1d}.

[0062] Application integration level. The value represents the number of applications in the business system that have integrated this dataset. The thresholds are {7; 5; 3; 1}.

[0063] Scenario coverage. The value represents the number of railway business scenarios that require this dataset. The threshold values ​​are {15; 10; 5; 1}.

[0064] Share popularity. The value represents the number of users who have submitted sharing requests to this dataset. The thresholds are {7; 5; 3; 1}.

[0065] Social popularity. The value represents the number of times this dataset is mentioned in news reports. The thresholds are {4; 3; 2; 1}.

[0066] Number of off-street inquiries. The value represents the number of times off-street user inquiries involved this dataset. The thresholds are {4; 3; 2; 1}.

[0067] Data universality. The calculation method is: number of non-railway industry-specific data elements / total number of data elements. The threshold values ​​are {0.9; 0.6; 0.3; 0.1}.

[0068] Citation count. The value represents the number of times this dataset is mentioned in papers or research reports. The increments are {4; 3; 2; 1}.

[0069] Market competition number. The value represents the number of similar or identical data products on the market. The price levels are {1; 2; 3; 4}.

[0070] Graphics processor computing power. Calculated as: total floating-point operations required during the training process / total floating-point operations per hour for a single graphics processor. The power levels are {512; 64; 8; 1}.

[0071] Central Processing Unit (CPU) computing power. Calculated as: total number of floating-point operations required during the training process / total number of floating-point operations per hour for a single reference CPU core. The power levels are {512; 64; 8; 1}.

[0072] Video memory capacity. Calculated as the total video memory required for training the model (in GB). The tiers are {640GB; 80GB; 24GB; 8GB}.

[0073] Storage size. The value represents the total memory used during training (in GB). Tiers are {512GB; 64GB; 8GB; 1GB}.

[0074] Cluster size. Calculated as the number of servers required to train the model. The thresholds are {8; 4; 2; 1}.

[0075] Data size. Calculated as the storage space (GB) occupied by the training set, test set, and intermediate data. The price tiers are {512GB; 64GB; 8GB; 1GB}.

[0076] Total number of model parameters. Calculated as: the total number of model parameters. Gear levels are {100B; 20B; 1B; 100M}.

[0077] Model accuracy. Calculated as the number of floating-point bytes for a single parameter. The range is {8; 4; 2; 1}.

[0078] Number of model layers. Calculated as the number of layers in the deep neural network. The thresholds are {10000; 1000; 100; 10}.

[0079] Computational complexity. The calculation method is: the number of floating-point operations required for each inference iteration of the model. The thresholds are {100B; 20B; 1B; 100M}.

[0080] Step 103 involves performing multi-level weighted calculations on multiple basic evaluation indicators to obtain multiple secondary evaluation indicators. These secondary evaluation indicators include: data quality, data scale, data activity, demand level, construction cost, operation and maintenance cost, management cost, and model complexity.

[0081] Step 104: Based on multiple secondary evaluation indicators, obtain multiple primary evaluation indicators. The primary evaluation indicators include: data cost value, application value-added coefficient, and data model value-added coefficient.

[0082] The railway dataset is generated, analyzed, and applied based on the railway business system. Therefore, it is strongly correlated with the cost of the business system, which consists of three parts: construction cost, operation and maintenance cost, and management cost.

[0083] The application value-added factor refers to the amplification of the cost value of railway data beyond its cost value, taking into account the continuous growth of data asset value, the application prospects and potential benefits of railway datasets, and the appropriate amplification of cost value based on these benefits.

[0084] Regarding the value-added factor of data models, some railway data products are products that bundle pre-trained machine learning models and related data to form intelligent services for trading, i.e., intelligent data sets. For such railway datasets, the additional costs incurred in training artificial intelligence models need to be considered, and the data cost value should be appropriately amplified based on these additional costs.

[0085] Step 105: Obtain the valuation results of the railway dataset to be valued based on the primary evaluation indicators.

[0086] In this embodiment of the invention, the railway dataset is valued from three major dimensions using three primary evaluation indicators. Each primary evaluation indicator is obtained through multi-level weighting. Figure 2 This is a schematic diagram of a multi-level weighted calculation provided in an embodiment of the present invention. For example... Figure 2 As shown, the multi-level weighted calculation involves 3 primary evaluation indicators, 8 secondary evaluation indicators, 15 tertiary evaluation indicators, and 50 basic evaluation indicators.

[0087] like Figure 2 As shown, the final valuation result P of the railway dataset to be valued is calculated as follows: P = TC × (1 + AC + MC). Where TC is the basic value of the data cost, AC is the application value-added coefficient, and MC is the data model value-added coefficient.

[0088] Specifically, the data cost value TC consists of the construction cost TC of the business system containing the dataset. B Operation and maintenance costs TC O and management costs TC M The weighted summation is used to obtain the result. The calculation method is: T C =w1 B TC B +w1 O TC O +w1 M TC M Among them, w1 B w1 O w1 M These are the weights corresponding to the construction cost, operation and maintenance cost, and management cost of the respective business systems.

[0089] Given the characteristics of railway business systems, the weights of operation and maintenance costs and management costs should be greater than those of construction costs. For example, the weight of construction costs should be 0.2, the weight of operation and maintenance costs should be 0.4, and the weight of management costs should be 0.4.

[0090] Assuming the railway dataset originates from N business systems, when calculating the value of the data cost, each item cost is the sum of the costs of that item corresponding to each of the involved business systems. Let Pn represent the proportion of the data contributed by the nth business system to the system's cost, where n∈[1,N]. Then we have:

[0091]

[0092] .

[0093] Among them, TCn B TCn O TCn M These represent the construction cost, operation and maintenance cost, and management cost of the nth business system, respectively.

[0094] In some embodiments, when calculating the cost value of data, it is also necessary to consider the costs of historical year Y. Assume that due to factors such as inflation, the currency adjustment factor for historical year Y relative to this year is IF. y y∈[1,Y]. Then we have:

[0095]

[0096]

[0097] Among them, TC n,y B ,TC n,y O ,TC n,y M These represent the construction cost, operation and maintenance cost, and management cost of the nth business system in year y, respectively.

[0098] Based on the above, the final formula for calculating the data cost value (TC) is as follows:

[0099] For the application of the added value coefficient AC, it can be based on Figure 2 The diagram illustrating multi-level weighted calculation shows the results obtained through multi-level weighted calculation. Specifically, multiple basic evaluation indicators can be weighted first to obtain multiple tertiary evaluation indicators. Then, the multiple tertiary evaluation indicators are weighted to obtain multiple secondary evaluation indicators.

[0100] like Figure 2 As shown, the weighted calculation of multiple tertiary evaluation indicators yields multiple secondary evaluation indicators, specifically: Data quality is determined by weighting the factors of standardization, completeness, accuracy, consistency, timeliness, and accessibility.

[0101] We calculate the data scale by weighting the data volume and growth rate.

[0102] Data activity is obtained by weighting access popularity and update rate.

[0103] The demand level is obtained by weighting the in-street and out-of-street applications.

[0104] The aforementioned tertiary evaluation indicators can be obtained by weighting the basic evaluation indicators. Specifically: The data standards, data models, metadata, business rules, authoritative reference sources, and security specifications are weighted and calculated to obtain the standardization.

[0105] The integrity of data elements and data records is calculated by weighting the data element integrity and data record integrity.

[0106] Accuracy is obtained by weighting the content correctness, format compliance, data duplication rate, data uniqueness, and the occurrence rate of dirty data.

[0107] Consistency is obtained by weighting the consistency of identical data and the consistency of related data.

[0108] The timeliness is obtained by weighting the accuracy of the time period, the timeliness of the time point, and the chronological order.

[0109] Accessibility is obtained by weighting accessibility and availability.

[0110] The data volume is obtained by weighting the storage space, the total number of records, and the total number of fields.

[0111] The growth rate is calculated by weighting the growth rate of records and the growth rate of fields.

[0112] The access popularity is calculated by weighting the query data volume, number of users, access duration, return rate, and proportion of cold data.

[0113] The update rate is obtained by weighting the data modification rate and the data update rate.

[0114] The application integration degree, scenario coverage and sharing popularity are weighted and calculated to obtain the in-road application.

[0115] We calculate the off-road applications by weighting social popularity, number of off-road inquiries, data universality, number of data citations, and number of market competitions.

[0116] Table 1-1 shows a weight allocation for a multi-level weighted calculation using a value-added coefficient.

[0117]

[0118]

[0119] Table 1-1 As shown above, the application value-added coefficient AC can be obtained by performing multi-level weighted calculations on each basic evaluation indicator according to the weights in Table 1-1.

[0120] The calculation of the data model's value-added coefficient (MC) is similar to that of the applied value-added coefficient (AC). Multiple tertiary evaluation indicators are obtained by weighting several basic evaluation indicators. These tertiary indicators are then weighted to obtain multiple secondary evaluation indicators. Finally, the data model's value-added coefficient (MC) is obtained by weighting these secondary indicators.

[0121] Specifically, multiple secondary evaluation indicators are obtained by weighting and calculating multiple tertiary evaluation indicators, as follows: The model complexity is obtained by weighting the training cost, dataset size, and model parameters.

[0122] Specifically, each tertiary evaluation indicator can be calculated by weighting the corresponding basic evaluation indicators. Specifically: The training cost is obtained by weighting the computing power of the graphics processing unit, the computing power of the central processing unit, the video memory capacity, the storage scale, and the cluster scale.

[0123] The dataset size is obtained by weighting the data size.

[0124] The model parameters are obtained by weighting the total number of model parameters, model accuracy, number of model layers, and computational complexity.

[0125] Table 1-2 shows the weight allocation for multi-level weighted calculation of the value-added coefficient of a data model.

[0126]

[0127] Table 1-2 As shown above, the data model value-added coefficient MC can be obtained by performing multi-level weighted calculations on each basic evaluation index according to the weights in Table 1-1.

[0128] Based on the weight configurations in Tables 1-1 and 1-2, and Figure 2 The weighting method shown above ultimately yields the following calculation method for the application value-added coefficient AC and the data model value-added coefficient MC of the railway dataset to be valued:

[0129] Among them, I 4 Let be the set of all the basic evaluation indicators in Tables 1-1 and 1-2, where i is a basic evaluation indicator in this set. i The preliminary results, such as ratios and capacities, are obtained from the calculation of the basic evaluation indicator i. j This refers to the j-th gear line corresponding to the basic evaluation indicator. H(x) is a step function, which takes a value of 1 when x≥0 and a value of 0 otherwise. R i Let w be the set of paths from the i-th basic evaluation indicator to its corresponding primary evaluation indicator, containing the three weight values ​​traversed by the path; k i It is R i The kth weight value in the middle.

[0130] In some embodiments, the circulation path of railway datasets includes both internal and external circulation. The main purpose of internal circulation is to enhance collaboration between enterprises and departments, improve data utilization efficiency, and promote the informatization and intelligent development of railways. Therefore, internal circulation can refer to obtaining valuation results based on the cost value of the data.

[0131] In some embodiments, a sensitivity check can be performed on the railway dataset first. Specifically, if the railway dataset is highly sensitive (e.g., involves non-public data related to internal railway operations and is not suitable for public release), then this portion of the dataset may not be valued. The railway dataset valuation method provided in this embodiment of the invention can be applied to scenarios where the data sensitivity is below a preset threshold (e.g., assessed as "medium sensitive" or "low sensitive" according to railway data protection or industry secret protection requirements).

[0132] For tradable railway datasets, the first step is to obtain a suggested price range. Evaluation metrics include data volume, variety, completeness, time span, real-time performance, depth, coverage, and scarcity. Combining historical transaction data, a suggested price range [Pmin, Pmax] is derived as a reference basis for subsequent pricing strategies, ensuring that the final valuation results are objectively grounded and avoiding confusion.

[0133] If the railway dataset possesses high value or highly specialized characteristics (e.g., assessed as highly scarce or irreplaceable in specific scenarios), applicable scenarios include B2B transactions, high-value data products, or transactions requiring customization. The pricing method starts with the suggested price range mentioned above, and combines the valuation results obtained through the valuation method provided in this embodiment of the invention to arrive at the final price.

[0134] If the buyer group is highly diverse (i.e., their needs, purchasing power, or application scenarios differ significantly), and considering that data value varies depending on user characteristics and can be categorized into different versions or groups, multiple prices can be derived based on the valuation results (such as tiered pricing or segmented pricing).

[0135] If the market is highly volatile or the data is highly time-sensitive, such as prices being greatly affected by changes in supply and demand or value changing rapidly over time, this approach is suitable for scenarios involving short-term, highly time-sensitive data, regularly updated long-term data, and markets with large supply and demand fluctuations. In this case, the price is dynamically adjusted based on market conditions or time factors, building upon the valuation results.

[0136] In some embodiments, if the data product can be sold in combination with other data, a combined pricing strategy is provided as an additional option. In this scenario, the buyer needs to cross-validate or analyze multi-source data, and purchasing in combination can enhance the overall value. Specifically, the combination can be regarded as a new data product, and the valuation results of each railway dataset are obtained based on the railway data valuation method provided in the embodiments of the present invention. The specific price of the data product sold in combination can be based on the sum of the valuation results of each railway dataset and a discount D applied. s Let P T =(1-D s ) ∑P i Among them, P i For the first in the combination i The estimation results for the railway dataset, P T This is the total price of the package. Discount rate D s It depends on the amount of data or the complementarity of each railway dataset.

[0137] The railway data valuation method provided in this invention comprehensively considers the multi-dimensional characteristics of railway data, such as quality, scale, and activity, and combines the characteristics of the railway industry, data product forms, and circulation paths to conduct value assessment, thereby achieving a scientific and reliable quantitative assessment of railway data.

[0138] Based on the above-described railway data valuation method, this embodiment of the invention provides a railway data valuation device. Figure 3 This is a schematic diagram of a railway data estimation device provided in an embodiment of the present invention. Figure 3As shown, the railway data valuation device includes: an acquisition module 301, an evaluation module 302, a processing module 303, and a valuation module 304.

[0139] Module 301 is used to obtain the dataset of railways to be valued.

[0140] Evaluation module 302 is used to obtain multiple basic evaluation indicators based on the railway dataset to be valued.

[0141] Processing module 303 is used to perform multi-level weighted calculations on multiple basic evaluation indicators to obtain multiple secondary evaluation indicators. The secondary evaluation indicators include: data quality, data scale, data activity, demand, construction cost, operation and maintenance cost, management cost, and model complexity.

[0142] The processing module 303 is also used to obtain multiple primary evaluation indicators based on multiple secondary evaluation indicators. The primary evaluation indicators include: data cost value, application value-added coefficient, and data model value-added coefficient.

[0143] Valuation module 304 is used to obtain the valuation results of the railway dataset to be valued based on the primary evaluation indicators.

[0144] Figure 3 The railway data estimation device provided in the illustrated embodiment can be used to execute this specification. Figures 1-2 The implementation principle and technical effects of the method embodiment shown can be further referred to the relevant description in the method embodiment.

[0145] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 4 As shown, the aforementioned electronic device may include at least one processor and at least one memory communicatively connected to the processor, wherein the memory stores program instructions executable by the processor, and the processor can execute this specification by calling the program instructions. Figures 1-2 The embodiment shown provides a method for estimating railway data.

[0146] like Figure 4 As shown, the electronic device is represented in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: one or more processors 410, communication interface 420 and memory 430, and a communication bus 440 connecting different system components (including memory 430, communication interface 420 and processor 410).

[0147] Communication bus 440 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) buses, Micro Channel Architecture (MAC) buses, Enhanced ISA buses, Video Electronics Standards Association (VESA) local buses, and Peripheral Component Interconnect (PCI) buses.

[0148] Electronic devices typically include a variety of computer-readable media. These media can be any available media that can be accessed by the electronic device, including volatile and non-volatile media, and removable and non-removable media.

[0149] Memory 430 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The electronic device may further include other removable / non-removable, volatile / non-volatile computer system storage media. Memory 430 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments described herein.

[0150] A program / utility having a set (at least one) of program modules may be stored in memory 430. Such program modules include—but are not limited to—an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules typically perform the functions and / or methods described in the embodiments of this specification.

[0151] Processor 410 executes various functional applications and data processing by running programs stored in memory 430, such as implementing the functions described in this specification. Figures 1-2 The embodiment shown provides a method for estimating railway data.

[0152] This specification provides a computer program product, which includes a computer program that, when executed by a processor, performs the functions described in this specification. Figures 1-2The embodiment shown provides a method for estimating railway data.

[0153] This specification provides a computer-readable storage medium storing computer instructions that cause a computer to execute this specification. Figures 1-2 The embodiment shown provides a method for estimating railway data.

[0154] The aforementioned computer-readable storage medium may be any combination of one or more computer-readable media. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program that may be used by or in connection with an instruction execution system, apparatus, or device.

[0155] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0156] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0157] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this specification, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0158] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this specification includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which the embodiments of this specification pertain.

[0159] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0160] It should be noted that the devices involved in the embodiments of this specification may include, but are not limited to, personal computers (PCs), personal digital assistants (PDAs), wireless handheld devices, tablet computers, mobile phones, MP3 displays, MP4 displays, etc.

[0161] In the several embodiments provided in this specification, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0162] Furthermore, the functional units in the various embodiments of this specification can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units.

[0163] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, a connector, or a network device, etc.) or a processor to execute some steps of the methods described in the various embodiments of this specification. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0164] The above description is merely a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.

[0165] The same or similar parts between the various embodiments in this specification can be referred to mutually. In particular, the device embodiments and terminal embodiments are basically similar to the method embodiments, so the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.

Claims

1. A method for estimating railway data, characterized in that, include: Obtain the dataset of railways to be valued; Based on the aforementioned dataset of railways to be valued, several basic evaluation indicators are obtained; Multiple secondary evaluation indicators are obtained by performing multi-level weighted calculations on the aforementioned basic evaluation indicators; The secondary evaluation indicators include: data quality, data scale, data activity, demand, construction cost, operation and maintenance cost, management cost, and model complexity. Multiple primary evaluation indicators are derived from the multiple secondary evaluation indicators; the primary evaluation indicators include: data cost value, application value-added coefficient, and data model value-added coefficient. The valuation results of the railway dataset to be valued are obtained based on the primary evaluation indicators.

2. The method according to claim 1, characterized in that, The process of performing multi-level weighted calculations on the multiple basic evaluation indicators yields multiple secondary evaluation indicators, including: The multiple basic evaluation indicators are weighted and calculated to obtain multiple tertiary evaluation indicators; the tertiary evaluation indicators include: standardization, completeness, accuracy, consistency, timeliness, accessibility, data volume, growth, access popularity, update rate, in-path application, out-of-path application, training cost, dataset size, and model parameters. The weighted calculation of the multiple tertiary evaluation indicators yields the multiple secondary evaluation indicators.

3. The method according to claim 2, characterized in that, The weighted calculation of the multiple tertiary evaluation indicators yields the multiple secondary evaluation indicators, including: The data quality is obtained by weighting the standardization, completeness, accuracy, consistency, timeliness, and accessibility.

4. The method according to claim 2, characterized in that, The weighted calculation of the multiple tertiary evaluation indicators yields the multiple secondary evaluation indicators, including: The data size is obtained by weighting the data volume and the growth rate.

5. The method according to claim 2, characterized in that, The weighted calculation of the multiple tertiary evaluation indicators yields the multiple secondary evaluation indicators, including: The data activity is obtained by weighting the access popularity and the update rate.

6. The method according to claim 2, characterized in that, The weighted calculation of the multiple tertiary evaluation indicators yields the multiple secondary evaluation indicators, including: The demand degree is obtained by weighting the in-road application and the out-of-road application.

7. The method according to claim 2, characterized in that, The weighted calculation of the multiple tertiary evaluation indicators yields the multiple secondary evaluation indicators, including: The model complexity is obtained by weighting the training cost, the dataset size, and the model parameters.

8. The method according to claim 2, characterized in that, The aforementioned basic evaluation indicators include: Data standards, data models, metadata, business rules, authoritative reference sources, security specifications, data element integrity, data record integrity, content correctness, format compliance, data duplication rate, data uniqueness, dirty data occurrence rate, consistency of identical data, consistency of related data, time period correctness, time point time timeliness, time sequence, accessibility, availability, storage space, total number of records, total number of fields, record growth rate, field growth rate, query data volume, number of users accessing the site, access duration, return visit rate, proportion of cold data, data modification rate, data update rate, application integration degree, scenario coverage, sharing popularity, social popularity, number of off-the-road inquiries, data universality, number of data citations, number of market competitors, graphics processor computing power, central processing unit computing power, video memory capacity, storage scale, cluster scale, data scale, total number of model parameters, model accuracy, number of model layers, and computational complexity.

9. The method according to claim 8, characterized in that, The weighted calculation of the multiple basic evaluation indicators yields multiple tertiary evaluation indicators, including: The standardization is obtained by weighting the data standard, the data model, the metadata, the business rules, the authoritative reference source, and the security specification. The integrity is obtained by weighting the integrity of the data elements and the integrity of the data records; The accuracy is obtained by weighting the correctness of the content, the compliance of the format, the data duplication rate, the data uniqueness, and the occurrence rate of dirty data. The consistency is obtained by weighting the consistency of the identical data and the consistency of the related data; The timeliness is obtained by weighting the correctness of the time period, the timeliness of the time point, and the temporality. The accessibility is obtained by weighting the accessibility and the availability. The data volume is obtained by weighting the storage space, the total number of records, and the total number of fields. The growth rate is obtained by weighting the growth rate of the record and the growth rate of the field. The access popularity is obtained by weighting the query data volume, the number of users accessing the service, the access duration, the return visit rate, and the proportion of cold data. The update rate is obtained by weighting the data modification rate and the data update rate. The in-road application is obtained by weighting the application integration degree, the scene coverage degree, and the sharing popularity. The off-road applications are obtained by weighting the social popularity, the number of off-road inquiries, the data universality, the number of data citations, and the number of market competitions. The training cost is obtained by weighting the computing power of the graphics processor, the computing power of the central processing unit, the video memory capacity, the storage scale, and the cluster scale. The dataset size is obtained by performing a weighted calculation on the data size. The model parameters are obtained by weighting the total number of model parameters, the model accuracy, the number of model layers, and the computational complexity.

10. A device for estimating railway data, characterized in that, include: The acquisition module is used to acquire the dataset of railways to be valued. The evaluation module is used to obtain multiple basic evaluation indicators based on the railway dataset to be valued; The processing module is used to perform multi-level weighted calculations on the multiple basic evaluation indicators to obtain multiple secondary evaluation indicators; The secondary evaluation indicators include: data quality, data scale, data activity, demand, construction cost, operation and maintenance cost, management cost, and model complexity. The processing module is also used to obtain multiple primary evaluation indicators based on the multiple secondary evaluation indicators; the primary evaluation indicators include: data cost value, application value-added coefficient, and data model value-added coefficient; The valuation module is used to obtain the valuation results of the railway dataset to be valued based on the primary evaluation indicators.

11. An electronic device, characterized in that, include: At least one processor; as well as At least one memory communicatively connected to the processor, wherein: The memory stores program instructions that can be executed by the processor, and the processor can execute the method as described in any one of claims 1 to 9 by calling the program instructions.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause the computer to perform the method as described in any one of claims 1 to 9.