Method for Determining Infrastructure Master Data Elements for the Full Lifecycle of Highways

By constructing the data-process matrix and joint matrix, the main data elements of highway infrastructure are determined, and the problem of repeated collection and separate storage of highway infrastructure data in different business scenarios is solved, realizing the interactive sharing of data and the improvement of application value.

CN119167030BActive Publication Date: 2025-05-30RES INST OF HIGHWAY MINIST OF TRANSPORT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411258442.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2025-05-30
Estimated Expiration
2044-09-09

AI Technical Summary

Technical Problem

There are problems of repeated collection and separate storage of highway infrastructure data in different business scenarios, resulting in unfavorable data integration and full life cycle data.

Method used

By obtaining data from multiple business scenarios and building a data-process matrix, filtering business master data that meets the source, utilization and storage requirements, and determining infrastructure master data elements based on the relationship between business data and infrastructure data.

Benefits of technology

It realizes the interactive sharing of highway infrastructure data among different business scenarios, and improves the application value of infrastructure master data elements in business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119167030B_ABST
    Figure CN119167030B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method for determining infrastructure master data elements for the entire life cycle of a highway, including: obtaining multiple business scenarios and business data in the entire life cycle of the highway; constructing a data - process matrix for characterizing the data flow of business data in each business scenario according to the sources and usage of each business data; screening multiple business master data whose data flow meets the requirements of source, utilization, and storage; constructing a joint matrix for characterizing the dual association relationship between business data and infrastructure data according to the first correlation coefficient between each business data and each infrastructure data, and the second correlation coefficient between each business data; determining infrastructure master data elements in units of business master data according to the joint matrix and the master data screening principle. This embodiment reasonably determines infrastructure master data elements starting from business scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of highway master data identification, and in particular, to a method for determining infrastructure master data elements for the entire life cycle of a highway. Background Art

[0002] Highway infrastructure involves a large number and wide types, and experiences a long life cycle from the early-stage design, construction to the later-stage maintenance operation, road administration management, etc. Highway infrastructure data refers to the data that highway infrastructure can generate or needs, and is the main support for highway operations.

[0003] Due to the different requirements for the format, accuracy, etc. of infrastructure data in different stages of business, the current highway infrastructure data has problems such as repeated collection and separate storage, which is time-consuming, laborious and not easy to maintain, easily causing "data islands" and being unfavorable for the penetration of infrastructure data throughout the life cycle. Summary of the Invention

[0004] The embodiments of the present invention provide a method for determining infrastructure master data elements for the entire life cycle of a highway, which reasonably determines infrastructure master data elements from the business scenario and promotes the interaction and sharing of infrastructure data among different business scenarios through the master data base.

[0005] In a first aspect, the embodiments of the present invention provide a method for determining infrastructure master data elements for the entire life cycle of a highway, including:

[0006] Obtain multiple business scenarios in the entire life cycle of the highway, and the business data related to infrastructure that each business scenario can produce and needs;

[0007] Construct a data-process matrix for characterizing the data flow of business data in each business scenario according to the source and usage of each business data;

[0008] According to the data-process matrix, screen multiple business master data whose data flows meet the requirements of source, utilization, and storage;

[0009] Construct a joint matrix for characterizing the dual association relationship between business data and infrastructure data according to the first correlation coefficient between each business data and each infrastructure data, and the second correlation coefficient between each business data;

[0010] Determine infrastructure master data elements in units of business master data according to the joint matrix and the master data screening principle, where the infrastructure data element is the smallest data unit of infrastructure data.

[0011] In a second aspect, the embodiments of the present invention provide an electronic device, and the electronic device includes:

[0012] One or more processors;

[0013] A memory for storing one or more programs,

[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining infrastructure master data elements for the entire life cycle of a highway according to any of the embodiments.

[0015] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for determining infrastructure master data elements for the entire life cycle of a highway according to any of the embodiments.

[0016] In summary, an embodiment of the present invention provides a method for determining infrastructure master data elements for the entire life cycle of a highway. According to the data requirements of typical services in the entire life cycle of a highway, the association relationships between service data and infrastructure data, as well as the association relationships between service data, are analyzed, and the highway infrastructure master data throughout the entire life cycle is reasonably determined to promote the interaction and sharing of infrastructure data among different service scenarios. Specifically, the method first constructs a data-process matrix according to the sources and usage of service data to represent the data flow of service data in each service scenario, and screens out the service master data that meets the requirements of sources, utilization, and storage from it; then, taking the service master data as a unit, respectively determines the infrastructure data directly associated with the service master data and the infrastructure data indirectly associated through other service data, and jointly determines the infrastructure data elements corresponding to the current service master data through the two types of associated infrastructure data. The entire method uses the service master data as a bridge to incorporate the association relationships between service data and infrastructure data, as well as the association relationships between service data, into the screening of infrastructure master data elements, and can further improve the application value of infrastructure master data elements in service scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a flowchart of a method for determining infrastructure master data elements for the entire life cycle of a highway provided by an embodiment of the present invention.

[0019] Figure 2It is a schematic diagram of a data - process matrix for business data related to infrastructure in the entire life - cycle business scenario provided by an embodiment of the present invention.

[0020] Figure 3 It is a schematic diagram of a joint matrix provided by an embodiment of the present invention, including the business data correlation coefficient and the correlation coefficient between business data and infrastructure data.

[0021] Figure 4 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0022] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be described clearly and completely below. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present invention.

[0023] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0024] In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0025] Figure 1 It is a flowchart of a method for determining the main data elements of infrastructure for the entire life - cycle of a highway provided by an embodiment of the present invention. This method is executed by an electronic device, as Figure 1 shown, and this method specifically includes:

[0026] S110. Obtain multiple business scenarios in the entire life - cycle of the highway, and the business data related to infrastructure that can be produced and required by each business scenario.

[0027] The business data here is a series of data categories divided from a business perspective, such as topography and line conditions, rather than specific data values. In this embodiment, according to the application characteristics of highway data, highway infrastructure data is sorted out through business scenarios and business data.

[0028] Optionally, the business scenarios in the whole life cycle of a highway can be divided into first-level business scenarios and second-level business scenarios. The first-level business scenarios are the major categories of business scenarios divided according to the basic stages of the whole life cycle of a highway, which may include highway construction scenarios, highway maintenance scenarios, and highway operation scenarios, etc. The second-level business scenarios are the business scenarios obtained by further subdividing the first-level business scenarios. For example, by subdividing the highway construction scenario, second-level business scenarios such as surveying, design, and construction are obtained; by subdividing the highway maintenance scenario, second-level business scenarios such as road condition inspection and evaluation, maintenance decision-making, maintenance design, and maintenance construction are obtained; by subdividing the highway operation scenario, second-level business scenarios such as operation monitoring and emergency response are obtained.

[0029] Optionally, the business data includes the data that each business scenario can provide and the data that each business scenario needs. This embodiment focuses on the data related to highway infrastructure in the business data, such as topography, line conditions, design information, construction delivery data, etc.

[0030] In practical applications, the above-mentioned business scenarios and business data can be determined manually according to experience or automatically according to the business systems in the whole life cycle of a highway. When determined automatically according to the business systems, this embodiment provides the following specific implementation manner:

[0031] First, according to the main functions of multiple business systems in the whole life cycle of a highway, multiple first-level business scenarios are determined. Exemplarily, the business systems involved in the whole life cycle of a highway include a highway construction system, a highway maintenance system, and a highway operation system, etc. Then, the system names or system descriptions can be read to understand the main functions of each business system, and these main functions can be summarized as each first-level business scenario. The specific understanding and summarization operations can be realized by means of natural language processing technology and will not be elaborated here.

[0032] Then, according to the function modules related to infrastructure in each business system, the second-level business scenarios included in each first-level business scenario are determined. Taking the highway maintenance system as an example, the function modules related to infrastructure in this system include a road condition inspection and evaluation module, a maintenance decision-making module, a maintenance design module, and a maintenance construction module, etc. Then, the module names or module descriptions can be read to understand the business content of each function module, and these business contents can be summarized as the second-level business scenarios under the highway maintenance scenario.

[0033] Next, from the input file sets of each functional module, crawl the infrastructure-related business data required for each secondary business scenario. Taking the design scenario as an example, since this secondary business scenario comes from the design module in the highway construction system, extract the input file set of this module, and crawl the file names or file descriptions related to the infrastructure from the input file set, and use the themes of these file names or file descriptions as the business data required for the design scenario, which includes topographical features.

[0034] Similarly, from the output file sets of each functional module, crawl the infrastructure-related business data that each secondary business scenario can produce. Still taking the design scenario as an example, extract the output file set of the design module, and crawl the file names or file descriptions related to the infrastructure from the output file set, and summarize these file names or file descriptions as the business data that the design scenario can produce, which includes line conditions, design information, etc.

[0035] S120. According to the sources and usage situations of each business data, construct a data-process matrix for characterizing the data flow of business data in each business scenario.

[0036] For a certain business data, the source of the data refers to the business scenario that can produce this business data, and the usage situation of the data includes the time span between the business scenarios that use this business data, and the usage frequency of each business scenario for this business data.

[0037] In this embodiment, considering comprehensively the sources and usage situations of business data, a matrix with each business data as rows and each business scenario as columns is constructed, and the role played by a certain business scenario in the data circulation of a certain business data is used as the matrix element to characterize the data circulation situation (abbreviated as data flow) of infrastructure-related business data in the business scenarios of the entire highway life cycle. This embodiment calls this matrix the data-process matrix. Among them, the role played by a certain business scenario in the data circulation of a certain business data includes producer (C), short-term user (SU), medium-term user (MU), and long-term user (LU), which are also the four element values of the data-process matrix, as Figure 2 shown. In a specific implementation manner, the construction of this matrix may include the following steps:

[0038] Step 1. Construct an empty data-process matrix with each business data as rows and each secondary business scenario as columns. Refer to Figure 2 , first construct a data-process matrix with empty elements. Each row of the matrix corresponds to a kind of infrastructure-related business data u, and each column corresponds to a secondary business scenario. All the u in the rows together form the set of infrastructure-related business data

[0039] Step 2. For each industry's business data, perform the following operations respectively:

[0040] S1-1. Assign the elements corresponding to the secondary business scenarios that can produce the current business data as producers. Combining Figure 2 , in this matrix, C is used to represent producers. Taking the second industry's business data line condition as an example, if the secondary business scenario that generates this business data is the design scenario, then assign the matrix element in the second row and second column as C.

[0041] S1-2. Assign the elements corresponding to the secondary business scenarios that require the current business data as users. Still referring to Figure 2 , U is used to represent users. Still taking the second industry's business data line condition as an example, the secondary business scenarios that use this business data include the construction scenario, the maintenance design scenario, the maintenance construction scenario, and the emergency response scenario. Then assign all the matrix elements in the second row with these scenarios as columns as U. It should be noted that this step is an intermediate step, Figure 2 and the intermediate matrix including element U is not shown in

[0042] S1-3. Modify the users who belong to the same primary business scenario as the producer to short-term users. Since the same business scenario corresponds to the same life stage in the entire life cycle of the highway, the time span of the secondary business scenarios belonging to the same primary business scenario is relatively short. Based on this rule, in this embodiment, if the business data produced by a certain secondary business scenario is used by other secondary business scenarios under the same primary business scenario, then this type of use is regarded as short-term use. Still referring to Figure 2 , SU is used to represent short-term users in this matrix. Still taking the second industry's business data as an example, verify each U element in the second row of the above intermediate matrix one by one. If the secondary business scenarios corresponding to the U element and the C element in the second row belong to the same primary business scenario, then modify this U element to SU.

[0043] S1-4. Modify the users who belong to different primary business scenarios from the producer and whose usage frequency of the current business data is less than the set threshold to medium-term users; modify the users who belong to different primary business scenarios from the producer and whose usage frequency of the current business data is greater than or equal to the set threshold to long-term users. Since different primary business scenarios correspond to different life stages in the entire life cycle of the highway, the time span between the secondary business scenarios under different primary business scenarios is relatively large. Based on this rule, in this embodiment, if the business data generated by a certain secondary business scenario is used by secondary business scenarios under other primary business scenarios, then this type of use is regarded as medium- and long-term use, and the medium-term use and long-term use are further distinguished through the threshold of the usage frequency. Optionally, this threshold can be set to 5 times per day. Still referring toFigure 2 , in this matrix, MU and LU are used to represent medium-term users and long-term users respectively. Taking the second industry business data as an example, after the operations of S1-3, the elements corresponding to the secondary business scenario maintenance design, maintenance construction, and emergency response are still U. Verify these U elements one by one. Among them, if the usage frequency of the line conditions in the maintenance design scenario is greater than 5 times per day, then modify the U element in this column to LU; if the usage frequencies of the line conditions in the maintenance construction scenario and the emergency response scenario are less than 5 times per day, then modify the U elements in these two columns to MU.

[0044] S130. According to the data-process matrix, screen each business master data whose data flow meets the requirements of source, utilization, and storage.

[0045] In this embodiment, the master data in the business data is screened from the perspective of data usage. The specific screening method can be skillfully integrated with the element characteristics of the data-process matrix, and it can be determined whether the business data meets the master data conditions through simple counting. Optionally, the business master data needs to meet at least one of the requirements of single source, high utilization, and stable storage.

[0046] Among them, single source means that each business data is allowed to have one producer. Combining with the data-process matrix, the business data that meets the requirement of single source can be screened by the number of producer elements in each row, that is, there is only one C in each row of the data-process matrix.

[0047] High utilization means that the business data has a high utilization rate. Combining with the data-process matrix, the business data that meets the high utilization requirement can be screened by counting LU, MU, and SU. For example, if a row of the data-process matrix includes 1 LU, or 2 MUs, or more than 2 SUs, then screen the business data in this row as the business data that meets the high utilization requirement. Further, this screening condition can be determined manually according to needs, or can be calculated and determined according to the usage situation of the three types of users for the business data. When it is calculated and determined according to the usage situation of the three types of users for the business data, this embodiment provides the following specific implementation manner:

[0048] Step 1. According to the time span between the three types of users and the producers and the usage frequency of the business data, convert each user element in the data-process matrix into a count value with short-term users as the counting unit.

[0049] Optionally, since the greater the time span between the user and the producer, the higher the effective utilization degree of the current business data, the effective utilization degree of the business data that MU can reflect is higher than that of the business data that SU can reflect. In order to quantitatively represent this difference between MU and SU, this embodiment is based on the minimum time span t between SU and C 1, and the minimum time span t between MU and C 2 , convert each MU into SUs, which is the count value in terms of short-term users. Since t 1 > t 2 , Furthermore, t 1 and t 2 can be determined in the following way: Arrange all the first-level business scenarios in chronological order, and calculate the difference between the central time points of every two adjacent first-level business scenarios to obtain the time span between every two adjacent first-level business scenarios; Take the average of all the time spans between adjacent first-level business scenarios, and the average time span can be obtained, which represents the minimum time span t between MU and C 1 . Similarly, based on the above arrangement, further arrange the second-level business scenarios under each first-level business scenario in chronological order to obtain the time sequence of the second-level business scenarios; Calculate the difference between the central time points of every two second-level business scenarios to obtain the time span between every two adjacent second-level business scenarios; Take the average of all the time spans between adjacent second-level business scenarios, and the average time span can be obtained, which represents the minimum time span t between SU and C 2 .

[0050] In addition, since the higher the usage frequency of business data by users, the higher the effective utilization degree of the current business data, the effective utilization degree of business data that LU can reflect is higher than that of MU. To quantitatively represent this difference between LU and MU, in this embodiment, according to the average usage frequency f 1 of LU for business data, and the average usage frequency f 2 of MU for business data, convert each LU into MUs, and then convert them into SUs, which is the count value in terms of short-term users. Since f 1 > f 2 , Furthermore, the usage frequencies of business data by all LUs in the data-process matrix for their respective industries can be statistically analyzed, and the average of the usage frequencies of all LUs can be obtained to get the average usage frequency f 1 of LU for business data; At the same time, the usage frequencies of business data by all SUs in the data-process matrix for their respective industries can be statistically analyzed, and the average of the usage frequencies of all MUs can be obtained to get the average usage frequency f 2 of MU for business data.

[0051] Step 2: Filter the business data that meets the high-utilization requirements according to the sum of the count values of each row. Through Step 1, the LU and MU of each row are converted into the count values of SU, and the count value of the SU element itself is 1. Add the count values within the same row. The business data corresponding to the row whose sum of count values is greater than the set threshold is the business data that meets the high-utilization requirements. Preferably, the set threshold is taken as 3.

[0052] Stable storage means that the business data has the characteristic of long-term stable storage. Combining with the data-process matrix, it is also possible to screen the business data that meets the stable storage by counting LU and MU. For example, if a row of the data-process matrix includes 2 LUs, or 4 MUs, or 1 LU and 2 MUs, then the business data of this row is screened as the business data that meets the long-term stable storage characteristic. Similarly, this screening condition can be determined manually according to needs, or can be calculated and determined according to the usage situation of the business data by medium- and long-term users. When calculated and determined according to the usage situation of the business data by medium- and long-term users, the present embodiment provides the following specific implementation manner:

[0053] Step 1: Convert each medium-term user element and long-term user element in the data-process matrix into a count value with the medium-term user as the counting unit according to the usage frequencies of the medium-term users and long-term users for the business data. Since stable storage requires long-term preservation of the business data, in this embodiment, the number of medium- and long-term users is concerned, and the number of SUs is not specifically limited. Specifically, the more the number of LUs and MUs, the more stable storage the business data requires; and the higher the usage frequency of the users for the business data, the stronger the demand for stable storage of the business data. Therefore, the degree of stable storage of the business data that LU can reflect is higher than that of the business data that MU can reflect. In order to quantitatively represent this difference between LU and MU, in this embodiment, still according to the average usage frequency f 1 of the business data by LU, and the average usage frequency f 2 of the business data by MU, each LU is converted into MUs, which is the count value with the medium-term user as the counting unit. For the convenience of distinction and description, in this embodiment, the count value with the SU as the counting unit generated when screening the high-utilization business data is called the first count value, and the count value with the MU as the counting unit generated when screening the stable storage business data here is called the second count value.

[0054] Step 2: Screen business data that meets the requirements of stable storage according to the sum of the second count values of each row. Through Step 1, the LU of each row is converted into the second count value of MU, and the second count value of the MU element itself is 1. Add the second count values of the same row. The business data corresponding to the row where the sum of the second count values is greater than the set threshold is the business data that meets the requirements of stable storage. Preferably, the set threshold is taken as 4.

[0055] When using the three conditions of single source, high utilization, and stable storage as the screening conditions for business master data at the same time, finally, the business data that meets all three requirements is taken as the business master data. Optionally, if there are requirements for data value in some cases, the screening condition of high data value can be added. High data value means that the data is indispensable in business applications and highly affects business management efficiency, operating costs, driver experience, energy conservation and environmental protection, etc. A data value function g(u) can be constructed to measure the data value of business data u, and finally, the data that meets the above four requirements is screened as the business master data. The screened business master data can jointly form a subset of the infrastructure business master data set. Exemplarily, For a subset of. Exemplarily, including construction delivery data, maintenance construction and delivery data, etc.

[0056] S140. Construct a joint matrix for characterizing the dual association relationship between business data and infrastructure data according to the association coefficients between each business data and each infrastructure data, and the association coefficients between each business data.

[0057] Infrastructure data refers to a series of data categories that highway infrastructure can provide. According to the type of infrastructure, infrastructure data can include several categories such as routes, structures, traffic safety facilities, management facilities, service facilities, etc. Each infrastructure data further includes multiple infrastructure data elements. For example, a route includes data elements such as route name, starting stake number, ending stake number, and route length. The infrastructure data element is the smallest data unit of infrastructure data. In this embodiment, multiple elements will be selected from all infrastructure data elements as the elements of infrastructure master data.

[0058] It can be seen that infrastructure data is the direct source of infrastructure master data elements. At the same time, considering that there are also data circulation relationships between business data, this step describes the data relationship from two perspectives: the association relationship between business data and infrastructure data, and the association relationship between business data, and simultaneously characterizes the two relationships through a special matrix form. In this embodiment, this special matrix form is called a joint matrix. In a specific implementation manner, the construction of the above joint matrix may include the following steps:

[0059] Step 1. Determine the correlation coefficients between each business data and each infrastructure data. Specifically, the infrastructure data set can be represented as For each business data u and each infrastructure data The correlation coefficient is denoted as a u,k . Each a u,k can be determined by means of a questionnaire or expert scoring, or can be determined by the interaction data between business systems.

[0060] Optionally, when determined by a questionnaire or expert scoring, for each pair of business data u and infrastructure data k, the correlation relationship between the two data can be expressed as two categories: strong and weak, corresponding to the correlation coefficients a u,k = 1 and a u,k = 2 respectively; multiple a u,k can be obtained through multiple questionnaires or multiple expert scorings. Take the average value of all a u,k as the final correlation coefficient.

[0061] Optionally, when determined by the interaction data between business systems, in the case where each business data is determined according to the input and output files of the functional modules in the business system, the correlation coefficients between each business data and each infrastructure data can be determined according to the functional modules in each business system and the data interfaces of each infrastructure. Specifically, first, the input files and output files of each functional module in each business system can be extracted, and the input file and output file with the current business data as the theme can be determined. The data in the input file jointly constitute the input data set F 1 , and the data in the output file jointly constitute the output data set F 2 ; at the same time, according to the data interfaces of each infrastructure, the input data set I 1 and the output data set I 2 of each infrastructure can be determined. Then, for each pair of business data u and infrastructure data k, the intersection P 1 of the input data set F 2 of the business data u and the output data set I 1 of the infrastructure k is determined respectively, and the intersection P 2 of the output data set F 1 of the business data u and the input data set I 2 of the infrastructure k; both of these intersections are data elements and are called data intersections; for the data intersections P 1 and P 2Take the union and determine the total number of elements in the union (i.e., the total amount of data). The larger this total is, the higher the degree of association between the business data u and the infrastructure data k. After performing the above operations on each pair of business data and infrastructure data, the total amount of data for each pair of business data and infrastructure data can be obtained. Normalize these total amounts of data and then scale the normalized values to the same interval, such as [0, 2], to obtain the correlation coefficient a between each pair of business master data and infrastructure data. u,k 。

[0062] Step 2: Construct a full matrix with each business data as a row, each infrastructure data as a column, and the correlation coefficient between the business data in the corresponding row and the infrastructure data in the corresponding column as an element. That is, the element in the i-th row and j-th column of the matrix is the a between the business data with index i and the infrastructure data with index j. u,k 。

[0063] Step 3: Determine the correlation coefficients between business data. Specifically, the correlation coefficient between business data u and business data can be denoted as c u,u' , and each c u,u' can also be determined by means of questionnaires or expert scoring, or by the interaction data between business systems. For the convenience of distinction and description, in this embodiment, the correlation coefficient between business data and infrastructure data in the above steps is referred to as the first correlation coefficient, and the correlation coefficient between business data in this step is referred to as the second correlation coefficient.

[0064] Optionally, when determined by questionnaires or expert scoring, for each pair of business data u and u', the second association relationship between the two data can be expressed as three categories: strong, medium, and weak, corresponding to the second correlation coefficients c u,u' = 1, c u,u' = 2, and c u,u' = 3 respectively; through multiple decomposition questionnaires or multiple expert scores for multiple c u,u' , take the average value of all c u,u' as the final second correlation coefficient.

[0065] Optionally, when determining through the interaction data between business systems, in the case where each business data is determined based on the input and output files of functional modules in the business system, the second correlation coefficient between each business data can be determined according to the functional modules in each business system that take each business master data as input or output. Specifically, first, in each business system, the set of functional modules with each business master data as the input file theme and the set of functional modules with each business master data as the output file theme can be located. For the convenience of distinction and description, the above two sets of functional modules are sequentially referred to as the first set of functional modules and the second set of functional modules. Then, for each pair of business master data u and u', the intersection Pu of the first set of functional modules of business master data u and the second set of functional modules of u' is determined respectively 1 , and the intersection Pu of the second set of functional modules of business master data u and the first set of functional modules of u' 2 ; both of these intersections take functional modules as elements and are called module intersections; take the union of the module intersections Pu 1 and Pu 2 , and determine the total number of elements (i.e., the total number of modules) in the union. The larger this total, the higher the degree of association between business data u and u'. After performing the above operations on each pair of business data, the total number of modules for each pair of business data can be obtained. Normalize these total numbers of modules and then scale the normalized values to the same interval, such as [0, 3], to obtain the second correlation coefficient c u,u' .

[0066] Step Four: Construct a symmetric matrix with each business data as rows and columns and the second correlation coefficient between the business data in the corresponding row and column as elements, that is, the element in the i-th row and j-th column of the matrix is the second correlation coefficient c between the business data with index i and the business data with index j u,u' .

[0067] Step Five: Concatenate and display the lower triangular matrix of the symmetric matrix on the side of the row index of the full matrix to obtain a joint matrix for characterizing the dual association relationship between business data and infrastructure data Figure 3 An exemplary joint matrix is shown, where the rectangular area in the lower right corner of the matrix is the full matrix, the triangular area on the left side of the matrix is the lower triangular matrix of the symmetric matrix, and the values of the elements in the matrix are only for illustration and not fixed values

[0068] S150. Determine infrastructure master data elements in units of business master data according to the joint matrix and the master data screening principle

[0069] Based on the above joint matrix, in this embodiment, the principle of master data screening is applied to respectively determine the set of infrastructure master data elements corresponding to each business master data, and the union of all sets of infrastructure master data elements is taken to obtain the final set of infrastructure master data elements. In a specific implementation manner, this process may include the following steps:

[0070] Step 1: For each industry business data (including each industry business master data) in the joint matrix, at least one type of infrastructure data with a first correlation coefficient greater than a certain threshold is respectively determined, and the at least one type of infrastructure data together constitutes a set of infrastructure data associated with the current business master data Optionally, in the above joint matrix, the set of infrastructure data in the same row corresponding to each industry business master data can be denoted as where D u,k is the set of data elements corresponding to infrastructure data k. For example, the infrastructure data set corresponding to design information includes the route data set D 2,1 、the structure data set D 2,2 、the traffic safety facility data set D 2,3 、the management facility data set D 2,4 、the service facility data set D 2,5 etc. Among them, the route data set further includes multiple data elements and can be expressed as D 2,1 ={route name, starting stake number, ending stake number, route length,...}. In this step, a threshold is set for the correlation coefficient a u,k , for example, 1.5, and the corresponding of each industry business data is updated through this threshold. Specifically, if a u,k ≥1.5, then D u,k remains unchanged, otherwise D u,k is set to be empty, and thus the new infrastructure data set corresponding to each industry business data is obtained where

[0071]

[0072] Step 2: For each business master data, the following operations are respectively performed:

[0073] S2-1: Determine other business data whose second correlation coefficient with the current business master data is greater than another threshold and is not a business master data. In this step, another threshold is set for the second correlation coefficient, and other business data that is closely related to the business master data and is not a business master data is screened according to this threshold. In this embodiment, the other business data is referred to as target business data. Optionally, the other threshold can be taken as 1. For any pair of business master data u and business data if cu,u' If it is ≥1, then use u' as the target business data of the business master data u.

[0074] S2-2. Use the infrastructure data set associated with the target business data as another infrastructure data set associated with the current business master data After Step 1, the current business master data already has an associated infrastructure data set Due to the strong correlation between the target business data and the current business data, in this step, the infrastructure data set associated with the target business data is used as another infrastructure data set associated with the current business master data The formula is expressed as follows:

[0075]

[0076] For the convenience of distinction and description, in this embodiment, the infrastructure data in is called the first infrastructure data of the current business master data, and the infrastructure data in is called the second infrastructure data of the current business master data.

[0077] S2-3. Take the union of the first infrastructure data and the second infrastructure data of the current business master data, and filter out the infrastructure master data elements that meet the master data filtering principle from the union. Optionally, the union can be expressed as where m represents the number of target business data with c u,u' ≥1, represents the second infrastructure data set corresponding to the target business data with index m. Select a certain master data filtering principle, such as compliance, uniqueness, and stability, etc., and filter out the infrastructure data elements that meet the master data filtering principle from the above union. Optionally, the master data filtering principle can be represented by the master data filtering function f(*), then the master data element M can be expressed as:

[0078]

[0079] Among them, uniqueness can be described by consistent data representation and consistent data values. Among them, the consistency of data representation represents the number of data elements with the same metadata of the infrastructure data in different business data. Stability can be described by the number of updates of the data. Among them, the number of updates represents the number of data elements that are frequently modified. Compliance can be described by data type division, granularity, etc. Among them, the data type division represents the number of data elements with accurate data type division. Exemplarily, if the master data filtering function is data representation consistency, then filter The data elements that are consistent with the infrastructure data metadata are used as the master data elements.

[0080] Performing the same operation on each business master data respectively can obtain the infrastructure master data elements corresponding to each business master data. Taking the union of the infrastructure master data elements corresponding to all business master data can obtain the final set of infrastructure master data elements.

[0081] In summary, this embodiment provides a method for determining infrastructure master data elements for the entire life cycle of a highway. According to the data requirements of typical highway life cycle services, the correlation between service data and infrastructure data, as well as the correlation between service data, are analyzed to reasonably determine the highway infrastructure master data throughout the life cycle, promoting the interactive sharing of infrastructure data among different service scenarios. Specifically, this method first constructs a data-process matrix based on the source and usage of service data to represent the data flow of service data in each service scenario, and filters out the business master data that meets the requirements of source, utilization, and storage; then, taking the business master data as a unit, respectively determines the infrastructure data directly associated with the business master data and the infrastructure data indirectly associated through other service data, and jointly determines the infrastructure data elements corresponding to the current business master data through the two types of associated infrastructure data. The entire method uses business master data as a bridge to incorporate the correlation between service data and infrastructure data, as well as the correlation between service data, into the screening of infrastructure master data elements, which can further improve the application value of infrastructure master data elements in service scenarios.

[0082] Figure 4 The following is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 4 shown, the device includes a processor 60, a memory 61, an input device 62, and an output device 63; the number of processors 60 in the device can be one or more, Figure 4 taking one processor 60 as an example; the processor 60, memory 61, input device 62, and output device 63 in the device can be connected through a bus or other means, Figure 4 taking connection through a bus as an example.

[0083] The memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for determining infrastructure master data elements for the entire life cycle of a highway in the embodiment of the present invention. The processor 60 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 61, that is, implements the above-mentioned method for determining infrastructure master data elements for the entire life cycle of a highway.

[0084] The memory 61 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 61 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 61 may further include a memory remotely provided with respect to the processor 60, and these remote memories may be connected to the device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0085] The input device 62 may be used to receive input digital or character information and generate key signal inputs related to user settings and function controls of the device. The output device 63 may include a display device such as a display screen.

[0086] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for determining infrastructure master data elements for the entire life cycle of a highway in any embodiment.

[0087] The computer storage medium of the embodiment of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0088] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0089] The program code contained on a computer-readable medium can be transmitted with any suitable medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.

[0090] The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the C language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A method for determining infrastructure master data elements for the entire life cycle of a highway, characterized in that: include: Obtain multiple business scenarios throughout the life cycle of a highway, as well as the infrastructure-related business data that can be produced and required by each business scenario; According to the source and usage of each business data, a data-process matrix is ​​constructed to characterize the data flow of business data in each business scenario; According to the data-process matrix, select various business master data whose data flows meet the requirements of source, utilization and storage; Constructing a joint matrix for representing the dual association relationship between the business data and the infrastructure data according to the first association coefficient between each business data and each infrastructure data, and the second association coefficient between each business data; For each row of business data in the joint matrix, respectively determine first infrastructure data having a first correlation coefficient greater than a first threshold; For each business master data, perform the following operations respectively: S2-1, determining other business data that has a second correlation coefficient with the current business master data greater than another threshold and is not the business master data; S2-2, determining the first infrastructure data of the other business data as the second infrastructure data associated with the current business master data; S2-3, taking a union of the first infrastructure data and the second infrastructure data of the current business master data, and filtering infrastructure master data elements that meet the master data screening principle from the union; The infrastructure master data elements corresponding to each business master data are combined as the final infrastructure master data element set.

2. The method according to claim 1, characterized in that The acquisition of multiple business scenarios in the entire life cycle of the highway, as well as the business data related to the infrastructure that can be produced and required by each business scenario, includes: Determine multiple primary business scenarios based on multiple business systems throughout the highway life cycle; Determine the secondary business scenarios included in each primary business scenario based on the functional modules related to infrastructure in each business system; Crawl the infrastructure-related business data required by each secondary business scenario from the input file collection of each functional module; From the output file collection of each functional module, crawl the infrastructure-related business data that can be produced by each secondary business scenario.

3. The method according to claim 1, characterized in that Business scenarios include primary business scenarios and secondary business scenarios; Accordingly, according to the source and usage of each business data, a data-process matrix is ​​constructed to characterize the data flow of business data in each business scenario, including: Build an empty data-process matrix with each business data as row and each secondary business scenario as column; For each row of business data, perform the following operations: S1-1. Mark the secondary business scenarios that can produce current business data as producers; S1-2, marking the secondary business scenario that requires the current business data as a user; S1-3. Re-mark users who belong to the same primary business scenario as the producer as short-term users; S1-4. Users who belong to different primary business scenarios from the producer and whose frequency of using the current business data is less than the set threshold are re-marked as mid-term users; S1-5. Re-mark users who belong to different primary business scenarios from the producer and whose usage frequency of the current business data is greater than or equal to the set threshold as long-term users; S1-6. Assign the final mark of each secondary business scenario to the element of the corresponding column of the current row.

4. The method according to claim 1, characterized in that: The data-process matrix has business data as rows and secondary business scenarios as columns; the elements in the data-process matrix include producers, short-term users, medium-term users and long-term users, which respectively represent the roles of the secondary business scenarios in the columns in the data flow of the business data in the rows; Accordingly, the data flow according to the data-process matrix is ​​screened to select each business master data that meets the source, utilization and storage requirements, including: Filtering business data that meets the single-source requirement according to the number of producer elements in each row of the data-process matrix; According to the time span between the three types of users and producers and the frequency of use of business data, each user element in the data-process matrix is ​​converted into a first count value with short-term users as the counting unit; according to the sum of the first count values ​​of each row, the business data that meets the high utilization requirement is screened; According to the frequency of use of the business data by the mid-term users and the long-term users, each mid-term user and long-term user element in the data-process matrix is ​​converted into a second count value with the mid-term user as the counting unit; and according to the sum of the second count values ​​of each row, the business data that meets the stable storage requirement is screened; Take the business data that meets the three requirements as the master data of each business.

5. The method according to claim 1, characterized in that The step of constructing a joint matrix for characterizing the dual association relationship between the business data and the infrastructure data according to the first association coefficient between each business data and each infrastructure data, and the second association coefficient between each business data, includes: Determine a first correlation coefficient between each business data and each infrastructure data according to the function modules of each business system and the data interface of each infrastructure; A full matrix is ​​constructed with each business data as a row, each infrastructure data as a column, and the first correlation coefficient between the business data in the row and the infrastructure data in the column as an element; Determine a second correlation coefficient between each business data according to a function module in each business system that takes each business data as input or output; A symmetric matrix is ​​constructed by taking each business data as a row and a column respectively and taking the second correlation coefficient between the business data in the row and the column as an element; The lower diagonal matrix of the symmetric matrix is ​​spliced ​​and displayed on the row index side of the full matrix to obtain a joint matrix for representing the dual correlation relationship between business data and infrastructure data.

6. The method according to claim 5, characterized in that Determining the first correlation coefficient between each business data and each infrastructure data according to the functional modules of each business system and the data interface of each infrastructure includes: According to the functional modules of each business system, determine the input files and output files with each business data as the theme; Determine the input data set and output data set of each infrastructure according to the data interface of each infrastructure; Determine a first data intersection of a data set in an input file with each business data as a subject and an output data set of each infrastructure, and a second data intersection of a data set in an output file with each business data as a subject and an input data set of each infrastructure; A union is obtained for the first data intersection and the second data intersection between each business data and each infrastructure, and the total amount of data in the union is converted into a first correlation coefficient between each business data and each infrastructure data.

7. The method according to claim 5, characterized in that Determining the correlation coefficient between each business data according to the functional module with each business data as input or output in each business system includes: In each business system, locate a first set of functional modules that use each business data as input, and a second set of functional modules that use each business data as output; Take the first module intersection of the first function module set of one type of business data and the first module intersection of the second function module set of the other type of business data, and the second module intersection of the second function module set of the one type of business data and the second module intersection of the first function module set of the other type of business data; A union is obtained for the intersection of the first module and the intersection of the second module of every two types of business data, and the total amount of modules in the union is converted into a correlation coefficient between every two types of business data.

8. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining infrastructure master data elements for the entire life cycle of a highway as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, which, when executed by a processor, implements the method for determining the master data elements of infrastructure for the entire life cycle of a highway as described in any one of claims 1-7.