Data service component management method and device
By calculating the integration degree of data service components, the problem of how to evaluate the construction quality of data service components is solved, the quantitative evaluation and quality feedback of the data model structure are realized, and the construction effect of data service components is improved.
Patent Information
- Application Number
- CN202110230147.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-02
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2041-03-02
AI Technical Summary
How to evaluate the construction status of data service components, especially how to evaluate the construction quality of data service components in data middle-office projects.
The quality of data service components is quantitatively assessed by calculating the degree of integration within their data model structures, including the first, second, and third integration degrees. The first integration degree reflects the degree of data flow and integration within the subject domain corresponding to the first dataset, the second integration degree reflects the degree of data flow and integration within multiple subject domains, and the third integration degree reflects the degree of data flow and integration between different subject domains.
It realizes the quantitative evaluation of data service components, can reflect the construction quality of data model structure, provide evaluation information and modification prompts, and help users understand and improve the construction of data service components.
Smart Images

Figure CN115081495B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a method and device for managing a data service component. Background Art
[0002] By building data service components and utilizing the data service capabilities provided by the data service components, (enterprise) data can be processed into data products or services that meet the needs, making operational decisions more efficient. At the same time, the accumulation and sedimentation of large amounts of data can be mined to realize their potential value.
[0003] Taking the data middle platform as an example, the data middle platform can be regarded as a middleware that aggregates and governs cross-domain data, abstracts and encapsulates data into services, and provides the logical concept of application value to the front end.
[0004] The data middle platform can provide data services through APIs instead of directly handing over the database to the front-end and allowing the front-end developers to use the data on their own.
[0005] Whether it is an Internet company, an enterprise, or traditional industry and agriculture, they are all carrying out data middle-office projects, starting from top-level design to sort out comprehensive data application scenarios, and gradually eliminating data silos and solving data chimney construction problems during implementation.
[0006] However, during the implementation of data service components (such as the data middle platform), the main technical problem faced is the construction status of the data service components, that is, how to evaluate the construction status of the data service components.
[0007] Therefore, a solution capable of evaluating data service components is needed. Summary of the Invention
[0008] A technical problem to be solved by the present disclosure is to provide a solution capable of evaluating data service components.
[0009] According to a first aspect of the present disclosure, a method for managing a data service component is provided, wherein the data service component is used to process acquired data according to a data model structure to obtain data that conforms to the data model structure, the data model structure comprising a first data set and a second data set, the first data set comprising one or more first data, the second data set comprising one or more second data, each type of first data corresponding to a data source, the first data coming from its corresponding data source, each type of second data corresponding to a subject domain, and the second data being data generated based on at least one type of first data and belonging to a subject in the subject domain corresponding to the second data. The method comprises: determining, for a type of first data, a first fusion degree of the first data based on a dependency relationship between second data corresponding to at least one subject domain and the first data, the first fusion degree being used to characterize a degree of data flow and integration of the first data within the at least one subject domain corresponding to the second data set; determining a second fusion degree based on the first fusion degree of the one or more first data, the second fusion degree being used to characterize a degree of data flow and integration of the first data set within the at least one subject domain corresponding to the second data set, the second fusion degree being positively correlated with the first fusion degree; and evaluating the data service component based on the second fusion degree.
[0010] According to a second aspect of the present disclosure, a management method for a data service component is provided, wherein the data service component is used to process acquired data according to a data model structure to obtain data that conforms to the data model structure, the data model structure includes a second data set and a third data set, the second data set includes one or more second data, each type of second data corresponds to a subject domain, the second data is data belonging to a subject in the subject domain corresponding to the second data, the third data set includes one or more third data, each type of third data corresponds to an application indicator, and the third data is data generated based on at least one second data for characterizing the application indicator corresponding to the third data, the method including: determining a third fusion degree based on a dependency relationship between the third data corresponding to at least one application indicator and the second data corresponding to at least two subject domains, the third fusion degree being used to characterize the degree of data flow and integration between different subject domains in the at least two subject domains corresponding to the second data set; and evaluating the data service component based on the third fusion degree.
[0011] According to a third aspect of the present disclosure, a method for managing a data service component is provided, wherein the data service component is used to process acquired data according to a data model structure to obtain data that conforms to the data model structure. The method includes: evaluating the data service component; and outputting evaluation information and / or modification prompts for the data service component based on the evaluation results.
[0012] According to a fourth aspect of the present disclosure, a data service component management device is provided, wherein the data service component is configured to process acquired data according to a data model structure to obtain data that conforms to the data model structure, the data model structure comprising a first data set and a second data set, the first data set comprising one or more first data, the second data set comprising one or more second data, each type of first data corresponding to a data source, the first data coming from its corresponding data source, each type of second data corresponding to a subject domain, and the second data being data generated based on at least one type of first data and belonging to a subject in the subject domain corresponding to the second data. The device comprises: a first fusion degree determination module configured to determine, for a type of first data, a first fusion degree of the first data based on a dependency relationship between second data corresponding to at least one subject domain and the first data, the first fusion degree being used to characterize a degree of data flow and integration of the first data within the at least one subject domain corresponding to the second data set; a second fusion degree determination module configured to determine, based on the first fusion degrees of the one or more first data, a second fusion degree being used to characterize a degree of data flow and integration of the first data within the at least one subject domain corresponding to the second data set, the second fusion degree being positively correlated with the first fusion degree; and an evaluation module configured to evaluate the data service component based on the second fusion degree.
[0013] According to a fifth aspect of the present disclosure, a data service component management device is provided, wherein the data service component is used to process the acquired data according to a data model structure to obtain data that conforms to the data model structure, the data model structure includes a second data set and a third data set, the second data set includes one or more second data, each type of second data corresponds to a subject domain, the second data is data belonging to a subject in the subject domain corresponding to the second data, the third data set includes one or more third data, each type of third data corresponds to an application indicator, and the third data is data generated based on at least one second data for characterizing the application indicator corresponding to the third data, the device includes: a third fusion degree determination module, used to determine a third fusion degree based on a dependency relationship between the third data corresponding to at least one application indicator and the second data corresponding to at least two subject domains, the third fusion degree being used to characterize the degree of data fusion between different subject domains in the at least two subject domains corresponding to the second data set; and an evaluation module, used to evaluate the data service component based on the third fusion degree.
[0014] According to the sixth aspect of the present disclosure, a data service component management device is provided, wherein the data service component is used to process acquired data according to a data model structure to obtain data that conforms to the data model structure. The device includes: an evaluation module for evaluating the data service component; and an output module for outputting evaluation information and / or modification prompts for the data service component based on the evaluation results.
[0015] According to a seventh aspect of the present disclosure, a computing device is provided, comprising: a processor; and a memory on which executable code is stored, and when the executable code is executed by the processor, the processor executes the method described in any one of the first to third aspects above.
[0016] According to an eighth aspect of the present disclosure, a non-temporary machine-readable storage medium is provided, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor executes the method described in any one of the first to third aspects above.
[0017] According to a ninth aspect of the present disclosure, a computer program product is provided, comprising an executable code. When the executable code is executed by a processor of an electronic device, the processor is caused to execute the method as described in any one of the first to third aspects above.
[0018] In an exemplary embodiment of the present disclosure, the second fusion degree obtained by calculation, which is used to characterize the degree of data flow and integration of the first data set in at least one subject domain corresponding to the second data set, can reflect the construction quality of the data model structure of the data service component to a certain extent, so that the data service component can be quantitatively evaluated based on the second fusion degree.
[0019] In some embodiments, the third fusion degree obtained by further calculation can reflect the degree of data flow and integration between at least two subject domains corresponding to the second data set, so that the calculated third fusion degree can also reflect the construction quality of the data model structure of the data service component to a certain extent, thereby enabling the data service component to be quantitatively evaluated based on the third fusion degree.
[0020] In some embodiments, the fourth fusion degree further determined based on the second fusion degree and the third fusion degree can not only reflect the degree of data flow integration of the first data set within at least one subject domain corresponding to the second data set, but also reflect the degree of data flow integration between different subject domains corresponding to the second data set, thereby enabling a comprehensive and accurate quantitative evaluation of the overall construction quality of the data model structure of the data service component based on the fourth fusion degree. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings, wherein like reference numerals generally represent like components in the exemplary embodiments of the present disclosure.
[0022] Figure 1 A schematic diagram showing a data model structure of a data service component and quantitative calculation of the quality of the data model structure according to an embodiment of the present disclosure is shown.
[0023] Figure 2 A schematic flowchart of a method for managing a data service component according to an embodiment of the present disclosure is shown.
[0024] Figure 3 A schematic diagram of the hierarchical structure of the data model structure of the data middle platform according to an embodiment of the present disclosure is shown.
[0025] Figure 4 A schematic flowchart of a data middle platform management method according to an embodiment of the present disclosure is shown.
[0026] Figure 5 A simulation diagram of the data model structure of the data middle platform according to an embodiment of the present disclosure is shown.
[0027] Figure 6A structural diagram of a data service component management device according to an embodiment of the present disclosure is shown.
[0028] Figure 7 A structural diagram of a data service component management device according to another embodiment of the present disclosure is shown.
[0029] Figure 8 A structural diagram of a data service component management device according to another embodiment of the present disclosure is shown.
[0030] Figure 9 A schematic structural diagram of a computing device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0031] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although preferred embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0032] The data service component is primarily used to process acquired data according to a pre-defined data model structure to obtain data that conforms to the data model structure. After the data service component is constructed or completed, its construction quality needs to be evaluated. The technical concept of this disclosure is to evaluate the data service component by quantitatively calculating the quality of its data model structure.
[0033] Figure 1 The following is a schematic diagram showing the data model structure of a data service component according to an embodiment of the present disclosure and the quantitative calculation of the quality of the data model structure. Figure 1 As shown, the data model structure of the data service component may include a first data set, a second data set, and a third data set.
[0034] The first data set includes one or more first data, each type of first data corresponds to a data source, and the first data comes from the corresponding data source.
[0035] Data sources provide raw data to data service components for processing. Each data source can correspond to a specific scenario, and all data generated within that scenario can be considered to originate from that data source. For example, if data source A corresponds to a marketing scenario, marketing activity data, prize redemption data, and other marketing-related data can all be considered to originate from data source A.
[0036] The second dataset includes one or more types of second data, each type of second data corresponding to a subject domain. The second data is generated based on at least one type of first data and belongs to a subject within the subject domain corresponding to the second data. The second dataset can be considered a collection of data divided according to the subject domain, obtained by processing the data in the first dataset, namely, the one or more types of second data.
[0037] In this disclosure, a topic is an abstract concept used to synthesize, categorize, and analyze the data in the first dataset at a higher level. Each topic essentially corresponds to a (macro) analytical domain. Logically, it corresponds to the analytical object involved in a specific (macro) analytical domain (within the enterprise). Simply put, one topic corresponds to one analytical object.
[0038] The subject domain is a division of topics, and each subject domain can correspond to a subject set. Usually, closely related topics can be divided into a subject domain. Alternatively, the subject domain can be divided according to the focus of the service object (such as an enterprise). The subject domain can include a first-level subject domain and a second-level subject domain. The second-level subject domain is a subject domain that is further divided from the first-level subject domain. Taking the data service component as the industrial data middle platform as an example, the subject domain (such as the first-level subject domain) can include but is not limited to one or more of the following: order (order, abbreviated as ord), raw materials (material, abbreviated as mat), equipment (machine, abbreviated as mc), organization (organization, abbreviated as org), production (production, abbreviated as prd), product (product, abbreviated as pro), customer (customer, abbreviated as cus), human resources (human_resouce, abbreviated as hr), finance (finance, abbreviated as fin).
[0039] The third data set includes one or more third data, each third data corresponds to an application indicator, and the third data is data generated based on at least one second data and used to characterize the application indicator corresponding to the third data.
[0040] The generation of the third data in the third data set depends on the second data in the second data set. The dependency mentioned here can include direct dependency and / or indirect dependency. Direct dependency means that the generation of the third data directly depends on one or more second data in the second data set, and indirect dependency means that the generation of the third data may indirectly depend on one or more second data in the second data set.
[0041] As an example, Figure 1The data model structure of the data service component shown may also include a fourth data set. The fourth data set may be a collection of data of common dimensions abstracted from analyzing the data in the second data set. For example, the fourth data set may be service data for analyzing a subject domain obtained by lightly summarizing the data in the second data set, generally a wide table. Thus, the generation of third data may depend on the data in the fourth data set, and the generation of data in the fourth data set in turn depends on the second data in the second data set. Therefore, the third data generated based on the data in the fourth data set can be considered to be indirectly dependent on the second data on which the generation of the data in the fourth data set depends.
[0042] For some or all of Figure 1 The data model structure shown processes the data service component of the data. The present disclosure proposes two fusion degree calculation methods: intra-domain fusion degree calculation method and inter-domain fusion degree calculation method.
[0043] You can choose either of these two calculation methods and perform a quantitative evaluation of the quality of the data service component based on the calculation results of one of the calculation methods; you can also choose both calculation methods and perform a quantitative evaluation of the quality of the data service component based on the calculation results of the two calculation methods.
[0044] Data fusion can be defined in both broad and narrow terms. Broadly speaking, data fusion refers to all data fusion activities. Narrowly, data fusion primarily refers to multi-sensor fusion. The data fusion discussed in this disclosure falls within the broad concept of data fusion. Data fusion encompasses both the flow and integration of data. This disclosure uses the concept of fusion degree to characterize the degree of data flow and integration.
[0045] The degree of intra-domain fusion (i.e., the second degree of fusion described below) describes the degree of data fusion of the first dataset within one or more subject domains corresponding to the second dataset. The degree of data fusion of the first dataset within one or more subject domains corresponding to the second dataset can reflect the data flow of one or more first data in the first dataset within one or more subject domains corresponding to the second dataset. If certain first data in the first dataset "flows" to a certain subject domain, the generation of the second data corresponding to the subject domain depends on this first data; if certain first data in the first dataset does not "flow" to a certain subject domain, the generation of the second data corresponding to the subject domain does not depend on this first data.
[0046] The degree of inter-domain integration (also referred to as the third degree of integration below) describes the degree of data integration across subject domains (i.e., between different subject domains corresponding to the second dataset). The degree of data flow and integration between different subject domains corresponding to the second dataset can reflect the data integration between the second data corresponding to different subject domains. Data integration exists between two subject domains, meaning that third data exists that is generated by relying on the second data corresponding to both subject domains. Data integration does not exist between two subject domains, meaning that third data does not exist that is generated by relying on the second data corresponding to both subject domains.
[0047] The data flow represented by the intra-domain fusion degree and the data integration between different subject domains represented by the inter-domain fusion degree are both related to the dependency relationship between data. Therefore, the "fusion degree" mentioned in this disclosure can also be expressed as "dependency".
[0048] The various aspects of the present disclosure are further described below with reference to specific embodiments.
[0049] Figure 2 A schematic flow chart of a method for managing a data service component according to one embodiment of the present disclosure is shown. The data service component is configured to process acquired data according to a data model structure to obtain data that conforms to the data model structure. The data model structure includes a first data set and a second data set. For details about the first and second data sets, see the relevant descriptions above. Figure 2 The method can be implemented in software by a computer program or executed by a specially configured computing device. Figure 2 The method shown.
[0050] See also Figure 2 In step S110, for a type of first data, a first fusion degree of the type of first data is determined based on a dependency relationship between the type of first data and second data corresponding to at least one subject domain.
[0051] The first integration degree is used to represent the degree of data flow and integration of the first data within the at least one subject domain corresponding to the second dataset. The at least one subject domain can be all categories of subject domains corresponding to the second dataset, or can be a subset of all categories of subject domains corresponding to the second dataset.
[0052] The dependency relationship between the second data corresponding to at least one subject domain and a type of first data is used to characterize whether the generation of the second data corresponding to the at least one subject domain is dependent on the first data. By performing a lineage analysis on the second data corresponding to the at least one subject domain, an upstream link of the second data can be obtained, and based on the upstream link, whether a dependency relationship exists between the second data and the first data is determined. The upstream link characterizes the link relationship in the generation process of the second data, which reflects the data on which the generation process of the second data depends.
[0053] The first integration degree is used to characterize the degree of data flow and integration of the first data within the at least one subject domain corresponding to the second data set. Obviously, if the generation of the second data corresponding to the at least one subject domain does not depend on the first data, the degree of data flow and integration of the first data within the at least one subject domain corresponding to the second data set can be considered zero, that is, the first integration degree of the first data is zero.
[0054] Therefore, as an optional (but not necessarily preferred) embodiment, the first fusion degree of the first data can be determined based on the degree of dependence of the second data corresponding to the at least one subject domain in the second data set on the first data, and the magnitude of the first fusion degree is positively correlated with the degree of dependence. The degree of dependence can be determined based on the amount of data (such as the number of data tables) in the first data on which the generation of the second data corresponding to the at least one subject domain depends, and the degree of dependence is positively correlated with the amount of data.
[0055] In step S120, a second fusion degree is determined based on the first fusion degree of one or more first data. The one or more first data may refer to all types of first data in the first data set or to some types of first data in the first data set.
[0056] The second integration degree is also known as the intra-domain integration degree. The second integration degree is used to characterize the degree of data flow and integration of the first dataset within the at least one subject domain corresponding to the second dataset. The first integration degree of one or more first data can be obtained based on the calculation method of step S110.
[0057] The second fusion degree is positively correlated with the first fusion degree. For example, corresponding weights can be configured for the fusion degrees of one or more first data, and the second fusion degree can be the cumulative result of all first fusion degrees under the corresponding weights.
[0058] In one embodiment of the present disclosure, a first data set may include one or more groups of first data tables, each group of first data tables corresponding to a data source, each group of first data tables including one or more first data tables, and the data in the first data tables originating from the data source corresponding to the group to which the first data tables belong. A second data set may include one or more groups of second data tables, each group of second data tables corresponding to a subject domain, each group of second data tables including one or more second data tables, and the generation of the second data tables is dependent (directly and / or indirectly) on at least one first data table. Each group of first data tables may be considered as a type of first data, and each group of second data tables may be considered as a type of second data.
[0059] When determining the first fusion degree in step S110, the first fusion degree of the set of first data tables can be determined based on the dependency relationships between each second data table and the first data table in the set of first data tables. The first fusion degree can represent the degree of data fusion within the subject domain indicated by each set of second data tables in the second data set.
[0060] When determining the second fusion degree in step S120, the second fusion degree may be determined based on the first fusion degree of one or more groups of first data tables. For example, the second fusion degree may be determined based on the first fusion degree of each group of first data tables in the first data set.
[0061] As an optional (but not necessarily preferred) embodiment, the degree of dependence of each second data table on the set of first data tables can be determined based on the dependency relationship between each second data table and the first data table in the set of first data tables. Based on this dependency, a first fusion degree of the set of first data tables can be determined, where the magnitude of the first fusion degree is positively correlated with the degree of dependence. The degree of dependence can be determined based on the number of first data tables in the set of first data tables that the generation of the second data table depends on, i.e., a greater number of first data tables indicates a greater degree of dependence.
[0062] Specifically, for a group of first data tables, a first dependency set can be determined for each second data table based on the dependency relationships between each second data table and the first data table in the group of first data tables. The elements in the first dependency set are the first data tables in the group of first data tables that the generation of the second data table depends on. Then, a first fusion degree of the group of first data tables can be determined based on the number of first data tables in the group of first data tables and the first dependency sets of each second data table.
[0063] For example, the first fusion degree of the group of first data tables can be determined based on the number of first data tables in the group of first data tables and the number of elements in the first dependency set of each second data table (e.g., the total number of elements in the first dependency sets of all second data tables). The first fusion degree can be negatively correlated with the number of first data tables and positively correlated with the total number of elements in the first dependency sets of all second data tables (which can be the number of non-repeating elements).
[0064] For another example, the number of first fusion directions of the group of first data tables can be determined based on the first dependency set of each second data table, where the first fusion direction is used to characterize the actual fusion direction of the first data table in the group of first data tables to the second data table in the second data set; the number of second fusion directions of the group of first data tables can be determined based on the number of first data tables in the group of first data tables, where the second fusion direction is used to characterize the feasible fusion direction of the first data table in the group of first data tables to the second data table in the second data set; the first fusion degree of the group of first data tables can be determined based on the number of first fusion directions and the number of second fusion directions, where the first fusion degree is positively correlated with the number of first fusion directions and negatively correlated with the number of second fusion directions.
[0065] When determining the number of first fusion directions of the group of first data tables, it is possible to first determine whether the first dependency set is valid based on a preset rule, and determine the number of first fusion directions of the group of first data tables based on the number of binary combinations of each valid first dependency set, wherein the number of first fusion directions is positively correlated with the number of binary combinations of each valid first dependency set, such as the number of first fusion directions can be the cumulative result of the number of binary combinations of each valid first dependency set. The number of binary combinations of the first dependency set, that is, the number of binary combinations of the number of elements in the first dependency set, can be set to 1 when the number of elements in the first dependency set is 1.
[0066] The preset rules may be set as follows: if the first dependency set is an empty set, the first dependency set is invalid; and / or if the first dependency set is a subset of another first dependency set, the first dependency set as a subset is invalid; and / or of multiple first dependency sets with the same set elements, only one remains valid. The above three conditions are preferably in an AND relationship.
[0067] When determining the number of second fusion directions of the group of first data tables according to the number of first data tables in the group of first data tables, the number of second fusion directions of the group of first data tables can be determined according to the number of combinations of two of the number of first data tables in the group of first data tables. The number of second fusion directions is positively correlated with the number of combinations of two of the number of first data tables in the group of first data tables. For example, the number of second fusion directions can be equal to the number of combinations of two of the number of first data tables in the group of first data tables.
[0068] When determining the second fusion degree based on the first fusion degree of one or more groups of first data tables, a weight of the first fusion degree can be set, the weight is positively correlated with the number of first data tables in the group corresponding to the first fusion degree, and negatively correlated with the number of all first data tables in the first data set, and the accumulated result of the first fusion degree of each group of first data tables under the corresponding weight is used as the second fusion degree.
[0069] As an example, the first fusion degree of a group of first data tables corresponding to the i-th data source in the first data set can be expressed as:
[0070]
[0071] Among them, n i is the number of first data tables corresponding to the i-th data source in the first data set; n i The number of combinations of two, that is k is the number of all second data tables in the second data set; a ij is the first dependency set of the jth second data table, that is, the set of first data tables corresponding to the ith data source on which the generation of the jth second data table depends. card(a ij ) is a set a ij The number of elements in, for example, a 56 ={ods 51 ,ods 55} indicates that the generation of the sixth second data table in the second data set depends on the two first data tables in the first data table set corresponding to the fifth data source, that is, For card(a ij ) takes the number of two combinations, that is f ij Used to represent the set a ij The effectiveness of f ij The value of f is 0 or 1. ij =1, it means set a ij Effective, f ij =0, it means set a ij invalid.
[0072] The number of second fusion directions can be regarded as the maximum value that the first fusion direction can theoretically reach, that is, theoretically the maximum number of first fusion directions of all first data tables corresponding to the same data source in the first data set is the number of two combinations of all first data tables corresponding to the data source.
[0073] According to the development specifications of the second data set, the generation of the second data table depends on the first data table. Therefore, the actual number of the first fusion direction can be obtained by performing a lineage analysis on the second data table and its upstream dependent first data table.
[0074] The calculation formula of the second fusion degree (i.e., the intra-domain fusion degree) can be expressed as:
[0075]
[0076] Wherein, m is the number of data source types corresponding to the first data set, and n is the number of all first data tables in the first data set. For the meaning of other symbols in the formula, please refer to the relevant description above.
[0077] When calculating the second fusion degree based on the above calculation formula, all second data tables can be scanned in sequence. For example, when scanning the j-th second data table, the first data table under the upstream i-th data source can be found through lineage analysis. The result set is assumed to be a ij ={ods ix ,ods iy}, f ij =1(set a ij valid flag bit).
[0078] When scanning the kth second data table (k>j), the result set is a ik , when the result set a of the kth second data table ik When included in any previous set, the valid mark bit of the new set is 0; when a ik When a contains (and is not equal to) a previous set, the valid mark bit of the previous set is 0; if a ik If the set does not contain the previous set, the mark bit defaults to 1.
[0079] Therefore, the effective rule of the set can be set as follows: if a ik is an empty set, then f ik =1; if and satisfy Then f ij =1,f ik =0; if and satisfy Then f ij =0,f ik =1, otherwise f ij =1,f ik =1.
[0080] Therefore, by analyzing the lineage of each second data table in the second data set in the above manner, the subset matrix A of the first data table that each second data table depends on upstream can be obtained. mk and the effective value matrix Fmk .
[0081] According to the subset matrix A mk and the effective value matrix F mk , the number of actual fusion directions (i.e., first fusion directions) of the first data table corresponding to a certain data source (e.g., the i-th data source) in the first data set in the second data set can be expressed as:
[0082]
[0083] Considering the number of fusion directions (i.e., second fusion directions) of the first data table corresponding to the i-th data source, the fusion degree of the first data table corresponding to a certain data source (e.g., the i-th data source) in the first dataset in the second dataset can be expressed as:
[0084]
[0085] Please note that card(a ij )=0, take 2 combinations as 0; card(a ij )=1, the number of 2 combinations is set to 1 (i.e. there is only one fusion direction); similarly, n i =0, the number of combinations of 2 is 0; n i =1, the number of combinations is set to 1 (i.e., there is only one fusion direction).
[0086] Furthermore, after calculating the fusion degree of the first data tables of different data sources, the weighted sum of the different fusion degrees is performed according to the proportion of the number of first data tables, and the calculation formula of the second fusion degree shown above can be obtained, that is:
[0087]
[0088] In step S130 , the data service component is evaluated based on the second fusion degree.
[0089] The greater the second fusion degree, the better the data flow and integration effect of the first data set in the at least one subject domain corresponding to the second data set; the smaller the second fusion degree, the worse the data flow and integration effect of the first data set in the at least one subject domain corresponding to the second data set.
[0090] Therefore, the second fusion degree can be used to reflect the construction quality of the first data set and the second data set, that is, the construction quality of the data model structure, and the construction quality of the data model structure can characterize the construction quality of the data service component. Therefore, the data service component can be evaluated based on the second fusion degree, that is, the data service component can be quantitatively evaluated based on the second fusion degree.
[0091] As an example, first evaluation information may be generated and output based on the second fusion degree, so that users (such as data service builders or acceptance personnel) can understand the construction quality of the data service component based on the first evaluation information.
[0092] The first evaluation information can be the overall evaluation information for the data service component or the data model structure of the data service component. The first evaluation information can be divided according to the level interval, such as the initial level, intra-domain level, inter-domain level, robust level, and perfect level can be pre-set. Each level interval corresponds to a certain degree of data fusion. The value of the second fusion degree can be used to determine which level interval the current data service component falls into, so that designers can understand the construction status of the data service component. The first evaluation information can also be information specifically used to characterize the construction quality of the data model structure, such as evaluation information that characterizes the data flow and integration effect of the first data set in the at least one subject domain corresponding to the second data set, so that users can understand the construction status of the data service component in detail.
[0093] Optionally, the first fusion degree of one or more first data obtained during the calculation process may also be output to the user so that the user can gain an in-depth understanding of the construction quality of the data model structure followed by the data service component when processing data.
[0094] The data model structure of the data service component may further include a third data set. For details about the third data set, please refer to the above description, which will not be repeated here.
[0095] The present disclosure may also determine a third degree of integration based on a dependency relationship between third data corresponding to at least one application indicator and second data corresponding to at least two subject domains. The third degree of integration is used to characterize the degree of data flow and integration between different subject domains in the at least two subject domains corresponding to the second data set. This third degree of integration is also referred to as the inter-domain integration degree.
[0096] For the at least two subject domains, if the generation of the third data in the third data set only depends on the second data corresponding to one of the at least two subject domains, or the subject domains corresponding to the second data on which the generation of the third data in the third data set depends do not belong to the at least two subject domains, then it can be considered that the degree of data flow and integration between different subject domains in the at least two subject domains is zero, that is, the third fusion degree is zero.
[0097] Based on the above considerations, the present disclosure proposes that, based on the dependency relationship between the third data and the second data in the second data set, it can be determined whether the generation of the third data depends on the second data corresponding to two or more of the at least two subject domains. If the generation of the third data does not depend on the second data corresponding to two or more of the at least two subject domains, then the contribution of the third data to the third fusion degree is considered to be zero. In this way, the third fusion degree can be determined based on the contribution of various third data to the third fusion degree. The contribution of the third data to the third fusion degree can be positively correlated with the number of categories belonging to the at least two subject domains in the subject domain corresponding to the second data on which the generation of the third data depends.
[0098] In one embodiment of the present disclosure, the second data set includes one or more sets of second data tables, each set of second data tables corresponding to a subject domain, and each set of second data tables includes one or more second data tables. The third data set includes one or more third data tables, each third data table corresponding to an application indicator, and the generation of the third data table depends on at least one second data table.
[0099] When determining the third fusion degree, the third fusion degree may be determined based on the dependency relationship between each third data table and the second data table.
[0100] As an example, based on the dependency relationship between each third data table and the second data table, the second dependency set of each third data table can be determined, and the elements in the second dependency set are the subject domains corresponding to the second data table on which the generation of the third data table depends; according to the number of subject domains corresponding to the second data set and the second dependency set of each third data table, the third degree of fusion is determined.
[0101] Specifically, the number of third fusion directions can be determined based on the second dependency set of each third data table, and the third fusion direction is used to characterize the actual fusion direction between the subject domains corresponding to each group of second data tables in the second data set; the number of fourth fusion directions can be determined based on the number of subject domains corresponding to each group of second data tables in the second data layer set (that is, the number of subject domains corresponding to the second data layer), and the fourth fusion direction is used to characterize the feasible fusion direction between the subject domains corresponding to each group of second data tables in the second data set; the third fusion degree is determined based on the number of third fusion directions and the number of fourth fusion directions, wherein the third fusion degree is positively correlated with the number of third fusion directions and negatively correlated with the number of fourth fusion directions.
[0102] When determining the number of third fusion directions, it is possible to first determine whether the second dependency set is valid based on a preset rule, and determine the number of third fusion directions based on the number of binary combinations of each valid second dependency set, wherein the number of third fusion directions is positively correlated with the number of binary combinations of each valid second dependency set. For example, the number of third fusion directions can be the cumulative result of the number of binary combinations of each valid second dependency set. The number of binary combinations of the second dependency set, that is, the number of binary combinations of the number of elements in the second dependency set, when the number of elements in the second dependency set is 1, its number of binary combinations can be set to 0.
[0103] The preset rules may be set as follows: if the first dependency set is an empty set, the first dependency set is invalid; and / or if the first dependency set is a subset of another first dependency set, the first dependency set as a subset is invalid; and / or of multiple first dependency sets with the same set elements, only one remains valid. The above three conditions are preferably in an AND relationship.
[0104] When determining the number of fourth fusion directions based on the number of subject domains corresponding to each group of second data tables in the second data set, the number of fourth fusion directions can be determined based on the number of two combinations of the number of subject domains corresponding to each group of second data tables in the second data set, wherein the number of fourth fusion directions is positively correlated with the number of two combinations of the number of subject domains corresponding to each group of second data tables in the second data set, such as the number of fourth fusion directions can be equal to the number of two combinations of the number of subject domains corresponding to each group of second data tables in the second data set.
[0105] The number of the fourth fusion direction can be regarded as the maximum value that the third fusion direction can theoretically reach, that is, theoretically, the maximum number of actual fusion directions between the subject domains corresponding to each group of second data tables in the second data set is the number of combinations of two of the number of subject domains of all categories.
[0106] According to the development specifications of the third data set, the generation of the third data table depends on the second data table. Therefore, the actual number of the third fusion direction can be obtained by performing a lineage analysis on the third data table and its upstream dependent second data table.
[0107] Therefore, the calculation formula of the third fusion degree can be expressed as:
[0108]
[0109] Where p is the number of all third data tables in the third data set; s is the number of categories of the subject domain (such as the first-level subject domain) corresponding to the second data set; b 1j is the second dependency set of the jth third data table, that is, the subject domain set corresponding to the second data table on which the generation of the jth third data table depends (including direct and indirect dependencies); card (b1j ) is set b 1j The number of elements in b 12 ={ord, prd} indicates that the second third data table is dependent on the data corresponding to the two subject domains "ord (order)" and "prd (production)" in the second data set, that is, it depends on the second data table corresponding to the two subject domains in the second data set. Then card(b 12 )=2;t 1j Represents set b 1j The effectiveness of t 1j The value is 0 or 1, t 1j =1, indicating set b 1j Effective, t 1j = 0, indicating set b 1j invalid.
[0110] The idea of calculating the third fusion degree based on the above calculation formula is similar to that of calculating the second fusion degree. For example, all third data tables can be scanned in sequence. For example, when scanning the jth third data table, the upstream dependent subject domain can be found through lineage analysis. The result set is assumed to be b 1j ={ord,prd},t 1j =1(set b 1j valid flag bit).
[0111] When scanning the kth third data table (k>j), the result set is b 1k , when the result set b of the kth third data table 1k When included in any previous set, the valid mark bit of the new set is 0; when b 1k When it contains (and is not equal to) a previous set, the valid mark bit of the previous set is 0; if b 1k If the set does not contain the previous set, the mark bit defaults to 1.
[0112] Therefore, the effective rule of the set can be set as follows: if b 1k is an empty set, then t 1k =1; if and satisfy Then t 1j =1,t 1k =0; if and satisfy Then t 1j =0,t 1k =1, otherwise t 1j =1,t 1k =1.
[0113] Therefore, by analyzing the lineage of each third data table in the third data set in the above manner, the subset matrix B of the subject domain that each third data table depends on in the upstream can be obtained. 1p And the effective value matrix T 1p , p is the number of the third data table.
[0114] According to the subset matrix B 1p And the effective value matrix T 1p , the number of actual fusion directions (i.e., third fusion directions) between the subject domains corresponding to each group of second data tables in the second data set can be expressed as:
[0115]
[0116] Taking into account the number of fusion directions (i.e., fourth fusion directions) between the subject domains corresponding to each group of second data tables in the second dataset, the degree of data fusion between different subject domains in the second dataset (i.e., third fusion degree) can be expressed as:
[0117]
[0118] It should be noted that the third degree of integration mentioned in this disclosure represents the degree of data flow and integration across subject domains. When the number of elements in the second dependency set of the third data table is 1, it indicates that the generation of the third data table only depends on the data of the same subject domain (i.e., the second data table corresponding to the same subject domain), and cannot reflect the data integration across subject domains. Therefore, card (b 1j )=1, the number of combinations can be set to 0.
[0119] The present disclosure can also obtain a fourth fusion degree based on the second fusion degree (intra-domain fusion degree) and the third fusion degree (inter-domain fusion degree) to reflect the overall data fusion effect of the data model structure of the data service component, wherein the fourth fusion degree is positively correlated with the second fusion degree and the third fusion degree respectively.
[0120] As an example, you can assign weights to the second and third fusion degrees, respectively, and use the sum of the second and third fusion degrees at the corresponding weights as the fourth fusion degree. For example, if the second and third fusion degrees are assigned equal weights, the fourth fusion degree can be equal to 0.5 × the second fusion degree + 0.5 × the third fusion degree.
[0121] The fourth degree of fusion can reflect the quality of the construction from the first dataset to the second dataset, and the quality of the construction from the second dataset to the third dataset. Therefore, the fourth degree of fusion can reflect the overall construction quality of the data model structure, and the construction quality of the data model structure can also represent the construction quality of the data service component. Therefore, the data service component can be evaluated based on the second degree of fusion, that is, the data service component can be quantitatively evaluated based on the second degree of fusion.
[0122] As an example, second evaluation information may be generated and output based on the fourth degree of integration, so that users (such as data service builders or acceptance personnel) can understand the construction quality of the data service component based on the first evaluation information.
[0123] The second evaluation information can be overall evaluation information for the data service component or its data model structure. The second evaluation information can be divided into different levels, such as initial, intra-domain, inter-domain, robust, and complete. Each level corresponds to a certain degree of data integration. The fourth integration degree can be used to determine which level the current data service component falls into, allowing designers to understand the construction status of the data service component.
[0124] Considering that the implementation of the disclosed solution requires that the data model structure of the data service component conform to certain design specifications, before executing the disclosed solution, it is possible to determine whether the data model structure of the data service component conforms to the design specifications. If the data model structure does not conform to the design specifications, a prompt to modify the data model structure is issued. For details about the design specifications, please refer to the relevant description below and will not be repeated here.
[0125] The present disclosure may also maintain the data model structure of the data service component, such as maintaining the first data set and the second data set, and / or maintaining the third data set. That is, the present disclosure may also include a process for creating the data model structure of the data service component.
[0126] The present disclosure may also adjust the data service component based on the evaluation results, and / or may also adjust the data service component based on external input. Among them, the external input may be the adjustment information set by the user for the data service component. Adjusting the data service component may refer to adjusting the data model structure of the data service component to improve the quality of the data model structure. For example, the data model structure may be adjusted with the purpose of making the fusion degree (second fusion degree and / or third fusion degree and / or fourth fusion degree) calculated based on the adjusted data model structure greater than that before the adjustment. At this time, the data service component may process the acquired data according to the adjusted data model structure.
[0127] Data service components can be but are not limited to data middle platforms, data warehouses, IData data fusion engines, and other systems or devices with data processing capabilities.
[0128] For example, the data service component of the data center platform can provide reusable data technology and data service capabilities. The purpose of establishing a data center platform is to integrate all enterprise data, break down data barriers, and eliminate inconsistent data standards and calibers.
[0129] The data middle platform can be applied to, but not limited to, various industries such as industrial and mining enterprises, retail, and cloud computing. It can be built for enterprises in these industries. For example, after an enterprise has accumulated a certain amount of data through production and operation activities, a data middle platform can be built based on its characteristics (such as its corporate framework and development model) to enhance its ability to convert data into value.
[0130] The data source of the data center can be the enterprise's global data, including databases, log data, embedded data, crawler data, external data, etc. The data source can be structured data, unstructured data, or semi-structured data.
[0131] The data model construction approach for the data middle platform is similar to that of a data warehouse. A data warehouse is a subject-oriented, integrated, relatively stable data collection that reflects historical changes. A data warehouse is a relatively specific functional concept, storing and managing a collection of data on one or more subjects, primarily providing analytical reporting. The data middle platform is an enterprise-level logical concept that embodies the enterprise's D2V (Data to Value) capabilities, primarily providing services through data APIs.
[0132] Figure 3 A schematic diagram of the hierarchical structure of the data model structure of the data middle platform according to an embodiment of the present disclosure is shown.
[0133] like Figure 3 As shown in the figure, the data model structure of the data middle platform can be implemented as a layered architecture, which can be divided into four layers from bottom to top: operational data layer (Operational Data Store, ODS), detailed data layer (Data Warehouse Detail, ODS), data light summary layer (Data Warehouse Summary, DWS) and tag / mart layer (Tag and Data Mart, TDM).
[0134] The operational data layer corresponds to the first dataset mentioned above; the data in the detail data layer corresponds to the second dataset mentioned above; the lightly aggregated data layer corresponds to the fourth dataset mentioned above; and the data in the tag / marketplace layer corresponds to the third dataset mentioned above. The lightly aggregated data layer may be omitted.
[0135] 1. Operational data layer
[0136] The operational data layer stores raw data generated during an enterprise's production and operations. Specifically, the operational data layer connects to various enterprise source systems, allowing data from these systems to be accessed. Data accessed to the operational data layer requires no processing; that is, the operational data layer retains the raw data from the source systems (including incremental and full details).
[0137] One or more types of source systems can be set up based on the company's business model and organizational structure, with each source system representing a data source. Therefore, the data in the operational data layer can be data corresponding to one or more source systems, i.e., one or more source system data, which is also referred to as the first data above.
[0138] Each type of source system (the data source mentioned above) can describe a larger scope of scenarios, and all data generated within that scope can be considered data from that source system. For example, if the source system is a marketing system, marketing-related data such as marketing activity data and prize redemption data all belong to the marketing system data.
[0139] Taking the industrial data middle platform as an example, the source system types may include but are not limited to one or more of the following: Manufacturing Execution System (MES), Enterprise Resource Planning System (ERP), Product Lifecycle Management System (PLM), Warehouse Management System (WMS), Safety Instrumented System (SIS), Distributed Control System / Distributed Control System (DCS), Supply Chain Management System (SCM), Customer Relationship Management System (CRM), Management Information System (MIS), Advanced Planning and Scheduling System (APS), Energy Management System (EMS), and Internet of Things (IOT).
[0140] When building an enterprise's data middle platform, you can customize the addition or reduction of source system types based on the enterprise's business model and organizational structure, such as the actual system used by industrial enterprises.
[0141] 2. Detailed data layer
[0142] The detail data layer can be used to process the data in the operation data layer, such as performing operations such as cleaning, filtering, and connecting on the data in the operation data layer.
[0143] The detailed data layer processes the data from the operational data layer to obtain data divided by subject domains, that is, one or more subject domain data. For example, the detailed data layer can abstract one or more data entities from the data operation layer, each data entity corresponding to a subject domain.
[0144] Taking the industrial data middle platform as an example, the subject domain (such as the first-level subject domain) can include but is not limited to one or more of the following: order (order, abbreviated as ord), raw material (material, abbreviated as mat), equipment (machine, abbreviated as mc), organization (organization, abbreviated as org), production (production, abbreviated as prd), product (product, abbreviated as pro), customer (customer, abbreviated as cus), human resources (human_resouce, abbreviated as hr), and finance (finance, abbreviated as fin).
[0145] 3. Light data aggregation layer
[0146] The data light aggregation layer is used to perform a preliminary summary of the data in the detailed data layer, abstracting some common dimensional data. For example, the data light aggregation layer can perform a light aggregation of the detailed data layer to obtain service data for analyzing the subject domain, typically in the form of a wide table.
[0147] In the present disclosure, the data light aggregation layer can statistically summarize the data in the detailed data layer according to different granularities and dimensions to generate public indicators that can be used by the upper layer (i.e., the label / market layer) when performing data analysis and algorithm application.
[0148] 4. Tag / Market Layer
[0149] The tag / mart layer can perform tag processing and / or generate application data marts based on the data in the detailed data layer and / or the data in the light data aggregation layer to provide personalized indicator (application indicator, i.e., tag) processing and / or directly application-oriented data.
[0150] Tags use raw data (i.e., DWD layer data and / or DWS layer data) and produce directly usable, readable, and valuable data through certain logical processing.
[0151] A data mart, also known as a data market, is designed to meet the needs of specific departments or users. It is stored in a multidimensional manner, including defining dimensions, indicators to be calculated, and dimensional hierarchies, to generate data cubes for decision-making analysis needs.
[0152] Providing external services through a unified tag / market layer can ensure indicator consistency and reduce logical data silos.
[0153] The data model structure of the data center can also include a dimensional data layer (Data Warehouse Dimension, DIM). DIM is independent of DWD, DWS, and TDM and is used to provide dimension field descriptions for DWD, DWS, and TDM.
[0154] The development specifications that the data center platform needs to follow during the construction process can be pre-set. The development specifications of the data center platform can include data model structure specifications and data model naming specifications. For more information about data model structure specifications, please refer to the above combined Figure 1 The data model naming conventions can include general conventions and naming rules for data tables at each layer.
[0155] General specifications may include: 1) table names must start with an English letter and not a number; 2) table names cannot contain special characters; 3) multiple words are separated by underscores.
[0156] Furthermore, the ODS layer naming rule may be: ods_{source system type}_{source system table name}_[{refresh cycle identifier}{incremental full identifier}].
[0157] The naming rule of the DWD layer can be: dwd_{first-level subject domain}_{second-level subject domain}_{name}_{refresh cycle identifier}{incremental full identifier}.
[0158] The naming rule of the DWS layer can be: dws_{first-level subject domain}_{second-level subject domain}_{data granularity}_{name}_{statistical period}.
[0159] The naming rule of the DIM layer can be: dim_{dimension definition}.
[0160] The TDM layer naming rule can be: tdm_{application layer data domain}_{name}_[{refresh cycle identifier}{incremental full identifier}]_[{statistical period}].
[0161] If the construction of the data middle platform follows the above development specifications (especially the naming specifications), by analyzing the data table names of the data middle platform, you can obtain the hierarchical structure of the data model of the data middle platform, the data layer where the data table is located, and other relevant detailed information.
[0162] against Figure 3 The data center of the data model structure shown, or the data model structure and Figure 3 Similar data middle platforms, such as those built in accordance with the above-mentioned development specifications, can use the data service component management method described above in this disclosure to evaluate the construction status of the data middle platform.
[0163] Figure 4 A schematic flow chart of a data middle platform management method according to one embodiment of the present disclosure is shown. The data middle platform management method can be executed during the construction of the data middle platform to evaluate the construction status of the data middle platform in real time and guide the construction of the data middle platform. It can also be executed after the construction of the data middle platform is completed to inspect and accept the data middle platform.
[0164] See also Figure 4 In step S310, the complete set of data models for the production environment of the data center is obtained. Here, the data model structure and all data of the data center can be obtained. For example, the table names of the data tables at each layer can be obtained to determine the data model structure of the data center by analyzing the data table names.
[0165] In step S320, it is determined whether the model naming complies with the specification.
[0166] Whether the model naming complies with the specification can be determined based on the name of the data table obtained in step S310. For the requirements of the naming specification, please refer to the relevant description above.
[0167] If the judgment result of step S320 is that the model naming does not comply with the specifications, the data model can be modified according to the specifications to stabilize the online operation of the data center.
[0168] If the judgment result of step S320 is that the model naming complies with the specifications, step S330 can be executed to evaluate the data fusion degree of the data center.
[0169] Steps S341 to S347 are used to represent the calculation process of the intra-domain fusion degree; steps S351 to S357 are used to represent the calculation process of the inter-domain fusion degree. The calculation process of the intra-domain fusion degree and the calculation process of the inter-domain fusion degree can be performed simultaneously without any particular order.
[0170] The calculation process of the domain fusion degree is as follows:
[0171] In step S341, the intra-domain fusion degree is calculated.
[0172] In step S343, it is determined whether both the ODS layer and the DWD layer are present. If one of them is missing, the degree of fusion within the domain is 0.
[0173] If both the ODS layer and the DWD layer are available, then step S345 is executed to count the number of ODS tables and DWD tables of all source system types, and the ODS table matrix upstream of the DWD table for lineage analysis (i.e., the A table matrix mentioned above). mk ).
[0174] In step S347, the intra-domain fusion degree is calculated according to the intra-domain fusion degree calculation formula.
[0175] The specific calculation process can be found in the relevant description above and will not be repeated here. The calculation results of the domain fusion degree are used to calculate the overall data fusion degree.
[0176] The calculation process of inter-domain fusion degree is as follows:
[0177] In step S351, the inter-domain fusion degree is calculated.
[0178] In step S353, it is determined whether both the DWD layer and the TDM layer are present. If one of them is missing, the inter-domain fusion degree is 0.
[0179] If both the DWD layer and the TDM layer are available, then step S355 is executed to count the number of first-level subject domains in all DWD tables, and the DWD table matrix upstream of the TDM table (i.e., the B mentioned above) is analyzed. 1p ).
[0180] In step S357, the inter-domain fusion degree is calculated according to the inter-domain fusion degree calculation formula.
[0181] The specific calculation process can be found in the relevant description above and will not be repeated here. The calculation results of the inter-domain fusion degree are used to calculate the overall data fusion degree.
[0182] The total data fusion degree = 0.5 × intra-domain fusion degree + 0.5 × inter-domain fusion degree.
[0183] Figure 5 The following is a simulation diagram of the data model structure of the data center according to an embodiment of the present disclosure, including the model list and lineage relationship of the data center. Figure 5 The process of calculating the degree of data fusion for the data model structure shown is as follows.
[0184] 1) First, you can verify whether the data model naming and development of the data center comply with the specifications. If they do, start calculating the intra-domain integration degree and inter-domain integration degree.
[0185] 2) ODS and DWD exist, and the degree of integration in the calculation domain is calculated
[0186] like Figure 4 As shown, the number of source system types in the ODS layer is 4, and the subset matrix obtained by lineage analysis of the DWD layer data table is A 45 .in,
[0187] a 11 ={ods_erp_1,ods_erp_2,ods_erp_3},a 12 ={},a 13 ={},a 14 ={}
[0188] a 21 ={},a 22 ={ods_mes_4,ods_mes_5},a 23 ={},a 24 ={},a 25 ={}
[0189] a 31 ={},a 32 ={},a 33 ={ods_crm_6,ods_crm_7,ods_crm_8,ods_crm_9}, a 34 ={ods_crm_8,ods_crm_9}
[0190] a 41 ={},a 42 ={},a 43 ={},a 44 ={},a 45 ={ods_asp_9}
[0191] The corresponding 0-1 valid value matrix is
[0192]
[0193]
[0194] The calculation results of the intra-domain fusion degree are consistent with intuitive judgment.
[0195] If the above data center is slightly adjusted, the CRM system adds two cloud tables ods_crm_11 and ods_crm_12, and the DWD layer has no blood relationship downstream, then the above subset matrix A 45 and the effective value matrix F 45 The total number of tables and the number of tables in the corresponding CRM system change.
[0196]
[0197] 3) If DWD and TDM exist (DWS can be missing), calculate the degree of inter-domain fusion
[0198] The calculation logic of inter-domain integration is concentrated in the DWD, DWS, and TDM layers. The subject domain of DWS will inherit the DWD layer, mainly analyzing the lineage of the TDM layer data table upstream to the DWD layer.
[0199] like Figure 5 As shown in the figure, there are four first-level subject domains in the DWD layer: ord, prd, cus, pro (order, production, customer, product), and three data tables in the TDM layer. The subset matrix obtained by lineage analysis of the TDM layer data tables is B 13 .
[0200] Among them, b 11 ={ord,prd},b 12 ={prd,cus,pro},b 13 ={cus,pro},
[0201] The corresponding 0-1 effective value matrix is F 12 =(110),
[0202]
[0203] Therefore, the overall data fusion degree = 0.5 × intra-domain fusion degree + 0.5 × inter-domain fusion degree
[0204] =0.5×1+0.5×0.67
[0205] =0.835
[0206] In summary, this disclosure focuses on the degree of data integration (i.e., the degree of elimination of data islands) during the construction of the data middle platform (especially the industrial data middle platform), and proposes the concepts and calculation methods of intra-domain integration and inter-domain integration, which quantitatively reflects the overall data integration of the data middle platform in the form of numerical values, and has guiding significance for the construction and acceptance of the data middle platform. In addition, through simulation with a small industrial data middle platform model, the results calculated based on the integration degree calculation method proposed in this disclosure are also in line with intuitive theoretical expectations.
[0207] Therefore, based on the above Figure 3 The data model system of the data middle platform (i.e., the IData model system) developed according to the data model specification shown in the figure uses the division of subject domains and data lineage in the top-level design process for calculation, and provides a fusion degree calculation method that can quantify the abstract concept of data fusion degree from the theoretical and computational levels.
[0208] Moreover, in a specific embodiment, 0.5*intra-domain fusion degree + 0.5*inter-domain fusion degree can be used to characterize the overall data fusion degree of the data model. The intra-domain fusion degree describes the degree of fusion of data within the same subject domain, and the inter-domain fusion degree describes the degree of data fusion across subject domains. The data fusion degree obtained by combining the two can simultaneously reflect the degree of data fusion of the source system within the subject domain, as well as the degree of data fusion across subject domains, so that when evaluating the data middle platform or guiding the construction of the data middle platform based on the data fusion degree obtained in this disclosure, the data middle platform finally obtained can better break the physical isolation islands between systems and reduce logical data islands.
[0209] It should be noted that the data middle platform mentioned in this disclosure can not only be an industrial data middle platform, but also an internal enterprise data middle platform, as well as data middle platforms in various other industries such as retail and cloud computing. Therefore, the data integration degree of the data middle platform mentioned above can also refer to the concept of data integration degree of the internal enterprise data middle platform, as well as data middle platforms in various other industries such as retail and cloud computing. That is, this disclosure can also be used to quantitatively calculate the data integration degree indicators of other internal enterprise data middle platforms, as well as data middle platforms in other industries such as retail and cloud computing, and evaluate the corresponding data middle platforms based on the quantitative calculation results.
[0210] In one embodiment of the present disclosure, a data service component may include a second data set and a third data set. The management method of the data service component of this embodiment may include: determining a third fusion degree based on a dependency relationship between third data corresponding to at least one application indicator and second data corresponding to at least two subject domains, the third fusion degree being used to characterize the degree of data flow and integration between the at least two subject domains corresponding to the second data set; and evaluating the data service component based on the third fusion degree.
[0211] The present disclosure also provides a method for managing a data service component, wherein the data service component processes acquired data according to a data model structure to obtain data that conforms to the data model structure. The method includes: evaluating the data service component; and outputting evaluation information and / or modification prompts for the data service component based on the evaluation results. The fusion evaluation data service component obtained by the method described above can be used. The specific implementation process can be found in the relevant description above.
[0212] The present disclosure also provides a data middle platform management method, which can be executed during the construction of the data middle platform to evaluate the construction status of the data middle platform in real time and guide the construction of the data middle platform. It can also be executed after the construction of the data middle platform is completed to inspect and accept the data middle platform.
[0213] The data center management method may include: evaluating the degree of data integration of the data center; and outputting modification prompts for the data model structure of the data center based on the evaluation results, so that relevant personnel can modify the data model structure of the data center according to the prompts to improve the degree of data integration of the data center. The specific implementation process of evaluating the degree of data integration of the data center can be found in the relevant description above and will not be repeated here.
[0214] This disclosure can be used for (or similar) Figure 1 In addition to evaluating the data fusion degree of the data center of the data model structure shown in the figure, based on the data fusion degree calculation concept disclosed in the present invention, other Figure 1 The data model structure shown or similar Figure 1 The data fusion degree of other data systems (such as data warehouse and IData data fusion engine) of the data model structure shown is evaluated.
[0215] The data service component management method disclosed herein may also be implemented as a data service component management device. Figures 6 to 8 The following is a schematic diagram of the structure of the data service component management device according to multiple embodiments of the present disclosure. The functional modules of the data service component management device can be implemented by hardware, software, or a combination of hardware and software that implements the principles of the present disclosure. It will be understood by those skilled in the art that Figures 6 to 8 The functional modules described herein may be combined or divided into sub-modules to implement the principles of the invention. Therefore, the description herein may support any possible combination, division, or further limitation of the functional modules described herein.
[0216] The following is a brief description of the functional modules that the data service component management device can have and the operations that each functional module can perform. For the details involved, please refer to the relevant description above and will not be repeated here.
[0217] See also Figure 6 The data service component management device 600 may include a first fusion degree determination module 610 , a second fusion degree determination module 620 and an evaluation module 630 .
[0218] The data service component is used to process the acquired data according to the data model structure to obtain data that conforms to the data model structure. The data model structure includes a first data set and a second data set. The first data set includes one or more first data, and the second data set includes one or more second data. Each type of first data corresponds to a data source, and the first data comes from its corresponding data source. Each type of second data corresponds to a subject domain, and the second data is data generated based on at least one first data and belongs to a subject in the subject domain corresponding to the second data.
[0219] The first fusion degree determination module 610 is used to determine a first fusion degree of a first data based on a dependency relationship between second data corresponding to at least one subject domain and the first data, wherein the first fusion degree is used to characterize the degree of data flow and fusion of the first data within at least one subject domain corresponding to the second data set.
[0220] The second fusion degree determination module 620 is used to determine a second fusion degree based on the first fusion degree of one or more first data. The second fusion degree is used to characterize the degree of data flow and integration of the first data set within at least one subject domain corresponding to the second data set. The second fusion degree is positively correlated with the first fusion degree.
[0221] The evaluation module 630 is configured to evaluate the data service component based on the second fusion degree. The evaluation module 630 may generate first evaluation information based on the second fusion degree and output the first evaluation information.
[0222] As an example, the first data set may include one or more groups of first data tables, each group of first data tables corresponds to a data source, each group of first data tables includes one or more first data tables, and the data in the first data tables comes from the data source corresponding to the group to which the first data tables belong. The second data set includes one or more groups of second data tables, each group of second data tables corresponds to a subject domain, each group of second data tables includes one or more second data tables, and the generation of the second data tables depends on at least one of the first data tables.
[0223] The first fusion degree determination module 610 can determine the first fusion degree of a group of first data tables based on the dependency relationship between each second data table and the first data table in the group of first data tables, and / or the second fusion degree determination module 620 can determine the second fusion degree based on the first fusion degree of one or more groups of first data tables.
[0224] The first fusion degree determination module 610 can determine, for a group of first data tables, a first dependency set of each second data table based on the dependency relationship between each second data table and the first data table in the group of first data tables, where the elements in the first dependency set are the first data tables in the group of first data tables on which the generation of the second data table depends; and determine the first fusion degree of the group of first data tables based on the number of first data tables in the group of first data tables and the first dependency set of each second data table.
[0225] Specifically, the first fusion degree determination module 610 can determine the number of first fusion directions of the group of first data tables based on the first dependency set of each second data table, where the first fusion direction is used to characterize the actual fusion direction of the first data table in the group of first data tables to the second data table in the second data set; determine the number of second fusion directions of the group of first data tables based on the number of first data tables in the group of first data tables, where the second fusion direction is used to characterize the feasible fusion direction of the first data table in the group of first data tables to the second data table in the second data set; determine the first fusion degree of the group of first data tables based on the number of the first fusion directions and the number of the second fusion directions, where the first fusion degree is positively correlated with the number of the first fusion directions and negatively correlated with the number of the second fusion directions.
[0226] The first fusion degree determination module 610 can determine whether the first dependency set is valid based on preset rules; and determine the number of first fusion directions for the group of first data tables based on the number of two-combinations of each valid first dependency set, where the number of first fusion directions is positively correlated with the number of two-combinations of each valid first dependency set. For details about the preset rules, please refer to the relevant description above and will not be repeated here.
[0227] The first fusion degree determination module 610 can determine the number of second fusion directions of the group of first data tables based on the number of combinations of two of the number of first data tables in the group of first data tables, wherein the number of second fusion directions is positively correlated with the number of combinations of two of the number of first data tables in the group of first data tables.
[0228] The second fusion degree determination module 620 can set a weight for the first fusion degree, where the weight is positively correlated with the number of first data tables in the group corresponding to the first fusion degree and negatively correlated with the number of all first data tables in the first data set; and the accumulated result of the first fusion degrees of one or more groups of first data tables under the corresponding weights is used as the second fusion degree.
[0229] The data model structure of the data service component may also include a third data set, wherein the third data set includes one or more third data, each of the third data corresponds to an application indicator, and the third data is data generated based on at least one of the second data for characterizing the application indicator corresponding to the third data. The data service component management device 600 may also include a third fusion degree determination module for determining the third fusion degree based on the dependency relationship between the third data corresponding to at least one application indicator and the second data corresponding to at least two subject domains. The third fusion degree is used to characterize the degree of data flow and integration between different subject domains in the at least two subject domains corresponding to the second data set.
[0230] The second data set includes one or more groups of second data tables, each group of second data tables corresponds to a subject domain, and each group of second data tables includes one or more second data tables. The third data set includes one or more third data tables, and each third data table corresponds to an application indicator. The generation of the third data table depends on at least one of the second data tables. The third fusion degree determination module can determine the third fusion degree based on the dependency relationship between each of the third data tables and each of the second data tables.
[0231] The third fusion degree determination module can determine the second dependency set of each third data table based on the dependency relationship, where the elements in the second dependency set are the subject domains corresponding to the second data table on which the generation of the third data table depends; the third fusion degree is determined according to the number of subject domains corresponding to the second data set and the second dependency set of each third data table.
[0232] The third fusion degree determination module can determine the number of third fusion directions based on the second dependency set of each third data table, where the third fusion direction is used to characterize the actual fusion direction between the subject domains corresponding to each group of second data tables in the second data set; determine the number of fourth fusion directions based on the number of subject domains corresponding to each group of second data tables in the second data set, where the fourth fusion direction is used to characterize the feasible fusion direction between the subject domains corresponding to each group of second data tables in the second data set; determine the third fusion degree based on the number of third fusion directions and the number of fourth fusion directions, where the third fusion degree is positively correlated with the number of third fusion directions and negatively correlated with the number of fourth fusion directions.
[0233] The third fusion degree determination module can determine whether the second dependency set is valid based on preset rules; determine the number of third fusion directions based on the number of two combinations of each valid second dependency set, wherein the number of third fusion directions is positively correlated with the number of two combinations of each valid second dependency set.
[0234] The data service component management device 600 may also include a fourth fusion degree determination module, which is used to obtain a fourth fusion degree based on the second fusion degree and the third fusion degree to reflect the overall data fusion effect of the data middle station, wherein the fourth fusion degree is positively correlated with the second fusion degree and the third fusion degree respectively.
[0235] The data service component management device 600 may further include a generation module and an output module. The generation module is configured to generate second evaluation information based on the fourth fusion degree; and the output module is configured to output the second evaluation information.
[0236] The data service component management device 600 may also include a judgment module and a prompt module. The judgment module is used to judge whether the data model structure of the data middle platform complies with the design specifications; the prompt module is used to issue a prompt to modify the data model structure if the data model structure does not comply with the design specifications.
[0237] The data service component management apparatus 600 may further include an adjustment module configured to adjust the data service component based on an evaluation result; and / or adjust the data service component based on an external input.
[0238] See also Figure 7 The data service component management device 700 may include a third fusion degree determination module 710 and an evaluation module 730 .
[0239] The data service component is used to process the acquired data according to the data model structure to obtain data that conforms to the data model structure. The data model structure includes a second data set and a third data set. The second data set includes one or more second data, each of which corresponds to a subject domain. The second data is data belonging to a subject in the subject domain corresponding to the second data. The third data set includes one or more third data, each of which corresponds to an application indicator. The third data is data generated based on at least one second data and used to characterize the application indicator corresponding to the third data.
[0240] Third fusion degree determination module 710 is configured to determine a third fusion degree based on a dependency relationship between third data corresponding to at least one application indicator and second data corresponding to at least two subject domains. The third fusion degree is used to represent the degree of data fusion between different subject domains in the at least two subject domains corresponding to the second dataset. The calculation process of the third fusion degree is described above and is not further elaborated here.
[0241] The evaluation module 730 is configured to evaluate the data service component based on the third fusion degree.
[0242] See also Figure 8 The data service component management device 800 may include an evaluation module 810 and an output module 820. The evaluation module 810 is used to evaluate the data service component; the output module 820 is used to output evaluation information and / or modification prompts for the data service component according to the evaluation results.
[0243] The data service component management device 800 may further include a first fusion degree determination module and a second fusion degree determination module. And / or the evaluation module 810 may further include a third fusion degree determination module. Optionally, the evaluation module 810 may further include a fourth fusion degree determination module.
[0244] For details on the operations that can be performed by the first fusion degree determination module, the second fusion degree determination module, the third fusion degree determination module, and the fourth fusion degree determination module, please refer to the relevant description above and will not be repeated here.
[0245] The evaluation module 810 may evaluate the data service component based on the calculated fusion degree.
[0246] Figure 9 A schematic diagram of the structure of a computing device that can be used to implement the above-mentioned data service component management method according to an embodiment of the present disclosure is shown.
[0247] See also Figure 9 , the computing device 900 includes a memory 910 and a processor 920 .
[0248] The processor 920 may be a multi-core processor or may include multiple processors. In some embodiments, the processor 920 may include a general-purpose main processor and one or more special coprocessors, such as a graphics processing unit (GPU) or a digital signal processor (DSP). In some embodiments, the processor 920 may be implemented using customized circuits, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs).
[0249] The memory 910 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. ROM may store static data or instructions required by the processor 920 or other modules of the computer. The permanent storage device may be a readable and writable storage device. The permanent storage device may be a non-volatile storage device that does not lose stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a large-capacity storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In other embodiments, the permanent storage device may be a removable storage device (such as a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory may store some or all instructions and data required by the processor during operation. In addition, the memory 910 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks may also be used. In some embodiments, the memory 910 may include a readable and / or writable removable storage device, such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and transient electronic signals transmitted wirelessly or wired.
[0250] The memory 910 stores executable codes. When the executable codes are processed by the processor 920 , the processor 920 can execute the data service component management method described above.
[0251] The data service component management method, apparatus, and device according to the present disclosure have been described above in detail with reference to the accompanying drawings.
[0252] In addition, the method according to the present disclosure may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing the above steps defined in the above method of the present disclosure.
[0253] Alternatively, the present disclosure may also be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium) on which executable code (or computer program, or computer instruction code) is stored. When the executable code (or computer program, or computer instruction code) is executed by a processor of an electronic device (or computing device, server, etc.), the processor is caused to perform the various steps of the above-mentioned method according to the present disclosure.
[0254] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or combinations of both.
[0255] The flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems and methods according to multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0256] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for managing a data service component, wherein: The data service component is used to process the acquired data according to the data model structure to obtain data that conforms to the data model structure. The data model structure includes a first data set and a second data set. The first data set includes one or more first data, and the second data set includes one or more second data. Each type of the first data corresponds to a data source, and the first data comes from its corresponding data source. Each type of the second data corresponds to a subject domain. The second data is data generated based on at least one type of the first data and belongs to a subject domain corresponding to the second data. The method includes: For a type of first data, determining a first fusion degree of the first data based on a dependency relationship between second data corresponding to at least one subject domain and the first data, the first fusion degree being used to characterize a degree of data flow and integration of the first data within the at least one subject domain corresponding to the second data set; determining a second fusion degree based on a first fusion degree of one or more first data, the second fusion degree being used to characterize the degree of data flow and integration of the first data set within the at least one subject domain corresponding to the second data set, the second fusion degree being positively correlated with the first fusion degree; and evaluating the data service component based on the second fusion degree; The data source includes one or more of the following: production process execution system; enterprise resource planning system; product lifecycle management system; warehouse management system; safety instrument system; distributed control system / distributed control system; supply chain management system; customer relationship management system; management information system; advanced production planning and scheduling system; enterprise energy management system; Internet of Things system; The subject domain includes one or more of the following: order, raw material, equipment, organization, production, customer, human resources, and finance.
2. The method according to claim 1, wherein The first data set includes one or more groups of first data tables, each group of first data tables corresponds to a data source, each group of first data tables includes one or more first data tables, and the data in the first data tables comes from the data source corresponding to the group to which the first data tables belong. The second data set includes one or more groups of second data tables, each group of second data tables corresponds to a subject domain, each group of second data tables includes one or more second data tables, and the generation of the second data tables depends on at least one of the first data tables. The step of determining a first fusion degree of a first data according to a dependency relationship between second data corresponding to at least one subject domain and the first data comprises: determining a first fusion degree of the group of first data tables according to a dependency relationship between each second data table and a first data table in the group of first data tables, and / or The step of determining the second fusion degree according to the first fusion degree of one or more first data includes: determining the second fusion degree according to the first fusion degree of one or more groups of first data tables.
3. The method according to claim 2, wherein: The step of determining a first fusion degree of a group of first data tables according to the dependency relationship between each second data table and the first data table in the group of first data tables includes: For a group of first data tables, determining a first dependency set for each second data table based on a dependency relationship between each second data table and the first data table in the group of first data tables, where an element in the first dependency set is the first data table in the group of first data tables on which generation of the second data table depends; A first fusion degree of the group of first data tables is determined according to the number of first data tables in the group of first data tables and the first dependency sets of each second data table.
4. The method according to claim 3, wherein: The step of determining the first fusion degree of the group of first data tables according to the number of first data tables in the group of first data tables and the first dependency set of each second data table includes: determining, according to the first dependency sets of the respective second data tables, the number of first fusion directions of the group of first data tables, the first fusion directions being used to represent actual fusion directions of the first data tables in the group of first data tables into the second data tables in the second data set; determining, according to the number of first data tables in the set of first data tables, the number of second fusion directions of the set of first data tables, the second fusion directions being used to represent feasible fusion directions of the first data tables in the set of first data tables into the second data table in the second data set; A first fusion degree of the set of first data tables is determined according to the number of the first fusion directions and the number of the second fusion directions, wherein the first fusion degree is positively correlated with the number of the first fusion directions and negatively correlated with the number of the second fusion directions.
5. The method according to claim 4, wherein The step of determining the number of first fusion directions of the group of first data tables according to the first dependency sets of the second data tables includes: Determining whether the first dependency set is valid based on preset rules; Based on the number of binary combinations of each valid first dependency set, the number of first fusion directions of the set of first data tables is determined, wherein the number of first fusion directions is positively correlated with the number of binary combinations of each valid first dependency set.
6. The method according to claim 5, wherein: The preset rules are: If the first dependency set is an empty set, the first dependency set is invalid; and / or If the first dependency set is a subset of another first dependency set, then the first dependency set as a subset is invalid; and / or Among multiple first-dependent collections with the same collection elements, only one remains valid.
7. The method according to claim 4, wherein: The step of determining the number of second fusion directions of a group of first data tables according to the number of first data tables in the group of first data tables includes: The number of second fusion directions of the group of first data tables is determined according to the number of combinations of two of the number of first data tables in the group of first data tables, wherein the number of second fusion directions is positively correlated with the number of combinations of two of the number of first data tables in the group of first data tables.
8. The method according to claim 2, wherein: The step of determining the second fusion degree according to the first fusion degree of one or more groups of first data tables comprises: Setting a weight for the first fusion degree, where the weight is positively correlated with the number of first data tables in the group corresponding to the first fusion degree and negatively correlated with the number of all first data tables in the first data set; The accumulated result of the first fusion degree of one or more groups of first data tables under corresponding weights is used as the second fusion degree.
9. The method according to claim 1, wherein The step of evaluating the data service component based on the second fusion degree includes: generating first evaluation information based on the second fusion degree; and The first evaluation information is output.
10. The method according to claim 1, wherein The data model structure of the data service component further includes a third data set, the third data set including one or more third data, each type of third data corresponding to an application indicator, the third data being generated based on at least one type of second data and used to characterize the application indicator corresponding to the third data, and the method further includes: A third fusion degree is determined based on a dependency relationship between third data corresponding to at least one application indicator and second data corresponding to at least two subject domains. The third fusion degree is used to characterize the degree of data flow and integration between different subject domains in the at least two subject domains corresponding to the second data set.
11. The method according to claim 10, wherein: The second data set includes one or more groups of second data tables, each group of second data tables corresponds to a subject domain, and each group of second data tables includes one or more second data tables. The third data set includes one or more third data tables, and each third data table corresponds to an application indicator. The generation of the third data table depends on at least one second data table. The step of determining the third fusion degree based on the dependency relationship between the third data corresponding to at least one application indicator and the second data corresponding to at least two subject domains includes: determining the third fusion degree based on the dependency relationship between each of the third data tables and each of the second data tables.
12. The method according to claim 11, wherein The step of determining the third fusion degree according to the dependency relationship between each of the third data tables and each of the second data tables includes: Based on the dependency relationship, determining a second dependency set for each of the third data tables, wherein the elements in the second dependency set are subject domains corresponding to the second data tables on which generation of the third data tables depends; The third fusion degree is determined according to the number of subject domains corresponding to the second data set and the second dependency set of each of the third data tables.
13. The method according to claim 12, wherein: The step of determining the third fusion degree according to the number of subject domains corresponding to the second data set and the second dependency set of each third data table includes: determining the number of third fusion directions according to the second dependency sets of the respective third data tables, wherein the third fusion directions are used to represent actual fusion directions between the subject domains corresponding to the respective groups of second data tables in the second data set; determining the number of fourth fusion directions according to the number of subject domains corresponding to each group of second data tables in the second data set, wherein the fourth fusion directions are used to represent feasible fusion directions between the subject domains corresponding to each group of second data tables in the second data set; A third fusion degree is determined according to the number of the third fusion directions and the number of the fourth fusion directions, wherein the third fusion degree is positively correlated with the number of the third fusion directions and negatively correlated with the number of the fourth fusion directions.
14. The method according to claim 13, wherein: The step of determining the number of third fusion directions according to the second dependency sets of each of the third data tables includes: Determining whether the second dependency set is valid based on preset rules; The number of third fusion directions is determined based on the number of two-way combinations of each valid second dependency set, wherein the number of third fusion directions is positively correlated with the number of two-way combinations of each valid second dependency set.
15. The method according to claim 14, wherein The preset rules are: If the second dependency set is an empty set, the second dependency set is invalid; and / or If the second dependency set is a subset of another second dependency set, the second dependency set as a subset is invalid; and / or Of the multiple second dependent collections with the same collection elements, only one remains valid.
16. The method according to claim 13, wherein: The step of determining the number of fourth fusion directions according to the number of subject domains corresponding to each group of second data tables in the second data set includes: The number of fourth fusion directions is determined based on the number of combinations of two of the number of subject domains corresponding to each group of second data tables in the second data set, wherein the number of fourth fusion directions is positively correlated with the number of combinations of two of the number of subject domains corresponding to each group of second data tables in the second data set.
17. The method according to claim 10, further comprising: Based on the second fusion degree and the third fusion degree, a fourth fusion degree is obtained for reflecting the overall data fusion effect of the data service component. The fourth fusion degree is positively correlated with the second fusion degree and the third fusion degree respectively.
18. The method according to claim 17, further comprising: generating second evaluation information based on the fourth fusion degree; as well as The second evaluation information is output.
19. The method of claim 1, further comprising: Determining whether the data model structure of the data service component complies with design specifications; If the data model structure does not comply with the design specifications, a prompt is issued to modify the data model structure.
20. The method of claim 1, further comprising: adjusting the data service component based on the evaluation result; and / or The data service component is adjusted based on external input.
21. The method according to any one of claims 1 to 19, wherein The data service component is the industrial data middle platform.
22. The method according to claim 1, wherein The data service component is used to process the acquired data according to the data model structure to obtain data that conforms to the data model structure. The data model structure includes a second data set and a third data set. The second data set includes one or more second data, each type of second data corresponds to a subject domain, and the second data is data generated based on at least one first data and belongs to a subject in the subject domain corresponding to the second data. Each type of first data corresponds to a data source, and the first data comes from its corresponding data source. The third data set includes one or more third data, each type of third data corresponds to an application indicator, and the third data is data generated based on at least one second data and used to characterize the application indicator corresponding to the third data. The method includes: determining a third fusion degree based on a dependency relationship between third data corresponding to at least one application indicator and second data corresponding to at least two subject domains, the third fusion degree being used to characterize a degree of data flow and integration between different subject domains in the at least two subject domains corresponding to the second data set; and evaluating the data service component based on the third degree of integration; The data source includes one or more of the following: production process execution system; enterprise resource planning system; product lifecycle management system; warehouse management system; safety instrument system; distributed control system / distributed control system; supply chain management system; customer relationship management system; management information system; advanced production planning and scheduling system; enterprise energy management system; Internet of Things system; The subject domain includes one or more of the following: order, raw material, equipment, organization, production, customer, human resources, and finance.
23. A method for managing a data service component, the data service component being configured to process acquired data according to a data model structure to obtain data conforming to the data model structure, the method comprising: evaluating the data service component; as well as Outputting evaluation information and / or modification prompts for the data service component according to the evaluation results; The step of evaluating the data service component includes: The data service component is evaluated based on the degree of integration obtained using the method of any one of claims 1 to 22.
24. A data service component management device, wherein: The data service component is used to process the acquired data according to the data model structure to obtain data that conforms to the data model structure. The data model structure includes a first data set and a second data set. The first data set includes one or more first data, and the second data set includes one or more second data. Each type of the first data corresponds to a data source, and the first data comes from its corresponding data source. Each type of the second data corresponds to a subject domain. The second data is data generated based on at least one type of the first data and belongs to a subject in the subject domain corresponding to the second data. The device includes: a first fusion degree determination module configured to determine, for a type of first data, a first fusion degree of the first data based on a dependency relationship between second data corresponding to at least one subject domain and the first data, wherein the first fusion degree is used to represent a degree of data flow and smoothness of the first data within the at least one subject domain corresponding to the second data set; a second fusion degree determination module, configured to determine a second fusion degree based on the first fusion degree of one or more first data, wherein the second fusion degree is used to represent the degree of data flow and integration of the first data set within the at least one subject domain corresponding to the second data set, and the second fusion degree is positively correlated with the first fusion degree; and an evaluation module, configured to evaluate the data service component based on the second fusion degree; The data source includes one or more of the following: production process execution system; enterprise resource planning system; product lifecycle management system; warehouse management system; safety instrument system; distributed control system / distributed control system; supply chain management system; customer relationship management system; management information system; advanced production planning and scheduling system; enterprise energy management system; Internet of Things system; The subject domain includes one or more of the following: order, raw material, equipment, organization, production, customer, human resources, and finance.
25. A data service component management device, wherein: The data service component is used to process the acquired data according to the data model structure to obtain data that conforms to the data model structure. The data model structure includes a second data set and a third data set. The second data set includes one or more second data, each type of second data corresponds to a subject domain, and the second data is data generated based on at least one first data and belongs to a subject in the subject domain corresponding to the second data. Each type of first data corresponds to a data source, and the first data comes from its corresponding data source. The third data set includes one or more third data, each type of third data corresponds to an application indicator, and the third data is data generated based on at least one second data and used to characterize the application indicator corresponding to the third data. The device includes: a third fusion degree determination module, configured to determine a third fusion degree based on a dependency relationship between third data corresponding to at least one application indicator and second data corresponding to at least two subject domains, the third fusion degree being used to represent a degree of data fusion between different subject domains in the at least two subject domains corresponding to the second data set; and an evaluation module, configured to evaluate the data service component based on the third fusion degree; The data source includes one or more of the following: production process execution system; enterprise resource planning system; product lifecycle management system; warehouse management system; safety instrument system; distributed control system / distributed control system; supply chain management system; customer relationship management system; management information system; advanced production planning and scheduling system; enterprise energy management system; Internet of Things system; The subject domain includes one or more of the following: order, raw material, equipment, organization, production, customer, human resources, and finance.
26. A data service component management device, wherein the data service component is used to process acquired data according to a data model structure to obtain data that conforms to the data model structure, the device comprising: An evaluation module, configured to evaluate the data service component; An output module, configured to output evaluation information and / or modification prompts for the data service component according to the evaluation results; Wherein, the evaluation module is further used for: The data service component is evaluated based on the degree of integration obtained using the method of any one of claims 1 to 22.
27. A computing device comprising: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to perform the method according to any one of claims 1 to 23.
28. A non-transitory machine-readable storage medium having executable codes stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the method according to any one of claims 1 to 23.
29. A computer program product comprising executable codes, which, when executed by a processor of an electronic device, causes the processor to perform the method according to any one of claims 1 to 23.
Citation Information
Patent Citations
Data center station system
CN112396404A