Business modeling method, medium and system based on enterprise data asset management platform

Through the business modeling method based on the enterprise data asset management platform and the use of multi-layer neural network models and other technical means, the problems of insufficient analysis of data operation mechanisms and insufficient automation capabilities in the existing technology are solved, and the comprehensive analysis and efficient utilization of data assets are achieved, and the efficiency of information construction is improved.

CN120123314APending Publication Date: 2025-06-10BEIJING NANCAL RUIYUAN DIGITAL TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510056228.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing data asset management technology lacks in-depth analysis of the internal data operation mechanism of the enterprise, making it difficult to fully explore the value of data assets, and there are shortcomings in data asset relationship discovery, business modeling automation and service interface generation.

Method used

Provide a business modeling method based on the enterprise data asset management platform, by reading the data asset catalog, extracting foreign key relationships between data tables, calculating entity correlation, using the K-means clustering algorithm to divide business domains, statistical data table fields, record operation sequences, constructing data table operation link diagrams, evaluating the importance of data table nodes, and using multi-layer neural network models to build business models.

Benefits of technology

It realizes comprehensive analysis and modeling of enterprise data assets, provides automated business modeling methods, and can automatically generate data service interface definitions that meet business needs, improving the efficiency of enterprise information construction and data asset utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123314A_ABST
    Figure CN120123314A_ABST
Patent Text Reader

Abstract

The invention provides a business modeling method, medium and system based on an enterprise data asset management platform, and belongs to the technical field of modeling methods.The business modeling method based on the enterprise data asset management platform comprises the following steps that an enterprise data asset catalog is read to generate a data asset list, the method comprises the steps of extracting a foreign key association relationship between data tables to construct an initial entity relationship table, calculating reference frequencies between the data tables to generate an entity association degree matrix, dividing a service domain by adopting a K-means clustering algorithm, counting data table fields in the service domain to generate a data table field mapping table, recording a data table operation sequence to construct a data table operation link diagram, and establishing a data table operation link diagram. The method comprises the following steps: generating a data table node evaluation matrix, dividing sub-domains and determining a core entity, and constructing and training a multilayer neural network model for constructing business affiliation, data table incidence relation and business interface definition of a new data table, so as to solve the defects in the aspects of data asset relation discovery, business modeling automation, service interface generation and the like in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of modeling methods, and in particular, relates to a business modeling method, medium and system based on an enterprise data asset management platform. Background Art

[0002] With the continuous advancement of informatization, the data assets within enterprises are characterized by continuous expansion in scale, increasing variety, and increasingly complex source channels. How to effectively manage and utilize these massive data assets has become the key to improving the core competitiveness of enterprises. Existing data asset management technologies mainly include data warehouses, data lakes, knowledge graphs, etc. These technologies can help enterprises collect, store and analyze data assets to a certain extent. However, the following problems still exist in actual applications:

[0003] First, the existing data asset management technology focuses more on the attributes and characteristics of the data itself, and lacks in-depth analysis of the internal data operation mechanism of the enterprise, such as data sources, data flows, and data processing processes. This makes it difficult for enterprises to truly understand the actual application scenarios and roles of data assets in the business, and it is difficult to fully explore the value of data assets. Secondly, the existing technology has limitations in the relationship discovery and business modeling of data assets. In most cases, professional analysts are required to manually judge and model the associations between data tables based on experience. This method is inefficient and easily affected by subjective factors. The lack of reliable automated modeling methods has brought obstacles to the value mining of enterprise data assets. In addition, the existing technology has shortcomings in converting data assets into executable business logic and service interfaces. Even if the relationship model of data assets can be established, it is difficult to automatically generate service definitions that meet business needs, which increases the cost and cycle of IT construction. In summary, the existing data asset management technology still has great limitations in supporting the digital transformation of enterprises. It is urgent to develop a full-process business modeling method based on enterprise data assets to improve the efficiency and quality of enterprise information construction. Summary of the invention

[0004] In view of this, the present invention provides a business modeling method, medium and system based on an enterprise data asset management platform to solve the problem that the prior art lacks in-depth analysis of the internal data operation mechanism and is difficult to fully explore the value of data assets.

[0005] The present invention is achieved in that:

[0006] The first aspect of the present invention provides a business modeling method based on an enterprise data asset management platform, which includes the following steps: reading the enterprise data asset directory to generate a data asset list, extracting the foreign key association relationship between data tables to construct an initial entity relationship table, calculating the reference frequency between data tables to generate an entity association matrix, using the K-means clustering algorithm to divide the business domain, counting the data table fields in the business domain to generate a data table field mapping table, recording the data table operation sequence to construct a data table operation link diagram, using the data table node evaluation equation group to generate a data table node evaluation matrix, dividing the subdomain and determining the core entity, and constructing and training a multi-layer neural network model for constructing the business affiliation, data table association relationship, and business interface definition of the new data table.

[0007] The specific description of the data table node evaluation equation group is as follows:

[0008] The data table node evaluation equation group includes a data table data inflow equation, a data table data outflow equation, a data table node centrality equation, and a data table time series association equation;

[0009] The data table data inflow equation is used to calculate the importance of the data table as a data receiver, and the input includes the number of data flow operations entering the data table recorded in the data table operation link diagram, the amount of data received per unit time recorded in the data table operation link diagram, and the data update frequency recorded in the data table operation link diagram. The output is the data table data inflow score, and the data table data inflow score is used to generate the data table node evaluation matrix;

[0010] The data table data outflow equation is used to calculate the importance of the data table as a data provider, and the input includes the number of data outputs from the data table recorded in the data table operation link diagram, the data output volume per unit time recorded in the data table operation link diagram, and the data reading frequency recorded in the data table operation link diagram. The output is the data table data outflow score, and the data table data outflow score is used to generate the data table node evaluation matrix;

[0011] The data table node centrality equation is used to calculate the position importance of the data table in the entire data table operation link diagram, the input includes the number of other data tables directly connected to the data table in the initial entity relationship table, the number of data tables indirectly associated in the initial entity relationship table, and the duration of each association in the data table operation link diagram, and the output is the data table position importance score, which is used to generate the data table node evaluation matrix;

[0012] The data table timing association equation is used to evaluate the timing dependency of the data table in the business processing process. The input includes the number of business processes in which the data table participates in the data table operation link diagram, the processing sequence number in each business process in the data table operation link diagram, and the execution frequency of each business process in the data table operation link diagram. The output is the data table timing criticality score, which is used to generate the data table node evaluation matrix.

[0013] The multi-layer neural network model adopts an improved hybrid deep learning structure, which specifically includes a neural network data encoding layer, a neural network feature extraction layer, a neural network pyramid processing layer, a neural network feature fusion layer, a neural network pattern recognition layer, a neural network business modeling layer, and a neural network output mapping layer;

[0014] The neural network data encoding layer processes input data using a one-hot encoding method, converting the data table operation sequence feature vector, the data table node importance vector, and the data table field mapping vector into standard numerical features;

[0015] The neural network feature extraction layer includes parallel long short-term memory network units and convolutional neural network units, wherein the long short-term memory network units are used to extract temporal features, and the convolutional neural network units are used to extract spatial features;

[0016] The neural network pyramid processing layer includes a neural network first layer pyramid structure, a neural network second layer pyramid structure, and a neural network third layer pyramid structure, wherein:

[0017] The first layer pyramid structure of the neural network is used to identify the basic business features of the data table, the input includes the temporal features and spatial features of the neural network feature extraction layer, and the output is the basic business type recognition result of the data table;

[0018] The second-layer pyramid structure of the neural network is used to identify the business association pattern between data tables, and the input includes the output result of the first-layer pyramid structure of the neural network and the importance characteristics of the data table nodes, and the output is the association relationship identification result between the data tables;

[0019] The third-layer pyramid structure of the neural network is used to generate a data service definition, the input includes the output results of the first-layer pyramid structure of the neural network and the second-layer pyramid structure of the neural network, and the output is a service interface definition of a data table;

[0020] The neural network feature fusion layer uses an attention mechanism to perform weighted fusion on the three outputs of the neural network pyramid processing layer to generate a comprehensive feature vector;

[0021] The neural network pattern recognition layer adopts a multi-layer perceptron structure to perform deep feature learning on the comprehensive feature vector and identify the business modeling pattern;

[0022] The neural network business modeling layer adopts a graph neural network structure to map the pattern recognition results of the neural network pattern recognition layer to a predefined business domain model template;

[0023] The neural network output mapping layer adopts a fully connected network structure to generate the final business model definition, including the business type label of the data table, the data table relationship type label, and the data table service interface specification;

[0024] The specific structure of the multi-layer neural network model finally obtained is: the input data is standardized by the neural network data encoding layer, the temporal features and spatial features are obtained by the neural network feature extraction layer, multi-level feature recognition is performed by the neural network pyramid processing layer, multi-dimensional features are integrated by the neural network feature fusion layer, business patterns are learned by the neural network pattern recognition layer, a model framework is constructed in the neural network business modeling layer, and finally a standardized business model definition is generated by the neural network output mapping layer.

[0025] Among them, the data table node evaluation equation group includes a data table data inflow equation, a data table data outflow equation, a data table node centrality equation, and a data table time series association equation.

[0026] Furthermore, the data table data inflow equation is used to calculate the importance of the data table as a data receiver, and the input includes the number of data flow operations entering the data table, the amount of data received per unit time, and the data update frequency, and the output is the data table data inflow score; the data table data outflow equation is used to calculate the importance of the data table as a data provider, and the input includes the number of data outputs from the data table, the amount of data output per unit time, and the data reading frequency, and the output is the data table data outflow score.

[0027] Furthermore, the data table node centrality equation is used to calculate the positional importance of the data table in the entire data table operation link diagram. The input includes the number of other data tables directly connected to the data table, the number of indirectly associated data tables, and the duration of each association. The output is the data table positional importance score; the data table timing association equation is used to evaluate the timing dependence of the data table in the business processing process. The input includes the number of business processes in which the data table participates, the processing sequence number in each business process, and the execution frequency of each business process. The output is the data table timing criticality score.

[0028] Furthermore, the multi-layer neural network model adopts an improved hybrid deep learning structure, including a neural network data encoding layer, a neural network feature extraction layer, a neural network pyramid processing layer, a neural network feature fusion layer, a neural network pattern recognition layer, a neural network business modeling layer, and a neural network output mapping layer.

[0029] Furthermore, the neural network data encoding layer uses one-hot encoding to process input data; the neural network feature extraction layer includes parallel long short-term memory network units and convolutional neural network units, wherein the long short-term memory network units are used to extract temporal features, and the convolutional neural network units are used to extract spatial features.

[0030] Furthermore, the neural network pyramid processing layer includes a first-layer pyramid structure for identifying basic business features of data tables, a second-layer pyramid structure for identifying business association patterns between data tables, and a third-layer pyramid structure for generating service interface definitions for data tables.

[0031] Furthermore, the neural network feature fusion layer adopts an attention mechanism to perform weighted fusion on the output of the neural network pyramid processing layer; the neural network pattern recognition layer adopts a multi-layer perceptron structure; the neural network business modeling layer adopts a graph neural network structure; and the neural network output mapping layer adopts a fully connected network structure.

[0032] A second aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores program instructions, and when the program instructions are run in a computer, they are used to execute the above-mentioned business modeling method based on an enterprise data asset management platform.

[0033] A third aspect of the present invention provides a business modeling system based on an enterprise data asset management platform including the above-mentioned computer-readable storage medium.

[0034] Compared with the prior art, the business modeling method based on the enterprise data asset management platform provided by the present invention has the following beneficial effects:

[0035] This paper proposes a business modeling method based on an enterprise data asset management platform, aiming to solve the deficiencies of existing technologies in data asset relationship discovery, business modeling automation, service interface generation, etc. This method makes full use of the existing data asset information within the enterprise, and constructs a multi-layer neural network model through a series of steps such as entity association analysis, business domain division, and node importance evaluation. This model can learn enterprise business processes and data semantics, and provide automated support for business modeling of new data tables;

[0036] Specifically, the main technical effects of the method of the present invention are embodied in the following aspects:

[0037] 1. Achieved comprehensive analysis and modeling of enterprise data assets. This method not only focuses on the attributes of the data table itself, but also focuses on analyzing the relationship between data tables, data flow and business processing flow, so as to more comprehensively understand the application scenarios and functions of data assets within the enterprise;

[0038] 2. Provides an automated business modeling method based on machine learning. This method uses clustering algorithms to identify business domains, uses node evaluation equations to analyze the importance of data tables, and constructs a multi-layer neural network model for business learning and modeling. It greatly improves the efficiency and accuracy of business modeling and reduces the need for manual intervention;

[0039] 3. It can automatically generate data service interface definitions that meet business requirements. Based on business modeling, this method further extracts information such as business types, associations, and service interface specifications of data tables to provide support for the rapid integration of new data tables;

[0040] 4. Improved the overall efficiency of enterprise information construction. Through the above-mentioned automated data asset analysis and business modeling capabilities, the business modeling and service definition cycle of newly added data tables has been greatly shortened, while improving the overall data asset utilization efficiency, injecting new momentum into the digital transformation of enterprises;

[0041] In general, the method of the present invention fully taps the value of internal data assets of enterprises, uses advanced machine learning algorithms to realize the automation of business modeling, and injects new vitality into the informatization construction of enterprises. This can not only improve the informatization level of enterprises themselves, but also provide a useful practical sample for the academic community in the fields of knowledge graph construction and enterprise informatization construction. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0043] Figure 1 The flowchart is a business modeling method based on the enterprise data asset management platform. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solution and advantages of the embodiments of the present invention more clear, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0045] The specific implementation mode of the present invention is as follows Figure 1As shown, the following steps are included:

[0046] S01. Read the enterprise data asset directory, obtain the data table name, data table field name, data table field type, data table data source, data table data format, and generate a data asset list;

[0047] S02, traversing each data table in the data asset list, extracting foreign key associations between data tables, and constructing an initial entity relationship table;

[0048] S03, calculating the reference frequency between the data tables, counting the number of foreign key references between the two data tables according to the foreign key association relationship, and generating an entity association matrix;

[0049] The purpose of this step is to calculate the correlation index between data tables based on the foreign key reference relationship between them, so as to provide a basis for the subsequent business domain division. First, traverse each data table relationship recorded in the initial entity relationship table and count the number of foreign key references between the two data tables. Then, based on factors such as the number of foreign key references, the strength of the bidirectional association between data tables, the degree of synchronization between data tables, and the field similarity, calculate the correlation index between data tables. The specific calculation formula is as follows:

[0050]

[0051] Among them, R ij Indicates the degree of association between the i-th data table and the j-th data table. is the rate of change of foreign key reference frequency, t is the time variable, λ is the time attenuation coefficient, and its value range is [0.1, 0.5]. is the second-order partial derivative of the bidirectional correlation strength between data tables, x and y represent the characteristic dimensions of the two data tables respectively. is the rate of change of synchronization between data tables, T is the size of the observation time window, and K ij is the field similarity between the two data tables, K max is the maximum field similarity between all table pairs, α, β, and γ are the weight coefficients of the corresponding factors, satisfying α+β+γ = 1. All the calculated data table association indicators are organized into an entity association matrix, which reflects the relationship between the data tables within the enterprise.

[0052] S04, clustering the entity association matrix, and using the K-means clustering algorithm to divide the data tables in the entity association matrix whose association values ​​are greater than a preset threshold into the same business domain; the purpose of this step is to use the clustering algorithm to divide highly associated data tables into the same business domain, providing a basis for subsequent business modeling. First, cluster the entity association matrix generated in step S03, using the K-means clustering algorithm. The distance calculation formula of the K-means clustering algorithm is as follows:

[0053]

[0054] Among them, D ij is the distance between the i-th data table and the j-th data table, R ik is the element in the i-th row and k-th column of the entity association matrix, and n is the total number of data tables. In the clustering process, data tables with association values ​​greater than a preset threshold (e.g. 0.7) are classified into the same business domain. The threshold can be adjusted according to specific circumstances. The final business domains contain highly related data tables, which provide a basis for subsequent business modeling.

[0055] S05. Count the data table field names of each data table in the business domain, extract data table field groups with the same naming prefix, and generate a data table field mapping table;

[0056] S06. Record the add, delete, modify and query operation sequence of each data table in the business domain, including the data table operation timestamp, data table operation type and data table data flow direction, and construct a data table operation link diagram;

[0057] S07, using a data table node evaluation equation group to calculate the importance of each data table node in the data table operation link diagram, and generating a data table node evaluation matrix;

[0058] The purpose of this step is to calculate the importance of each data table node based on the role and status of the data table in the entire business process, providing a basis for subsequent business domain segmentation. Specifically, the following four equations are used to calculate the importance of data table nodes:

[0059] Data table data flows into the equation:

[0060]

[0061] Among them, S in is the data inflow score, N in is the number of data inflow operations, N max is the maximum number of operations, is the rate of change of data receiving rate, τ is the integral variable, V in is the amount of data received per unit time, Vmax is the maximum amount of data received, θ 1 is the nonlinear measurement index of the impact of data volume, F in is the data update frequency, F max is the maximum update frequency, ω 1 ,ω 2 ,ω 3 is the weight coefficient, satisfying ω 1 +ω 2 +ω 3 =1,ε 1 is the error term.

[0062] Data table data flow equation:

[0063]

[0064] Among them, S out is the data outflow score, N out is the number of data output operations, β is the time attenuation coefficient, V out is the data output per unit time, is the acceleration of the output frequency, θ 2 is the nonlinear influence index of data output, F out is the data reading frequency, λ 1 , 2 , 3 is the weight coefficient, satisfying λ 1 +λ 2 +λ 3 =1,ε 2 is the error term.

[0065] Data table node centrality equation:

[0066]

[0067] Among them, C is the position importance score, D direct is the number of directly connected data tables, D indirect is the number of indirectly related data tables, D max is the maximum number of associated data tables, is the spatial second-order partial derivative of the node centrality, is the time rate of change of correlation, T duration is the duration of the association, T max is the maximum duration, ω is the periodic change frequency, μ 1 , μ 2 , μ 3 is the weight coefficient, satisfying μ 1 +μ 2 +μ 3 =1,ε3 is the error term.

[0068] Data table timing correlation equation:

[0069]

[0070] Where T is the timing criticality score, P num is the number of participating business processes, P max is the maximum number of business processes, is the acceleration of the number of business processes, O avg is the average processing sequence number, O max is the maximum processing sequence number, is the rate of change of the processing order, F process is the business process execution frequency, F max is the maximum execution frequency, i is the imaginary unit, α is the time decay coefficient, η 1 , η 2 , η 3 is the weight coefficient, satisfying η 1 +η 2 +η 3 =1,ε 4 is the error term.

[0071] The calculation results of the above four equations are integrated into a data table node evaluation matrix, which records the importance of each data table node in the business process. This data table node evaluation matrix will provide a key basis for subsequent business domain segmentation.

[0072] S08, using the data table node evaluation matrix to divide the business domain into sub-domains, and determining the data table with the highest importance score as the core entity of the sub-domain;

[0073] S09, converting the data table operation link diagram, the data table node evaluation matrix, and the data table field mapping table into a training data set, wherein each piece of training data includes a data table operation sequence feature vector, a data table node importance vector, and a data table field mapping vector;

[0074] The purpose of this step is to integrate the various data obtained in the previous steps into a unified training data set to provide a basis for subsequent deep learning model training. First, extract the operation sequence information of each data table from the data table operation link diagram in step S06 and construct the operation sequence feature vector X op =[op 1 ,op 2 ,...,op n ] T , where op iThen, the importance index of each data table is extracted from the data table node evaluation matrix in step S07 to construct the node importance vector X imp =[S in ,S out ,C,T] T Finally, the field mapping information of each data table is extracted from the data table field mapping table in step S05 to construct the field mapping vector X map =[m 1 ,m 2 ,...,m k ] T , where m i Indicates the feature code of the i-th field mapping. Integrate the above three feature vectors into a complete training data sample as the input of the neural network model. Organize the data table information in all business domains into a unified training data set to provide a basis for subsequent deep learning model training.

[0075] S10, constructing a multi-layer neural network model, wherein the multi-layer neural network model comprises a neural network input layer, a neural network hidden layer, and a neural network output layer, wherein the number of neurons in the neural network input layer is the same as the feature dimension of the training data, and the number of neurons in the neural network output layer is the same as the number of business object types;

[0076] The purpose of this step is to build a multi-layer neural network model to learn the business relationships and semantic associations between internal data tables of the enterprise, and to provide support for subsequent business modeling. The number of neurons in the neural network input layer is the same as the feature dimension of the training data in step S09, and the received operation sequence feature vector X op , node importance vector X imp and the field mapping vector X map The hidden layer of the neural network adopts an improved hybrid deep learning structure, including:

[0077] Data encoding layer: Use one-hot encoding to standardize the input data;

[0078] Feature extraction layer: contains parallel long short-term memory networks and convolutional neural networks to extract temporal features and spatial features respectively;

[0079] Pyramid processing layer: Through the three-layer pyramid structure, the basic business characteristics of the data table, the association mode between data tables, and the data service interface definition are identified respectively;

[0080] Feature fusion layer: The attention mechanism is used to perform weighted fusion on the outputs of the above three layers;

[0081] Pattern recognition layer: uses a multi-layer perceptron structure to perform deep feature learning and identify business modeling patterns;

[0082] Business modeling layer: uses a graph neural network structure to map pattern recognition results into predefined business domain model templates;

[0083] Output mapping layer: uses a fully connected network structure to generate the final business model definition.

[0084] The number of neurons in the output layer of the neural network is the same as the number of business object types (such as customers, orders, products, etc.), which is used to generate the final business model definition.

[0085] S11. Use the training data set to train the multi-layer neural network model, and use the back propagation algorithm to update the parameters of the multi-layer neural network model until the multi-layer neural network model converges or reaches a preset training round. The training is completed to obtain a business domain model for constructing the business attribution of the new data table, the data table association relationship, and the business interface definition, specifically including identifying the business type of the data table, determining the association relationship of the data table, and generating a data service interface.

[0086] The specific implementation methods of the above steps are described in detail below:

[0087] Step S01: Read the enterprise data asset directory, obtain the data table name, data table field name, data table field type, data table data source, data table data format, and generate a data asset list.

[0088] First, the purpose of this step is to extract the basic information of all data tables within the enterprise, laying the foundation for subsequent business modeling. The specific implementation is as follows:

[0089] Read the data asset catalog from the enterprise's data asset management platform, which contains all the data table information within the enterprise.

[0090] Traverse each data table in the data asset directory and obtain metadata information such as the data table name, field name, field type, data source, and data format.

[0091] Organize the data table information obtained above into a unified data asset list, including data table name, field name and type, data source and data format, etc.

[0092] This data asset list provides the necessary basic data for subsequent steps such as entity relationship identification and business domain division.

[0093] Step S02: traverse each data table in the data asset list, extract the foreign key association relationship between the data tables, and construct an initial entity relationship table.

[0094] The purpose of this step is to construct a preliminary entity relationship table by analyzing the foreign key association relationship between data tables. The specific implementation is as follows:

[0095] Traverse each data table in the data asset list, check the field information of the table, and identify the fields that serve as foreign keys.

[0096] For each foreign key field, find its associated target data table and record the association relationship between the two data tables.

[0097] All the data table associations identified above are integrated into an initial entity relationship table, which contains data table name, association type (one-to-one, one-to-many, etc.), association fields, etc.

[0098] This initial entity relationship table provides basic data for subsequent entity association analysis, business domain division and other steps.

[0099] Step S03: Calculate the reference frequency between the data tables, count the number of foreign key references between the two data tables according to the foreign key association relationship, and generate an entity association matrix.

[0100] The purpose of this step is to calculate the correlation index between data tables based on the foreign key reference relationship between them, so as to provide a basis for the subsequent business domain division. The specific implementation is as follows:

[0101] Traverse each data table association relationship recorded in the initial entity relationship table and count the number of foreign key references between the two data tables.

[0102] The correlation index between data tables is calculated based on factors such as the number of foreign key references, the strength of the two-way association between data tables, the degree of synchronization between data tables, and the similarity of fields. The specific formula is as follows:

[0103]

[0104] Among them, R ij It represents the association degree between the i-th data table and the j-th data table. α, β, and γ are the weight coefficients of the corresponding factors, satisfying α+β+γ=1.

[0105] All calculated data table association indicators are organized into an entity association matrix, which reflects the relationship between the data tables within the enterprise.

[0106] Step S04: performing clustering calculation on the entity association matrix, and using a K-means clustering algorithm to divide data tables in the entity association matrix whose association values ​​are greater than a preset threshold into the same business domain.

[0107] The purpose of this step is to use clustering algorithms to divide highly related data tables into the same business domain, providing a basis for subsequent business modeling. The specific implementation is as follows:

[0108] The entity association matrix generated in step S03 is clustered using a K-means clustering algorithm.

[0109] The distance calculation formula of the K-means clustering algorithm is as follows:

[0110]

[0111] Among them, D ij is the distance between the i-th data table and the j-th data table, R ik is the element in the i-th row and k-th column in the entity association matrix, and n is the total number of data tables.

[0112] In the clustering process, data tables with correlation values ​​greater than a preset threshold (eg, 0.7) are grouped into the same business domain. The threshold can be adjusted according to specific circumstances.

[0113] The final business domains contain highly related data tables, which provide a basis for subsequent business modeling.

[0114] Step S05: Count the data table field names of each data table in the business domain, extract data table field groups with the same naming prefix, and generate a data table field mapping table.

[0115] The purpose of this step is to analyze the field information of each data table in the business domain, extract field groups with the same naming prefix, and generate a data table field mapping table to provide a reference for subsequent business modeling.

[0116] The specific implementation is as follows:

[0117] Traverse each business domain obtained in step S04 and perform field analysis on the data table in each business domain.

[0118] For each data table in the business domain, extract all its field name information.

[0119] According to the naming rules of field names, identify the field groups with the same prefix. For example, fields such as "Customer Table_Customer Number" and "Order Table_Customer Number" belong to the same field group.

[0120] All identified field groups are organized into a data table field mapping table, which records the field mapping relationship between data tables in each business domain.

[0121] The data table field mapping table provides valuable information for subsequent business modeling and helps to identify the semantic associations between data tables.

[0122] Step S06: Record the add, delete, modify and query operation sequences of each data table in the business domain, including data table operation timestamp, data table operation type and data table data flow, and construct a data table operation link diagram.

[0123] The purpose of this step is to analyze the operation sequence information of each data table in the business domain, build a data table operation link diagram, and provide a basis for subsequent node importance evaluation. The specific implementation is as follows:

[0124] Each business domain obtained in step S04 is traversed, and an operation sequence analysis is performed on the data table in each business domain.

[0125] Record the add, delete, modify and query operation information of each data table, including the operation timestamp, operation type (add, delete, modify, query) and data flow direction.

[0126] Based on the above recorded data table operation information, a data table operation link diagram is constructed, which reflects the data flow and operation timing relationship between the data tables in the business domain.

[0127] The data table operation link diagram provides necessary input data for subsequent node importance assessment and helps identify business-critical data tables.

[0128] Step S07: using the data table node evaluation equation group to calculate the importance of each data table node in the data table operation link diagram, and generating a data table node evaluation matrix.

[0129] The purpose of this step is to calculate the importance of each data table node based on the role and status of the data table in the entire business process, and provide a basis for subsequent business domain segmentation. The specific implementation is as follows:

[0130] The importance of the data table nodes is calculated using the following four equations:

[0131] Data table data flows into the equation:

[0132]

[0133] Data table data flow equation:

[0134]

[0135] Data table node centrality equation:

[0136]

[0137] Data table timing correlation equation:

[0138]

[0139] The calculation results of the above four equations are integrated into a data table node evaluation matrix, which records the importance of each data table node in the business process.

[0140] This data table node evaluation matrix will provide key basis for subsequent business domain segmentation.

[0141] Step S08: using the data table node evaluation matrix to divide the business domain into sub-domains, and determining the data table with the highest importance score as the core entity of the sub-domain.

[0142] The purpose of this step is to further subdivide the business domain, identify the core data tables of each subdomain, and provide a basis for business modeling. The specific implementation is as follows:

[0143] Based on the data table node evaluation matrix generated in step S07, the previously divided business domains are further subdivided.

[0144] The data table nodes in each business domain are sorted according to their importance scores in the evaluation matrix.

[0145] The data tables with the highest importance scores are identified as the core entities of the subdomain.

[0146] Centered on the core entity, other data tables with high correlation are divided into the same subdomain.

[0147] Through the above method, the original business domain is further subdivided into smaller subdomains, and each subdomain has a clear core data table.

[0148] These subdomains and their core entities provide refined basic data for subsequent business modeling.

[0149] Step S09: converting the data table operation link diagram, the data table node evaluation matrix, and the data table field mapping table into a training data set, wherein each piece of training data contains a data table operation sequence feature vector, a data table node importance vector, and a data table field mapping vector.

[0150] The purpose of this step is to integrate the various data obtained in the previous steps into a unified training data set to provide a basis for subsequent deep learning model training. The specific implementation is as follows:

[0151] Extract the operation sequence information of each data table from the data table operation link diagram in step S06, and construct the operation sequence feature vector X op .

[0152] Extract the importance index of each data table from the data table node evaluation matrix in step S07, and construct the node importance vector X imp .

[0153] Extract the field mapping information of each data table from the data table field mapping table in step S05, and construct the field mapping vector X map .

[0154] The above three feature vectors are integrated into a complete training data sample as the input of the neural network model.

[0155] Organize the data table information in all business domains into a unified training data set to provide a basis for subsequent deep learning model training.

[0156] Step S10: Construct a multi-layer neural network model, which includes a neural network input layer, a neural network hidden layer, and a neural network output layer. The number of neurons in the neural network input layer is the same as the feature dimension of the training data, and the number of neurons in the neural network output layer is the same as the number of business object types.

[0157] The purpose of this step is to build a multi-layer neural network model to learn the business relationships and semantic associations between internal data tables of the enterprise, and provide support for subsequent business modeling. The specific implementation is as follows:

[0158] Neural network input layer: The number of neurons in the input layer is the same as the feature dimension of the training data in step S09, and the received operation sequence feature vector X op , node importance vector X imp and the field mapping vector X map .

[0159] Neural network hidden layer: The hidden layer adopts an improved hybrid deep learning structure, including:

[0160] Data encoding layer: Use one-hot encoding to standardize the input data;

[0161] Feature extraction layer: contains parallel LSTM networks and convolutional networks to extract temporal features and spatial features respectively;

[0162] Pyramid processing layer: Through the three-layer pyramid structure, the basic business characteristics of the data table, the association mode between data tables, and the data service interface definition are identified respectively;

[0163] Feature fusion layer: The attention mechanism is used to perform weighted fusion on the outputs of the above three layers;

[0164] Pattern recognition layer: uses a multi-layer perceptron structure to perform deep feature learning and identify business modeling patterns;

[0165] Neural network output layer: The number of neurons in the output layer is the same as the number of business object types (such as customers, orders, products, etc.), which is used to generate the final business model definition.

[0166] Step S11: Use the training data set to train the multi-layer neural network model, and use the back propagation algorithm to update the multi-layer neural network model parameters until the multi-layer neural network model converges or reaches a preset training round. The training is completed to obtain a business domain model for constructing the business attribution of the new data table, data table association relationship, and business interface definition.

[0167] The purpose of this step is to use the multi-layer neural network model constructed above to learn and model the business relationships and data semantics within the enterprise, and provide support for the business attribution, association relationship and service interface definition of the new data table. The specific implementation is as follows:

[0168] The multi-layer neural network model is trained using the training data set constructed in step S09.

[0169] The back propagation algorithm is used to continuously adjust the parameters of each layer of the neural network to minimize the model loss function. The training process continues until the model converges or reaches the preset training rounds.

[0170] After the training is completed, a fully learned business domain model is obtained. Based on the input data table information, the model can automatically identify the business type of the data table, determine the relationship between data tables, and generate a standardized data service interface definition.

[0171] By using this business domain model, the business affiliation and relationship of newly added data tables can be quickly determined, and corresponding service interface specifications can be generated, greatly improving the efficiency of enterprise information construction.

[0172] In order to better understand and implement the present invention, an embodiment of a specific application scenario of the present invention is provided below: A large pharmaceutical company is undergoing digital transformation and upgrading, and its goal is to fully explore its huge internal data assets and improve business operation efficiency. The company has data assets covering multiple business areas such as R&D, production, marketing, and finance, with a total of 80 interrelated data tables. In order to realize the automated modeling of business processes, the company decided to adopt the business modeling method based on the data asset management platform proposed in the present invention.

[0173] The first step is to generate a data asset list. Through the data asset management platform, the company obtained the names, fields, data types, sources, and formats of all data tables, and compiled them into a complete data asset list. The list contains 80 data tables, covering multiple business areas such as clinical trials, drug production, raw material procurement, sales management, and financial accounting. For example, in the field of clinical trials, there are "clinical trial plan table", "clinical trial record table", "subject information table", etc.; in the field of drug production, there are "production plan table", "production process record table", "product qualification certificate table", etc.; in the field of raw material procurement, there are "raw material purchase order table", "raw material storage table", "supplier information table", etc.; in the field of sales management, there are "sales order table", "sales invoice table", "customer information table", etc.; in the field of financial accounting, there are "accounts receivable table", "accounts payable table", "general ledger detail table", etc. Through this detailed data asset list, a good foundation has been laid for subsequent entity association analysis and business modeling.

[0174] Next is entity association analysis. Based on the field information in the data asset list, the company identified the foreign key associations between the data tables. For example, the "Subject Number" field in the "Clinical Trial Record Table" is associated with the "Subject Number" primary key of the "Subject Information Table"; the "Product Number" field in the "Production Process Record Table" is associated with the "Product Number" primary key of the "Product Qualification Certificate Table"; the "Customer Number" field in the "Sales Order Table" is associated with the "Customer Number" primary key of the "Customer Information Table", and so on. By sorting out these foreign key associations, the company constructed an initial entity association table, recording the association types and association fields between the data tables.

[0175] On this basis, the company calculated the reference frequency between each data table, and counted the mutual reference times between two data tables based on the foreign key relationship, thus generating an entity association matrix. The specific calculation formula is as follows:

[0176]

[0177] Among them, R ij Indicates the correlation between the i-th data table and the j-th data table, is the rate of change of foreign key reference frequency, t is the time variable (unit: day), K ij is the field similarity between the two data tables (value range 0-1), is the second-order partial derivative of the bidirectional correlation strength between data tables, is the rate of change of synchronization between data tables, T is the observation time window (30 days), T ijis the temporal correlation between two data tables (value range 0-1). After calculation, the enterprise obtained an 80*80 entity correlation matrix, which reflects the relationship between the data tables.

[0178] Based on the above entity association matrix, the company uses the K-means clustering algorithm to divide the data table into business domains. The cluster distance calculation formula is as follows:

[0179]

[0180] Among them, D ij is the distance between the i-th data table and the j-th data table, R ik is the element in the i-th row and k-th column of the entity association matrix. After clustering, the company divided the 80 data tables into five business domains, namely:

[0181] The clinical trial management business domain includes 9 tables, including "Clinical Trial Plan Table", "Clinical Trial Record Table", "Subject Information Table" and so on; the manufacturing business domain includes 12 tables, including "Production Plan Table", "Production Process Record Table", "Product Qualification Certificate Table" and so on; the raw material procurement business domain includes 7 tables, including "Raw Material Purchase Order Table", "Raw Material Receipt Table", "Supplier Information Table" and so on; the sales management business domain includes 11 tables, including "Sales Order Table", "Sales Invoice Table", "Customer Information Table" and so on; the financial accounting business domain includes 9 tables, including "Accounts Receivable Table", "Accounts Payable Table", "General Ledger Detail Table" and so on.

[0182] Next, the company further analyzed the data table fields in each business domain and identified field groups with the same naming prefix. For example, in the clinical trial management domain, the "Protocol Number" field of the "Clinical Trial Protocol Table", the "Protocol Number" field of the "Clinical Trial Record Table", and the "Protocol Number" field of the "Subject Information Table" all belong to the "Protocol Number" field group; in the manufacturing domain, the "Product Number" field of the "Production Plan Table", the "Product Number" field of the "Production Process Record Table", and the "Product Number" field of the "Product Qualification Certificate Table" all belong to the "Product Number" field group, and so on. Through this field analysis, the company constructed a data table field mapping table to record the field correspondence between the data tables in each business domain.

[0183] After completing the business domain division and field mapping, the company further analyzed the role and status of each data table in the business process. Specifically, the company used the following four equations to calculate the importance of data table nodes:

[0184] Data table data flows into the equation:

[0185]

[0186] Among them, S in is the data inflow score, N in is the number of data inflow operations (value range 0-120), is the rate of change of data receiving rate, V in is the amount of data received per unit time (value range 0-2100), F in The data update frequency (range: 0-0.8 times / day).

[0187] Data table data flow equation:

[0188]

[0189] Among them, S out is the data outflow score, N out is the number of data output operations (value range 0-90), V out is the data output per unit time (value range 0-1800), is the acceleration of the output frequency, F out is the data reading frequency (range 0-0.7 times / day).

[0190] Data table node centrality equation:

[0191]

[0192] Among them, C is the position importance score, D direct is the number of directly connected data tables (value range 0-15), D indirect is the number of indirectly associated data tables (value range 0-35), is the spatial second-order partial derivative of the node centrality, is the time rate of change of correlation, T duration The duration of the association (range: 0-90 days).

[0193] Data table timing correlation equation:

[0194]

[0195] Where T is the timing criticality score, P num is the number of business processes involved (value range 0-8), is the acceleration of the number of business processes, O avg is the average processing sequence number (value range 1-5), is the rate of change of the processing order, F process The frequency of business process execution (range: 0-0.6 times / day).

[0196] Based on the calculation results of the above four sets of equations, the company generated an 80*4 data table node evaluation matrix, recording the importance of each data table node in the business process. For example, in the clinical trial management domain, the scores of "Clinical Trial Plan Table" are (0.82, 0.75, 0.78, 0.71), the scores of "Clinical Trial Record Table" are (0.91, 0.68, 0.85, 0.75), and the scores of "Subject Information Table" are (0.78, 0.63, 0.72, 0.66); in the manufacturing domain, the scores of "Production Plan Table" are (0.85, 0.82, 0.88, 0.79), the scores of "Production Process Record Table" are (0.92, 0.88, 0.92, 0.83), and the scores of "Product Qualification Certificate Table" are (0.86, 0.79, 0.84, 0.77); in the raw material procurement domain, the scores of "Raw Material Purchase Order Table" are (0.75, 0.71, 0.73, 0.69), and "Raw Material Storage Table" are (0.86, 0.79, 0.84, 0.77). The scores of "Sales Order Table" are (0.88, 0.84, 0.87, 0.82), "Sales Invoice Table" are (0.91, 0.86, 0.90, 0.84), "Customer Information Table" are (0.68, 0.62, 0.65, 0.61); in the sales management domain, the scores of "Sales Order Table" are (0.88, 0.84, 0.87, 0.82), "Sales Invoice Table" are (0.91, 0.86, 0.90, 0.84), "Customer Information Table" are (0.68, 0.62, 0.65, 0.61). The scores of "Account Information Table" are (0.79, 0.72, 0.76, 0.68); in the financial accounting domain, the scores of "Accounts Receivable Table" are (0.83, 0.78, 0.81, 0.75), the scores of "Accounts Payable Table" are (0.85, 0.81, 0.83, 0.77), and the scores of "General Ledger Detail Table" are (0.79, 0.72, 0.76, 0.69).

[0197] With the above data table node evaluation matrix, the company further subdivided each business domain, identified the data table with the highest importance score as the core entity of the subdomain, and then divided other data tables with higher relevance into the same subdomain. For example, in the clinical trial management domain, the "clinical trial record form" is determined as the core entity, and the "clinical trial plan form" and "subject information form" with a high degree of correlation are divided into the same subdomain; in the manufacturing domain, the "production process record form" is determined as the core entity, and the "production plan form" and "product qualification certificate form" with a high degree of correlation are divided into the same subdomain; in the raw material procurement domain, the "raw material purchase order form" is determined as the core entity, and the "raw material storage form" and "supplier information form" with a high degree of correlation are divided into the same subdomain; in the sales management domain, the "sales order form" is determined as the core entity, and the "sales invoice form" and "customer information form" with a high degree of correlation are divided into the same subdomain; in the financial accounting domain, the "accounts receivable form" is determined as the core entity, and the "accounts payable form" and "general ledger details form" with a high degree of correlation are divided into the same subdomain.

[0198] After completing the above data asset analysis and business domain division, the company began to build a multi-layer neural network model to learn the business relationships and semantic connections within the company. Specifically, the model includes:

[0199] 1. Input layer: receiving data table operation sequence feature vector X op , node importance vector X imp and the field mapping vector X map Among them, X op Contains information such as the type of add, delete, modify, and query operations, timestamp, and data flow direction of each data table; X imp Contains the four scores of each data table in the node evaluation matrix; X map Contains the field group code to which each data table field belongs.

[0200] 2. Data encoding layer: Use one-hot encoding to standardize the input data.

[0201] 3. Feature extraction layer: Contains parallel long short-term memory networks and convolutional neural networks, which extract the temporal features and spatial features of the input data respectively.

[0202] 4. Pyramid processing layer:

[0203] The first layer: Identify the basic business characteristics of the data table, such as clinical trials, manufacturing, sales management, etc.

[0204] The second layer: Identify the association patterns between data tables, such as one-to-one, one-to-many and other relationship types.

[0205] The third layer: Generate the service interface definition of the data table, including input parameters, output parameters, etc.

[0206] 5. Feature fusion layer: The attention mechanism is used to perform weighted fusion on the outputs of the above three layers to generate a comprehensive business feature vector.

[0207] 6. Pattern recognition layer: Using a multi-layer perceptron structure, deep feature learning is performed on business feature vectors to identify potential business modeling patterns.

[0208] 7. Business modeling layer: Using graph neural network structure, the pattern recognition results are mapped to predefined business domain model templates to generate specific business model definitions.

[0209] 8. Output mapping layer: uses a fully connected network structure to convert the business model definition into an executable data service interface specification.

[0210] In terms of training data preparation, the company extracted key features from the aforementioned data table operation link diagram, node evaluation matrix, and field mapping table to construct training samples. Specifically, it includes:

[0211] Operation sequence feature vector X op : Extract the add, delete, modify and query operation sequence of each data table from the data table operation link diagram and encode it into a feature vector. For example, the operation sequence of the "clinical trial record table" can be expressed as X op =[2,1,3,2,1,3,2] T , where 2 represents an insert operation, 1 represents a delete operation, and 3 represents an update operation.

[0212] Node importance vector X imp : Extract the scores of each data table in four dimensions from the node evaluation matrix to form an importance vector. For example, the node importance vector of "Clinical Trial Record Table" is X imp =[0.91,0.68,0.85,0.75] T .

[0213] Field mapping vector X map : Extract the field group code of each data table field from the field mapping table to form a field mapping vector. For example, the field mapping vector of the "Clinical Trial Record Table" is X map =[1,2,3,4,5,6,7] T , where 1 corresponds to the “Protocol Number” field group, 2 corresponds to the “Subject Number” field group, and so on.

[0214] Using the above three types of feature vectors as input, the company trained an 8-layer deep neural network model. The training process uses the back propagation algorithm to continuously adjust the parameters of each layer to minimize the loss function until the model converges or reaches the preset 1,000 training rounds. After the training is completed, the company obtains a fully learned business domain model that can automatically identify the business type of the data table, determine the relationship between the data tables, and generate a standardized data service interface definition based on the input data table information.

[0215] For example, when the company needs to create a new "drug batch quality record table", it only needs to input the metadata information of the table (table name, field name, field type, etc.) into the trained model. The model can automatically recognize that the table belongs to the manufacturing business domain and is closely related to the "production process record table" and "product qualification certificate table", and generate the following business model definition:

[0216] Business type: Manufacturing; Relationship: One-to-many relationship with "Production Process Record Table", the associated field is "Product Number", one-to-one relationship with "Product Qualification Certificate Table", the associated field is "Product Number"; Service interface: Input parameters: Product number, batch number, test items, test results; Output parameters: Product number, batch number, qualified status;

[0217] With the above business model definition, the company's IT department can quickly complete the database design and service interface development of the "drug batch quality record table", greatly improving the integration efficiency of new data assets. At the same time, the company can also automatically adjust data services according to changes in the business model to ensure rapid response to business needs.

[0218] Table 1 Variables of the present invention and their explanations

[0219]

[0220] Specifically, the principle of the present invention is:

[0221] The business modeling method proposed in this invention is based on the enterprise internal data asset management platform, which mainly includes the following key steps: data asset list generation, entity association analysis, business domain division, node importance evaluation, training data set construction and neural network model training. These steps together constitute a complete data-driven business modeling process, which contains rich technical principles;

[0222] First, in the data asset list generation stage, this method analyzes the enterprise data asset catalog and extracts the basic information of each data table, laying the foundation for subsequent association analysis and business modeling. This step mainly uses the technical means of data extraction and organization;

[0223] Secondly, in the entity association analysis stage, this method calculates the association index between data tables by counting the foreign key reference relationship between them. Here, a multi-factor-based association calculation formula is adopted, which comprehensively considers factors such as foreign key reference frequency, data table feature similarity, and data synchronization degree, and more comprehensively reflects the internal connection between data tables. This association calculation method originates from the edge weight calculation principle in complex network theory;

[0224] Next, in the business domain division stage, this method uses the K-means clustering algorithm to divide highly related data tables into the same business domain. The clustering distance calculation formula based on the correlation matrix is ​​used here to effectively identify the business boundaries within the enterprise. The purpose of business domain division is to lay the foundation for subsequent fine-grained business modeling. This idea originates from the domain division technology in the construction of knowledge graphs;

[0225] In the node importance evaluation stage, this method constructs a set of data table node evaluation equations, which comprehensively evaluates the importance of each data table node in the entire business process from multiple dimensions such as data inflow, data outflow, node centrality, and time series correlation. This node evaluation method based on the topological structure of the business process and data operation characteristics is derived from the research results in the field of complex network analysis and business process mining;

[0226] Finally, in the neural network model training phase, the method constructs a multi-level neural network architecture, including data encoding layer, feature extraction layer, pyramid processing layer, feature fusion layer, pattern recognition layer, business modeling layer, and output mapping layer. This hybrid deep learning structure can fully learn multi-dimensional features such as data table operation sequence, node importance, and field mapping, and finally output a model definition that meets business needs. This technical solution is based on the latest research results in the fields of knowledge graph representation learning and end-to-end business modeling.

Claims

1. A business modeling method based on an enterprise data asset management platform, characterized in that: The following steps are involved: Read the enterprise data asset catalog to generate a data asset list, extract the foreign key relationship between data tables to build an initial entity relationship table, calculate the reference frequency between data tables to generate an entity association matrix, use the K-means clustering algorithm to divide the business domain, count the data table fields in the business domain to generate a data table field mapping table, record the data table operation sequence to build a data table operation link diagram, use the data table node evaluation equation group to generate a data table node evaluation matrix, divide the subdomain and determine the core entity, build and train a multi-layer neural network model to construct the business affiliation, data table association relationship, and business interface definition of the new data table.

2. A business modeling method based on an enterprise data asset management platform according to claim 1, characterized in that: The data table node evaluation equation group includes a data table data inflow equation, a data table data outflow equation, a data table node centrality equation, and a data table time series association equation.

3. A business modeling method based on an enterprise data asset management platform according to claim 2, characterized in that: The data table data inflow equation is used to calculate the importance of the data table as a data receiver. The input includes the number of data flow operations entering the data table, the amount of data received per unit time, and the data update frequency. The output is the data table data inflow score. The data table data outflow equation is used to calculate the importance of the data table as a data provider. The input includes the number of data outputs from the data table, the amount of data output per unit time, and the data reading frequency. The output is the data table data outflow score.

4. A business modeling method based on an enterprise data asset management platform according to claim 3, characterized in that: The data table node centrality equation is used to calculate the position importance of the data table in the entire data table operation link diagram. The input includes the number of other data tables directly connected to the data table, the number of indirectly associated data tables, and the duration of each association. The output is the data table position importance score. The data table timing association equation is used to evaluate the timing dependency of the data table in the business processing process. The input includes the number of business processes in which the data table participates, the processing sequence number in each business process, and the execution frequency of each business process. The output is the data table timing criticality score.

5. A business modeling method based on an enterprise data asset management platform according to claim 4, characterized in that: The multi-layer neural network model adopts an improved hybrid deep learning structure, including a neural network data encoding layer, a neural network feature extraction layer, a neural network pyramid processing layer, a neural network feature fusion layer, a neural network pattern recognition layer, a neural network business modeling layer, and a neural network output mapping layer.

6. A business modeling method based on an enterprise data asset management platform according to claim 5, characterized in that: The neural network data encoding layer uses one-hot encoding to process input data; the neural network feature extraction layer includes parallel long short-term memory network units and convolutional neural network units, wherein the long short-term memory network units are used to extract temporal features, and the convolutional neural network units are used to extract spatial features.

7. A business modeling method based on an enterprise data asset management platform according to claim 6, characterized in that: The neural network pyramid processing layer includes a first-layer pyramid structure for identifying basic business features of data tables, a second-layer pyramid structure for identifying business association patterns between data tables, and a third-layer pyramid structure for generating service interface definitions for data tables.

8. A business modeling method based on an enterprise data asset management platform according to claim 7, characterized in that: The neural network feature fusion layer adopts the attention mechanism to perform weighted fusion on the output of the neural network pyramid processing layer; the neural network pattern recognition layer adopts a multi-layer perceptron structure; the neural network business modeling layer adopts a graph neural network structure; and the neural network output mapping layer adopts a fully connected network structure.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, and when the program instructions are run in a computer, they are used to execute the business modeling method based on an enterprise data asset management platform as described in any one of claims 1 to 8.

10. A business modeling system based on an enterprise data asset management platform, characterized in that: A computer-readable storage medium comprising the computer-readable storage medium of claim 9.

Citation Information

Cited By

  • Associable table discovery method based on dynamic data, electronic equipment and storage medium

    CN120849479A

  • Methods for discovering associationable tables based on dynamic data, electronic devices and storage media

    CN120849479B

  • Enterprise data asset information management method based on knowledge graph

    CN121544398A