A business data graph generation method and system combined with data lake warehouse
By combining the business data map generation method of the data lake warehouse, the problems of weak readability, low semantic level and low update efficiency in the existing technology are solved, and efficient and intuitive business data map generation and update are achieved.
Patent Information
- Application Number
- CN202510138189.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-08
AI Technical Summary
In the prior art, business data has weak readability, low semantic level, and low efficiency in updating business scenarios.
By combining the business data map generation method of the data lake warehouse, hierarchical analysis of business data attributes is carried out in response to the request of the user, and a business data attribute tree is generated, and a cascading sub-network is extracted and configured, a cascading total network is generated, and training is performed to generate a graph mapping function, and finally a business data map is constructed.
Improve the readability and semantic hierarchy of data, allowing business personnel to understand and apply data more intuitively, and quickly update the map when business scenarios change, improving efficiency.
Smart Images

Figure CN119578518B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a method and system for generating a business data graph in combination with a data lake warehouse. Background Art
[0002] With the rapid development of information technology, enterprises have accumulated a massive amount of business data. Existing technologies usually adopt the method of centrally storing and managing data from different sources and formats. Among them, structured data is used for storage and analysis, while unstructured or semi-structured data is directly stored in its original form.
[0003] However, most business data exists in the form of underlying parameters or intermediate values and lacks business context, resulting in poor readability of business data and difficulty for business personnel to directly understand and apply it. Although data from different sources and formats are stored centrally, they lack semantic associations with each other, resulting in a low data semantic level and inability to fully reflect business logic and internal connections. There is also modeling and analysis based on different business data, but for new business scenarios, data analysts are required to reconfigure the calculation logic to generate corresponding high-level semantic data, resulting in inefficient updating of business scenarios and difficulty in flexibly adapting to changes. Summary of the invention
[0004] In order to solve the technical problems in the prior art such as weak readability of business data, low semantic level and low efficiency in updating business scenarios, the present invention provides a business data graph generation method and system combined with a data lake warehouse.
[0005] The technical solution of the present invention to solve the above technical problems is as follows:
[0006] In a first aspect, the present invention provides a method for generating a business data graph in combination with a data lake warehouse, including: in response to a business data graph generation request initiated by a user terminal, obtaining a business data attribute set for hierarchical analysis, and generating a business data attribute tree; extracting the N-level first group of leaf nodes and the N-level first parent node of the business data attribute tree until the N-level M-th group of leaf nodes and the N-level M-th parent node; configuring the N-level first cascade subnetwork based on the N-level first group of leaf nodes and the N-level first parent node, and configuring the N-level M-th cascade subnetwork based on the N-level M-th group of leaf nodes and the N-level M-th parent node; until the secondary leaf nodes and the first node of the business data attribute tree are extracted, and the secondary cascade subnetwork is configured; based on the business data attribute tree, connecting the N-level first cascade subnetwork until the N-level M-th cascade subnetwork, until the secondary cascade subnetwork is generated; based on the business data attribute tree, training the cascade total network to generate a graph mapping function; and constructing a business data graph based on the graph mapping function and the business data attribute tree and feeding it back to the user terminal.
[0007] In a second aspect, the present invention provides a business data graph generation system combined with a data lake warehouse, including: a data attribute analysis module, which is used to respond to a business data graph generation request initiated by a user terminal, obtain a business data attribute set for hierarchical analysis, and generate a business data attribute tree; a node hierarchical extraction module, which is used to extract the N-level first group of leaf nodes and the N-level first parent node of the business data attribute tree until the N-level M-th group of leaf nodes and the N-level M-th parent node; a cascade network configuration module, which is used to configure the N-level first cascade sub-network based on the N-level first group of leaf nodes and the N-level first parent node, based on the N-level M-th group of leaf nodes and the N-level The Mth parent node of the Nth level configures the Mth cascade subnetwork of the Nth level; the secondary relationship construction module is used to extract the secondary leaf nodes and the first-level nodes of the business data attribute tree, and configure the secondary cascade subnetwork; the network level connection module is used to connect the Nth first cascade subnetwork to the Mth cascade subnetwork of the Nth level, and to the second-level cascade subnetwork based on the business data attribute tree, to generate a cascade total network; the graph function training module is used to perform training on the cascade total network based on the data lake warehouse to generate a graph mapping function; the data graph construction module is used to build a business data graph based on the graph mapping function and the business data attribute tree and feedback it to the user end.
[0008] The beneficial effects of the present invention are:
[0009] In response to the business data graph generation request initiated by the user end, the business data attribute set is obtained for hierarchical analysis, and a business data attribute tree is generated, which realizes the hierarchical combing and structured processing of business data, which helps to improve the readability and semantic level of data. From the business data attribute tree, the first group of leaf nodes and the first parent node of level N are extracted from the top to the bottom, until the Mth group of leaf nodes and the Mth parent node of level N, and the first to the Mth cascade sub-networks of level N are configured based on the extracted nodes, until the second-level leaf nodes and the first-level nodes are extracted, and the second-level cascade sub-network is configured, and the business data attribute tree is mapped to a cascade network structure, which is conducive to subsequent network training and graph generation. Next, based on the business data attribute tree, the second-level cascade sub-network is connected to the Mth cascade sub-network of level N to generate a complete cascade total network, covering the complex associations between various levels and dimensions of business data, and laying a network foundation for graph generation. Based on the massive multi-source heterogeneous data of the data lake warehouse, the generated cascade total network is trained to obtain the graph mapping function, which characterizes the mapping rules from the attribute tree to the graph, so that the corresponding graph representation can be quickly generated according to the business data. According to the graph mapping function and the original business data attribute tree, a business data graph is constructed and fed back to the user end. The generated graph not only retains the hierarchical structure of the data, but also integrates rich semantic associations, allowing business personnel to intuitively understand and apply the data, while being traceable to the original attributes, improving interpretability. When the business scenario changes, you only need to update the business data attribute tree and re-execute the graph mapping to efficiently generate a new business data graph, meeting the flexible and changeable needs of the business. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 A schematic diagram of a process flow of a method for generating a business data graph in combination with a data lake warehouse provided by the present invention;
[0011] Figure 2 A structural schematic diagram of a business data graph generation system combined with a data lake warehouse provided by the present invention.
[0012] In the accompanying drawings, the components represented by the reference numerals are as follows:
[0013] Data attribute analysis module 11, node hierarchical extraction module 12, cascade network configuration module 13, secondary relationship construction module 14, network level connection module 15, graph function training module 16, data graph construction module 17. DETAILED DESCRIPTION
[0014] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0015] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0016] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or explanation". Any embodiment described as "for example" in the present invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any technician in the field to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in the present invention.
[0017] Embodiment 1:
[0018] like Figure 1 As shown, an embodiment of the present invention provides a method for generating a business data graph combined with a data lake warehouse, which is applied to a server.
[0019] Specifically, a method for generating a business data graph in combination with a data lake warehouse provided in an embodiment of the present invention is applied to a server side, wherein the server may be an independent physical server, a cloud server, or a server cluster.
[0020] The business data graph generation method includes:
[0021] S100: In response to a business data graph generation request initiated by a user terminal, a business data attribute set is obtained for hierarchical analysis to generate a business data attribute tree.
[0022] Specifically, when the user needs to generate a business data graph, it sends a business data graph generation request to the server. After receiving the request, the server first obtains a business data attribute set. The business data attribute set includes various data attributes related to the business data graph to be generated. The data attributes can be business parameters, business indicators, calculation rules, etc.
[0023] After obtaining the business data attribute set, the data attributes in the set are hierarchically analyzed. The purpose of hierarchical analysis is to identify and establish the hierarchical relationship and association relationship between various data attributes, so as to form a business data attribute tree with a clear hierarchical structure. This attribute tree can intuitively reflect the subordinate relationship and association degree between various business data attributes.
[0024] By converting the originally discrete business data attributes into a hierarchical tree structure, the foundation is laid for subsequent graph generation. This tree structure makes the semantic hierarchy of business data clearer and the relationship between data more intuitive.
[0025] S200: Extracting the first group of leaf nodes at level N and the first parent node at level N up to the Mth group of leaf nodes at level N and the Mth parent node at level N of the business data attribute tree.
[0026] Specifically, after obtaining the business data attribute tree, the business data attribute tree is analyzed hierarchically. For the Nth level, the leaf nodes of the level are grouped based on the structural relationship of the business data attribute tree itself, and each group contains several leaf nodes and their corresponding parent nodes. Among them, the leaf nodes refer to the nodes in the current level, and the parent nodes are the nodes of the previous level directly connected to these leaf nodes.
[0027] Then, the node information of the first group of leaf nodes and the first parent node of level N to the Mth group of leaf nodes and the Mth parent node of level N is obtained in sequence by group extraction. For each group, the leaf node set of the group and its corresponding parent node are included, thereby maintaining the hierarchical association relationship between the data, laying the foundation for the subsequent configuration of the cascade sub-network.
[0028] The above grouping extraction process fully complies with the inherent structure of the business data attribute tree, ensuring that the extraction results can accurately reflect the hierarchical relationship of the original data. In this way, the complex tree structure can be converted into a grouping form that is easy to process, providing a data foundation for subsequent network construction.
[0029] S300: Based on the N-level first group of leaf nodes and the N-level first parent node, configure the N-level first cascade sub-network; based on the N-level M-th group of leaf nodes and the N-level M-th parent node, configure the N-level M-th cascade sub-network.
[0030] Specifically, after obtaining the first group of leaf nodes of level N and the first parent node of level N until the Mth group of leaf nodes of level N and the Mth parent node of level N, a corresponding cascade subnetwork is configured for each group of nodes. For the first group of level N, the first cascade subnetwork of level N is configured using the leaf nodes and parent nodes contained therein; and so on, until the Mth cascade subnetwork of level N is configured using the leaf nodes and parent nodes of the Mth group of level N.
[0031] During the configuration process, the input of each cascade sub-network is the set of leaf nodes of the group, and the output is the corresponding parent node. This configuration method enables each cascade sub-network to learn and express the mapping relationship from leaf nodes to parent nodes, thereby capturing the intrinsic connection between data.
[0032] By configuring an independent cascade sub-network for each group of nodes, it is possible to achieve refined modeling of the relationships between different groups of data. This group modeling method not only improves the expressiveness of the model, but also provides a basis for subsequent network fusion.
[0033] S400: until the secondary leaf nodes and the primary nodes of the service data attribute tree are extracted, a secondary cascade sub-network is configured.
[0034] Specifically, after completing the cascade sub-network configuration of the N-level nodes, a bottom-up processing method is adopted to process the node configuration of the N-1 level, N-2 level, and so on in sequence, until the secondary node of the business data attribute tree is reached. The layer-by-layer upward processing method ensures the continuity and integrity of the hierarchical relationship and effectively guarantees the semantic transmission of business data.
[0035] When processing the second-level node, extract the second-level leaf node and its corresponding first-level node, and configure the second-level cascade sub-network based on these nodes. The first-level node, as the root node of the entire business data attribute tree, carries the highest-level business semantic information, while the second-level leaf node directly connected to it transmits the lower-level business attribute information. The second-level cascade sub-network realizes the conversion of business data from a lower semantic level to a higher semantic level by establishing a mapping relationship from the second-level leaf node to the first-level node. As the top-level network in the entire network structure, the second-level cascade sub-network takes the second-level leaf node set as input and the first-level node as output, ensuring that the final generated business data graph has a clear semantic hierarchical structure.
[0036] The completion of the configuration of the secondary cascade sub-network marks the completion of the construction of the entire cascade network structure. This bottom-up configuration method not only realizes the complete expression of the hierarchical relationships in the business data attribute tree, but also provides a reliable network foundation for generating high-quality business data graphs through layer-by-layer semantic enhancement.
[0037] S500: Based on the service data attribute tree, connect the N-level first cascade subnetworks to the N-level M-th cascade subnetwork, and to the second cascade subnetwork to generate a cascaded total network.
[0038] Specifically, after obtaining the N-level first cascade subnetwork to the N-level M-th cascade subnetwork, and the second-level cascade subnetwork, all configured cascade subnetworks are connected in order to build a complete cascade total network. The connection process is based on the structure of the business data attribute tree to ensure that the connection relationship between the networks corresponds to the hierarchical relationship of the business data.
[0039] During the connection process, the cascade subnetworks within the same level are first connected in parallel. For example, at the N-level level, the first cascade subnetwork of level N to the M-th cascade subnetwork of level N are fully connected in parallel to form the overall network structure of this level. Similarly, each cascade subnetwork of level N-1 is also fully connected in parallel, and so on, until the connection of the second-level cascade subnetwork is completed. The connection between each level follows the hierarchical transmission relationship, that is, the subnetwork in the k-level network is connected to the corresponding subnetwork in the k-1-level network, where the output of the k-1-level network serves as the input of the k-level network. This connection method enables business data to be transmitted and converted from bottom to top in the network.
[0040] The generation of the cascaded total network fully considers the hierarchical characteristics of business data. Through parallel connection and hierarchical transmission, the business data is mapped from the bottom layer to the top layer step by step, laying a structural foundation for subsequent network training and the generation of graph mapping functions.
[0041] S600: Based on the data lake warehouse, the cascaded total network is trained to generate a graph mapping function.
[0042] Specifically, after obtaining the cascaded total network, the generated cascaded total network is trained using the large-scale data resources stored in the data lake warehouse. As the source of training data, the data lake warehouse not only provides rich business data samples, but also ensures the diversity and representativeness of the training data.
[0043] During the training process, first, cascade training data is collected from the data lake warehouse according to the structure of the business data attribute tree. The collected training data needs to cover the node relationships at all levels in the business data attribute tree to ensure the comprehensiveness of the training. Then, a step-by-step training strategy is adopted to train each cascade sub-network separately to generate an initial set of graph mapping functions. This pre-training method enables each sub-network to better learn the local mapping relationship. After that, cascade training is performed on the initial set of graph mapping functions, and the final graph mapping function is obtained through overall optimization. This function can accurately map the input business data to the corresponding graph structure to achieve a visual expression of the data relationship.
[0044] By leveraging the data advantages of the data lake warehouse and the step-by-step training strategy, the accuracy and generalization ability of the graph mapping function are ensured, providing a reliable function foundation for the generation of business data graphs.
[0045] S700: Construct a business data graph based on the graph mapping function and the business data attribute tree and feed it back to the user end.
[0046] Specifically, after obtaining the graph mapping function, the obtained graph mapping function is applied to the business data attribute tree to automatically construct a business data graph. The graph mapping function can accurately map the hierarchical relationship and association relationship in the business data attribute tree to the graph structure, thereby forming an intuitive data visualization expression.
[0047] During the construction process, the nodes of the business data graph are composed of business data attributes, and the lines between the nodes represent the processing logic rules. In this way, the semantic relationship of the business data is clearly presented. Each node not only contains the original business data attribute information, but also reflects the semantic level of the data through its position in the graph.
[0048] After the construction is completed, the business data graph is fed back to the user who initiated the request. Users can intuitively understand the relationship between business data through the graph and quickly grasp the semantic content of the data. This visual expression method significantly improves the readability of business data and facilitates users to analyze data and make decisions.
[0049] The automatic conversion from the business data attribute tree to the business data graph is realized through the graph mapping function, so that the graph generation needs of users can be quickly responded to, and the efficiency of business scenario updates is improved. At the same time, the generated graph has a higher semantic level, providing users with more valuable data insights.
[0050] Take the business data analysis of e-commerce enterprises as an example. An e-commerce platform needs to analyze the association between multi-dimensional data such as product sales, user behavior, and logistics distribution. When the business analyst initiates a data graph business data graph generation request through the user end, first, obtain the business data attribute set, including product data (such as product ID, product category, brand, price, inventory, etc.), sales data (such as order ID, sales time, sales quantity, sales amount, etc.), user data (such as user ID, user level, purchase frequency, browsing history, etc.) and logistics data (such as logistics order number, delivery method, delivery time, logistics status, etc.). Subsequently, these attributes are hierarchically analyzed, with the product sales business as the first-level node, product information, transaction information, user information, and logistics information as the second-level nodes, and specific data attributes as the third-level nodes to generate a business data attribute tree. Taking the product information branch as an example, the specific parameters such as product ID, category, brand, etc. are used as N-level leaf nodes, and the basic attributes of the product are used as N-level parent nodes. This hierarchical structure clearly shows the subordinate relationship between product-related data.
[0051] In the node extraction phase, processing starts from the bottom layer. For the product information branch, extract the basic attributes of leaf nodes such as product ID, category, brand and their parent node products, and configure the corresponding cascade sub-network. The network takes specific product parameters as input and outputs product feature representation. In this way, other branches such as transaction information, user information, logistics information, etc. are processed in sequence until all levels of nodes are processed. When generating a cascaded total network, the sub-networks of the same level are first connected in parallel. For example, N-level networks such as product attributes and transaction records are fully connected. Then, according to the hierarchical relationship of the business data attribute tree, the network is connected upward level by level, and finally connected to the root node of the product sales business to form a complete cascaded total network structure. This network structure ensures that data features can be passed from the bottom layer to the top layer step by step. After that, the historical transaction data stored in the data lake warehouse is used for training. These data contain complete business records such as product information, transaction records, user behavior, logistics distribution, etc. First, pre-train each sub-network such as the product information network so that it can accurately extract the features of this level. The entire network is then trained end-to-end to obtain a graph mapping function that can accurately map business relationships. After that, a business data graph containing all business entities and relationships is constructed based on the trained graph mapping function. In this graph, products, orders, users, logistics, etc. are used as graph nodes, and the lines between nodes show the association between various attributes, such as the sales relationship between products and orders, and the distribution relationship between orders and logistics. At the same time, the graph shows each level from specific parameters to business decisions through a clear hierarchical structure.
[0052] Through this graph, business analysts can intuitively view the full-link data of product sales, analyze the relationship between product sales and user behavior, discover the impact of logistics and distribution on sales, and quickly respond to the data analysis needs of new business scenarios. When an enterprise needs to analyze other business scenarios, it only needs to update the corresponding business data attributes, and the system can quickly generate the corresponding business data graph, greatly improving the efficiency of data analysis and business decision-making. It can be seen from this example that the business data graph generation method provided in the embodiment of the present application can effectively improve the readability and semantic level of the data, enabling business personnel to better understand and apply the data, and can also respond quickly when the business scenario changes.
[0053] Furthermore, the embodiment of the present application also includes:
[0054] S110: extracting a first service data attribute to a Qth service data attribute according to the service data attribute set;
[0055] S120: Based on the first service data attribute, perform correlation analysis on the second service data attribute to the Qth service data attribute to obtain a first service data attribute correlation tree;
[0056] S130: Based on the second service data attribute, performing association analysis on the first service data attribute, the third service data attribute, and up to the Qth service data attribute to obtain a second service data attribute association tree;
[0057] S140: performing association analysis on the first business data attribute to the Q-1th business data attribute based on the Qth business data attribute to obtain a Qth business data attribute association tree;
[0058] S150: Merge the first service data attribute association tree, the second service data attribute association tree, and the Qth service data attribute association tree to generate the service data attribute tree.
[0059] Specifically, first, each business data attribute is extracted from the business data attribute set in sequence and marked as a first business data attribute, a second business data attribute, and finally a Qth business data attribute, where Q is the total number of business data attributes. This extraction method ensures complete processing of the business data attributes.
[0060] Then, taking the first business data attribute as a benchmark, a correlation analysis is performed on all other business data attributes (i.e., the second business data attribute to the Qth business data attribute), thereby constructing a first business data attribute association tree. The association tree reflects the association relationship between the first business data attribute and other attributes. Subsequently, taking the second business data attribute as a benchmark, a correlation analysis is performed on all other attributes including the first business data attribute to construct a second business data attribute association tree. This analysis method ensures that the association relationship between data attributes is captured from different perspectives. According to the same analysis method, each business data attribute is processed in turn until the first to Q-1th business data attributes are analyzed for correlation based on the Qth business data attribute to obtain the Qth business data attribute association tree. Afterwards, all the obtained business data attribute association trees are fused to generate the final business data attribute tree. This fusion process comprehensively considers each business data attribute association tree to ensure that the generated business data attribute tree can fully reflect the hierarchical structure of the business data.
[0061] Through the processing method based on correlation analysis, the relationship between business data attributes is fully explored and expressed, laying a solid data foundation for subsequent graph generation.
[0062] Furthermore, the embodiment of the present application also includes:
[0063] S121: Collecting a reference data sequence according to the first service data attribute;
[0064] S122: Collecting and comparing data sequences according to the second service data attribute to the Qth service data attribute;
[0065] S123: performing grey correlation analysis according to the reference data sequence and the comparison data sequence to obtain a second business data attribute correlation until a Qth business data attribute correlation;
[0066] S124: extracting business data attribute sets and association sets that are greater than or equal to the association threshold from the second business data attribute association to the Qth business data attribute association, and constructing the first business data attribute association tree.
[0067] In a preferred embodiment,
[0068] Taking the construction of the first business data attribute association tree as an example, the process of constructing each business data attribute association tree is explained. First, taking the first business data attribute as a benchmark, a data sequence related to the attribute is collected from the data source as a benchmark data sequence. The sequence reflects the data performance of the first business data attribute in different business scenarios. For example, if the first business data attribute is the sales volume of a product, the benchmark data sequence can be the sales data sequence of the product at different time points. Then, data sequences are collected for the second business data attribute to the Qth business data attribute respectively as comparison data sequences. These sequences have the same collection dimensions as the benchmark data sequence to ensure the comparability of subsequent analysis. Continuing with the above example, the comparison data sequence may include data sequences of related attributes such as product price, promotion intensity, and market competition index.
[0069] Subsequently, the grey correlation analysis method is used to calculate the degree of association between the benchmark data sequence and each comparison data sequence. For example, the benchmark data sequence and the comparison data sequence are dimensionless; the difference information sequence between the comparison sequence and the benchmark sequence is calculated; the correlation coefficient is solved; and the correlation degree of each comparison data sequence relative to the benchmark sequence is calculated. Through this analysis, the correlation value of the second business data attribute to the Qth business data attribute relative to the first business data attribute can be obtained. Through grey correlation analysis, not only can the situation of incomplete data and insufficient information be effectively handled, but also complex business scenarios with irregular data distribution can be handled.
[0070] After that, a relevance threshold is set, which can be adjusted according to specific business needs. Business data attributes with a relevance greater than or equal to the threshold and their corresponding relevance are extracted to form a business data attribute set and a relevance set, respectively. Based on these two sets, a first business data attribute relevance tree is constructed. In the relevance tree, the first business data attribute is located at the root node of the tree, and other attributes are organized according to the hierarchical relationship according to the size of the relevance, forming a relevance tree with a clear hierarchical structure.
[0071] Through the grey correlation analysis method, the correlation between business data attributes can be accurately identified, and attribute combinations with significant correlation can be screened out. This method based on quantitative analysis not only improves the accuracy of correlation identification, but also provides reliable data support for the subsequent construction of business data attribute trees. At the same time, by setting the correlation threshold, the scale of the correlation tree can be effectively controlled to ensure that the generated correlation tree can reflect important business relationships without becoming too complicated due to the introduction of too many weak correlations.
[0072] Furthermore, the embodiment of the present application also includes:
[0073] S151: extracting the i-th business data attribute association tree and the j-th business data attribute association tree from the first business data attribute association tree, the second business data attribute association tree, and up to the Q-th business data attribute association tree;
[0074] The i-th business data attribute association tree includes an i-th association attribute set and an i-th association degree set, and the j-th business data attribute association tree includes a j-th association attribute set and a j-th association degree set;
[0075] S152: Configure association tree fusion rules:
[0076] When the j-th business data attribute belongs to the i-th associated attribute set, and the i-th business data attribute does not belong to the j-th associated attribute set, the i-th business data attribute is set as a first-level node, the non-intersecting associated attributes of the i-th associated attribute set and the j-th associated attribute set are set as second-level nodes, the j-th associated attribute set is set as a third-level node, and the third-level node is a child node of the j-th business data attribute, and the j-th business data attribute belongs to the non-intersecting associated attributes;
[0077] When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute does not belong to the i-th associated attribute set, the j-th business data attribute is set as a first-level node, the non-intersecting associated attributes of the j-th associated attribute set and the i-th associated attribute set are set as second-level nodes, the i-th associated attribute set is set as a third-level node, and the third-level node is a child node of the i-th business data attribute, and the i-th business data attribute belongs to the non-intersecting associated attributes;
[0078] When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is equal to the number of second associated attributes of the j-th associated attribute set, and when the j-th business data attribute association degree of the i-th association degree set is greater than the i-th business data attribute association degree of the j-th association degree set, the i-th business data attribute is set as a first-level node, the non-intersection associated attributes of the i-th associated attribute set and the j-th associated attribute set are set as second-level nodes, and the j-th associated attribute set is set as a third-level node;
[0079] When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is equal to the number of second associated attributes of the j-th associated attribute set, and when the j-th business data attribute association degree of the i-th association degree set is less than the i-th business data attribute association degree of the j-th association degree set, set the j-th business data attribute as a first-level node, set the non-intersection associated attributes of the j-th associated attribute set and the i-th associated attribute set as a second-level node, and set the i-th associated attribute set as a third-level node;
[0080] When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is greater than the number of second associated attributes of the j-th associated attribute set, the i-th business data attribute is set as a first-level node, the non-intersection associated attributes of the i-th associated attribute set and the j-th associated attribute set are set as second-level nodes, and the j-th associated attribute set is set as a third-level node;
[0081] When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is less than the number of second associated attributes of the j-th associated attribute set, the j-th business data attribute is set as a first-level node, the non-intersection associated attributes of the j-th associated attribute set and the i-th associated attribute set are set as second-level nodes, and the i-th associated attribute set is set as a third-level node;
[0082] When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is equal to the number of second associated attributes of the j-th associated attribute set, and when the j-th business data attribute association degree of the i-th association degree set is equal to the i-th business data attribute association degree of the j-th association degree set, the i-th business data attribute is set as the first-level first node, the j-th business data attribute is set as the first-level second node, and the intersection association attribute of the i-th associated attribute set and the j-th associated attribute set is set as the common second-level node of the first-level first node and the second-level second node;
[0083] S153: According to the association tree fusion rule, fuse the first business data attribute association tree, the second business data attribute association tree until the Qth business data attribute association tree to generate the business data attribute tree.
[0084] In a preferred embodiment, after obtaining the first business data attribute association tree, the second business data attribute association tree, and the Qth business data attribute association tree, first, any two association trees are selected from all generated business data attribute association trees (including the first business data attribute association tree to the Qth business data attribute association tree), and are respectively recorded as the i-th business data attribute association tree and the j-th business data attribute association tree. The i-th business data attribute association tree includes the i-th association attribute set (recording attributes related to the i-th business data attribute) and the i-th association degree set (recording corresponding association degree values), and the j-th business data attribute association tree includes the j-th association attribute set (recording attributes related to the j-th business data attribute) and the j-th association degree set (recording corresponding association degree values).
[0085] Next, configure detailed association tree fusion rules, which specify the hierarchical layout of the fused nodes based on the relationship characteristics between the two association trees.
[0086] The first case is a one-way association case. Specifically, when the j-th business data attribute belongs to the i-th association attribute set, but the i-th business data attribute does not belong to the j-th association attribute set, it means that there is a one-way association relationship. At this time, the i-th business data attribute is set as a first-level node; the non-intersection parts of the two association attribute sets are set as second-level nodes; the j-th association attribute set is set as a third-level node, and these nodes are all child nodes of the j-th business data attribute; the j-th business data attribute is classified as a non-intersection association attribute. When the opposite one-way association situation occurs (that is, the i-th business data attribute belongs to the j-th association attribute set, but the j-th business data attribute does not belong to the i-th association attribute set), the corresponding symmetric processing method is adopted.
[0087] The second case is a two-way association with an equal number of associated attributes. Specifically, when the i-th business data attribute and the j-th business data attribute belong to each other's associated attribute set, and the number of attributes in the two associated attribute sets is equal, the hierarchical relationship is determined by comparing the degree of association. When the degree of association of the j-th business data attribute in the i-th association set is greater than the degree of association of the i-th business data attribute in the j-th association set, the i-th business data attribute is set as a first-level node; the non-intersection part of the two associated attribute sets is set as a second-level node; and the j-th associated attribute set is set as a third-level node. When an opposite degree of association occurs, the corresponding symmetrical processing method is adopted.
[0088] The third case is a bidirectional association with unequal numbers of associated attributes. Specifically, when two business data attributes are associated with each other, but the sizes of associated attribute sets are different, if the number of attributes in the i-th associated attribute set is greater than that in the j-th associated attribute set, the i-th business data attribute is set as a first-level node; the non-intersection part of the two associated attribute sets is set as a second-level node; and the j-th associated attribute set is set as a third-level node. If the number of attributes in the i-th associated attribute set is less than that in the j-th associated attribute set, the corresponding symmetric processing method is adopted.
[0089] The fourth case is a two-way association and a completely equal situation. Specifically, when the following conditions are met at the same time, a special equal processing method is adopted: the i-th business data attribute and the j-th business data attribute belong to each other's associated attribute set; the number of attributes in the two associated attribute sets is equal; the correlation between the two attributes is also equal (that is, the correlation of the j-th business data attribute in the i-th correlation set is equal to the correlation of the i-th business data attribute in the j-th correlation set). In this case, the i-th business data attribute is set as the first-level first node; the j-th business data attribute is set as the second-level second node; the intersection of the two associated attribute sets is set as the shared second-level node of the two first-level nodes. This processing method reflects the equal relationship between the two business data attributes.
[0090] After that, all business data attribute association trees are fused according to the association tree fusion rules configured above. The specific fusion process adopts an iterative method. First, any two association trees are fused, and the fusion result is fused with the next association tree. This process is repeated until all association trees are processed, thereby generating a complete business data attribute tree that accurately reflects the hierarchical relationship and association relationship between each business data attribute.
[0091] Furthermore, the embodiment of the present application also includes:
[0092] S510: The N-level first cascade subnetworks to the N-level M-th cascade subnetworks are fully connected in parallel to obtain a one-layer network, and the N-1-level first cascade subnetworks to the N-1-th L-th cascade subnetworks are fully connected in parallel to obtain a two-layer network, until the two-level cascade subnetworks are used as the N-1-layer network to construct the cascaded total network;
[0093] Among them, any sub-network of the k-layer network has a sub-network in the k-1-layer network, and the output of the sub-network in the k-1-layer network is the input of the sub-network in the k-layer network.
[0094] In a preferred embodiment, after obtaining the N-level first cascade subnetwork to the N-level M-th cascade subnetwork, and then to the second-level cascade subnetwork, a hierarchical connection method is used to construct a cascaded total network. First, all cascaded subnetworks at the bottom layer (N level) (the N-level first cascade subnetwork to the N-level M-th cascade subnetwork) are fully connected in parallel. Specifically, the N-level first cascade subnetwork to the N-level M-th cascade subnetwork are fully connected in parallel to form the first layer of the cascaded total network.
[0095] Then, all cascade subnetworks of the penultimate layer (N-1 level) are fully connected in parallel. Specifically, the first cascade subnetwork of level N-1 to the Lth cascade subnetwork of level N-1 are fully connected in parallel to form the second layer of the cascaded total network. In this way, the process is carried out layer by layer until the second-level cascade subnetwork is used as the N-1th layer network, and finally a complete cascaded total network structure is constructed.
[0096] Among them, for the k-th layer network (k>1) in the cascaded total network, any sub-network in it has a corresponding sub-network in the k-1 layer network, and the output of the corresponding sub-network in the k-1 layer network will be used as the input of the sub-network in the k-1 layer network. This inter-layer connection method ensures that data can be effectively transmitted between network layers.
[0097] Through the parallel fully connected hierarchical network construction method, not only the hierarchical structure of the original business data attribute tree is maintained, but also the layer-by-layer mapping and conversion of data is realized through the connection relationship between networks. The cascaded total network constructed in this way provides a complete network foundation for subsequent network training and the generation of graph mapping functions.
[0098] Furthermore, the embodiment of the present application also includes:
[0099] S610: Collect cascade training data according to the data lake warehouse and the business data attribute tree;
[0100] S620: According to the cascade training data, the N-level first cascade subnetwork to the N-level M-th cascade subnetwork to the second cascade subnetwork are trained separately to generate an initial graph mapping function set;
[0101] S630: Perform cascade training on the initial atlas mapping function set according to the cascade training data to generate the atlas mapping function.
[0102] In a feasible implementation, after obtaining the cascaded total network, the cascaded total network is trained to generate a final graph mapping function.
[0103] First, the data required for training is collected from the data lake warehouse. The collection process is based on the business data attribute tree to ensure that the collected data can cover the data features of each level and each node in the attribute tree. The rich data resources in the data lake warehouse provide sufficient data support for training. The collected cascade training data includes both the original business data and the relationship between the data.
[0104] Subsequently, a step-by-step training strategy is adopted to first train each sub-network in the cascaded total network independently. Specifically, starting from the bottom layer, the N-level first cascaded sub-network to the N-level M-th cascaded sub-network, and finally to the second-level cascaded sub-network, are trained one by one. The training of each sub-network is based on the part of the cascaded training data collected above that is related to the network. Through independent training, each sub-network can better learn the local mapping relationship and form an initial set of graph mapping functions.
[0105] After that, based on the initial graph mapping function set obtained in the previous steps, the entire network is cascaded trained using the cascade training data. This process will take into account the association between each sub-network, and through overall optimization, the network can better express the overall mapping relationship of the business data, and finally generate a complete graph mapping function.
[0106] Through the local-first-global training strategy, each sub-network is first trained separately to fully learn local features, and then the overall mapping relationship is optimized through cascade training. This training method not only improves the training efficiency, but also ensures that the final generated graph mapping function has good performance.
[0107] Furthermore, the embodiment of the present application also includes:
[0108] S631: Construct a first cascade training loss function, wherein the first cascade training loss function is used to count the proportion of sub-networks whose loss values are greater than or equal to a loss threshold in the N-level first cascade sub-networks to the N-level M-th cascade sub-networks to the second cascade sub-networks;
[0109] S632: Construct a second cascade training loss function, wherein the second cascade training loss function is used to count the mean loss values in the N-level first cascade subnetwork to the N-level M-th cascade subnetwork to the second cascade subnetwork;
[0110] S633: If, in at least 95% of consecutive preset number of trainings, the first cascade training loss function is less than or equal to the first loss threshold, and the second cascade training loss function is less than or equal to the second loss threshold, output the graph mapping function.
[0111] In a preferred embodiment, when the initial graph mapping function set is cascaded, first, a first cascade training loss function is constructed to monitor the number of sub-networks with poor performance in the network. The loss function counts the proportion of sub-networks whose loss values exceed or are equal to the loss threshold from the first cascade sub-network of level N to the Mth cascade sub-network of level N, until the second cascade sub-network, so that the weak links in the network can be discovered and paid attention to in time, providing a basis for subsequent optimization.
[0112] At the same time, the second cascade training loss function is constructed, which focuses on evaluating the overall performance of the entire network. By calculating the mean loss value of all sub-networks, a global view of the network training status is obtained. The mean statistics are not only simple and intuitive, but also effectively reflect the overall training level of the network.
[0113] Then, set strict training completion conditions. Continuously monitor the training process for a preset number of consecutive times, requiring that in at least 95% of the training times, the value of the first cascade training loss function is less than or equal to the first loss threshold, and the value of the second cascade training loss function is less than or equal to the second loss threshold. When these conditions are met, the final graph mapping function can be output.
[0114] By adopting a dual loss function, the first cascade training loss function ensures the local quality of the network and avoids too many sub-networks with poor performance; the second cascade training loss function ensures the overall performance level of the network. The synergy of the two loss functions, the setting of a 95% qualified ratio requirement and a preset number of consecutive times, jointly ensure that the final output graph mapping function has stable and reliable performance, providing reliable functional support for the generation of business data graphs.
[0115] The method for generating a business data graph in combination with a data lake warehouse provided by an embodiment of the present invention has at least the following technical effects:
[0116] In response to the business data graph generation request initiated by the user end, the business data attribute set is obtained for hierarchical analysis, and a business data attribute tree is generated, which realizes the hierarchical combing and structured processing of business data, lays the foundation for subsequent steps, and helps to improve data readability and semantic hierarchy. Extract the first group of leaf nodes and the first parent node of the N-level business data attribute tree until the M-th group of leaf nodes and the M-th parent node of the N-level. Based on the first group of leaf nodes and the first parent node of the N-level, configure the first cascade sub-network of the N-level. Based on the M-th group of leaf nodes and the M-th parent node of the N-level, configure the M-th cascade sub-network of the N-level, until the second-level leaf nodes and the first-level nodes of the business data attribute tree are extracted, and the second-level cascade sub-network is configured, and the business data attribute tree is mapped into a multi-level cascade network structure, which depicts the association between nodes at different levels and prepares for graph generation. Based on the business data attribute tree, connect the first cascade sub-network of the N-level until the M-th cascade sub-network of the N-level, until the second-level cascade sub-network, and generate a cascade total network, covering the association between each level and dimension of the business data, and laying a network foundation for graph generation. Based on the data lake warehouse, the cascaded total network is trained to generate a graph mapping function, which describes the mapping rules from the attribute tree to the graph, so that the graph can be quickly generated according to the business data. According to the graph mapping function and the business data attribute tree, the business data graph is constructed and fed back to the user end. The generated graph not only retains the hierarchical structure of the data, but also integrates rich semantic information, allowing business personnel to intuitively understand and apply the data. When the business scenario changes, the graph can be updated efficiently to flexibly adapt to changes in demand.
[0117] Embodiment 2:
[0118] like Figure 2 As shown, based on the same inventive concept as the method for generating a business data graph in combination with a data lake warehouse provided in Embodiment 1, an embodiment of the present invention further provides a system for generating a business data graph in combination with a data lake warehouse, including:
[0119] The data attribute analysis module 11 is used to respond to the business data graph generation request initiated by the user end, obtain the business data attribute set for hierarchical analysis, and generate a business data attribute tree;
[0120] A node hierarchical extraction module 12, used for extracting the first group of leaf nodes at level N and the first parent node at level N up to the Mth group of leaf nodes at level N and the Mth parent node at level N of the business data attribute tree;
[0121] The cascade network configuration module 13 is used to configure the first N-level cascade sub-network based on the first N-level leaf node group and the first N-level parent node, and to configure the M-level N-level cascade sub-network based on the M-level leaf node group and the M-level parent node;
[0122] A secondary relationship construction module 14 is used to extract the secondary leaf nodes and the primary nodes of the business data attribute tree and configure a secondary cascade sub-network;
[0123] A network level connection module 15, configured to connect the N-level first cascade subnetworks to the N-level M-th cascade subnetworks to the second cascade subnetworks based on the service data attribute tree, to generate a cascaded total network;
[0124] A graph function training module 16, used to perform training on the cascaded total network based on the data lake warehouse to generate a graph mapping function;
[0125] The data graph construction module 17 is used to construct a business data graph based on the graph mapping function and the business data attribute tree and feed it back to the user end.
[0126] Furthermore, the data attribute analysis module 11 includes the following execution steps:
[0127] Extracting, according to the service data attribute set, a first service data attribute to a Qth service data attribute;
[0128] Based on the first business data attribute, performing association analysis on the second business data attribute up to the Qth business data attribute to obtain a first business data attribute association tree;
[0129] Based on the second business data attribute, performing association analysis on the first business data attribute, the third business data attribute, and up to the Qth business data attribute to obtain a second business data attribute association tree;
[0130] Until, based on the Qth business data attribute, correlation analysis is performed on the first business data attribute to the Q-1th business data attribute to obtain a Qth business data attribute correlation tree;
[0131] The first service data attribute association tree, the second service data attribute association tree, and finally the Qth service data attribute association tree are integrated to generate the service data attribute tree.
[0132] Furthermore, the data attribute analysis module 11 also includes the following execution steps:
[0133] According to the first service data attribute, collecting a reference data sequence;
[0134] According to the second service data attribute to the Qth service data attribute, collecting and comparing data sequences;
[0135] Performing grey correlation analysis on the reference data sequence and the comparison data sequence to obtain a second business data attribute correlation degree up to a Qth business data attribute correlation degree;
[0136] A business data attribute set and a correlation degree set that are greater than or equal to a correlation degree threshold value are extracted from the second business data attribute correlation degree to the Qth business data attribute correlation degree, and the first business data attribute correlation tree is constructed.
[0137] Furthermore, the data attribute analysis module 11 also includes the following execution steps:
[0138] Extracting the i-th business data attribute association tree and the j-th business data attribute association tree from the first business data attribute association tree, the second business data attribute association tree, and up to the Q-th business data attribute association tree;
[0139] The i-th business data attribute association tree includes an i-th association attribute set and an i-th association degree set, and the j-th business data attribute association tree includes a j-th association attribute set and a j-th association degree set;
[0140] Configure association tree fusion rules:
[0141] When the j-th business data attribute belongs to the i-th associated attribute set, and the i-th business data attribute does not belong to the j-th associated attribute set, the i-th business data attribute is set as a first-level node, the non-intersecting associated attributes of the i-th associated attribute set and the j-th associated attribute set are set as second-level nodes, the j-th associated attribute set is set as a third-level node, and the third-level node is a child node of the j-th business data attribute, and the j-th business data attribute belongs to the non-intersecting associated attributes;
[0142] When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute does not belong to the i-th associated attribute set, the j-th business data attribute is set as a first-level node, the non-intersecting associated attributes of the j-th associated attribute set and the i-th associated attribute set are set as second-level nodes, the i-th associated attribute set is set as a third-level node, and the third-level node is a child node of the i-th business data attribute, and the i-th business data attribute belongs to the non-intersecting associated attributes;
[0143] When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is equal to the number of second associated attributes of the j-th associated attribute set, and when the j-th business data attribute association degree of the i-th association degree set is greater than the i-th business data attribute association degree of the j-th association degree set, the i-th business data attribute is set as a first-level node, the non-intersection associated attributes of the i-th associated attribute set and the j-th associated attribute set are set as second-level nodes, and the j-th associated attribute set is set as a third-level node;
[0144] When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is equal to the number of second associated attributes of the j-th associated attribute set, and when the j-th business data attribute association degree of the i-th association degree set is less than the i-th business data attribute association degree of the j-th association degree set, set the j-th business data attribute as a first-level node, set the non-intersection associated attributes of the j-th associated attribute set and the i-th associated attribute set as a second-level node, and set the i-th associated attribute set as a third-level node;
[0145] When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is greater than the number of second associated attributes of the j-th associated attribute set, the i-th business data attribute is set as a first-level node, the non-intersection associated attributes of the i-th associated attribute set and the j-th associated attribute set are set as second-level nodes, and the j-th associated attribute set is set as a third-level node;
[0146] When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is less than the number of second associated attributes of the j-th associated attribute set, the j-th business data attribute is set as a first-level node, the non-intersection associated attributes of the j-th associated attribute set and the i-th associated attribute set are set as second-level nodes, and the i-th associated attribute set is set as a third-level node;
[0147] When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is equal to the number of second associated attributes of the j-th associated attribute set, and when the j-th business data attribute association degree of the i-th association degree set is equal to the i-th business data attribute association degree of the j-th association degree set, the i-th business data attribute is set as the first-level first node, the j-th business data attribute is set as the first-level second node, and the intersection association attribute of the i-th associated attribute set and the j-th associated attribute set is set as the common second-level node of the first-level first node and the second-level second node;
[0148] According to the association tree fusion rule, the first business data attribute association tree, the second business data attribute association tree, and finally the Qth business data attribute association tree are fused to generate the business data attribute tree.
[0149] Furthermore, the network level connection module 15 includes the following execution steps:
[0150] The N-level first cascade subnetworks to the N-level M-th cascade subnetworks are fully connected in parallel to obtain a one-layer network, and the N-1-level first cascade subnetworks to the N-1-th L-th cascade subnetworks are fully connected in parallel to obtain a two-layer network, until the two-level cascade subnetworks are used as the N-1-layer network to construct the cascaded total network;
[0151] Among them, any sub-network of the k-layer network has a sub-network in the k-1-layer network, and the output of the sub-network in the k-1-layer network is the input of the sub-network in the k-layer network.
[0152] Furthermore, the graph function training module 16 includes the following execution steps:
[0153] According to the data lake warehouse, based on the business data attribute tree, cascade training data is collected;
[0154] According to the cascade training data, the N-level first cascade subnetwork to the N-level M-th cascade subnetwork to the second cascade subnetwork are trained separately to generate an initial graph mapping function set;
[0155] According to the cascade training data, cascade training is performed on the initial atlas mapping function set to generate the atlas mapping function.
[0156] Furthermore, the graph function training module 16 also includes the following execution steps:
[0157] Constructing a first cascade training loss function, wherein the first cascade training loss function is used to count the proportion of subnetworks whose loss values are greater than or equal to a loss threshold in the N-level first cascade subnetworks to the N-level M-th cascade subnetworks to the second cascade subnetworks;
[0158] Constructing a second cascade training loss function, wherein the second cascade training loss function is used to count the mean loss values in the N-level first cascade subnetwork to the N-level M-th cascade subnetwork to the second cascade subnetwork;
[0159] If, during a preset number of consecutive trainings, at least 95 percent of the first cascade training loss function is less than or equal to a first loss threshold, and the second cascade training loss function is less than or equal to a second loss threshold, the graph mapping function is output.
[0160] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0161] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0162] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0163] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0164] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0165] Although preferred embodiments of the present invention have been described, additional changes and modifications may occur to these embodiments once those skilled in the art understand the basic inventive concepts.
[0166] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention belong to the scope of the present invention and its equivalent technologies, the present invention is also intended to include these changes and variations.
Claims
1. A method for generating a business data graph in combination with a data lake warehouse, characterized in that: Applied to a server, the method comprises: In response to a business data graph generation request initiated by a user terminal, a business data attribute set is obtained for hierarchical analysis to generate a business data attribute tree; Extracting the first group of leaf nodes at level N and the first parent node at level N up to the Mth group of leaf nodes at level N and the Mth parent node at level N of the business data attribute tree; Based on the first group of leaf nodes at level N and the first parent node at level N, configure a first cascade sub-network at level N; based on the Mth group of leaf nodes at level N and the Mth parent node at level N, configure an Mth cascade sub-network at level N; Until the secondary leaf nodes and the primary nodes of the business data attribute tree are extracted, and the secondary cascade sub-network is configured; Based on the service data attribute tree, connecting the N-level first cascade subnetworks to the N-level M-th cascade subnetworks to the second-level cascade subnetworks to generate a cascaded total network; Based on the data lake warehouse, the cascaded total network is trained to generate a graph mapping function; According to the graph mapping function and the business data attribute tree, a business data graph is constructed and fed back to the user end; Wherein, based on the service data attribute tree, connecting the N-level first cascade subnetworks to the N-level M-th cascade subnetworks to the second-level cascade subnetworks to generate a cascaded total network, including: The N-level first cascade subnetworks to the N-level M-th cascade subnetworks are fully connected in parallel to obtain a one-layer network, and the N-1-level first cascade subnetworks to the N-1-th L-th cascade subnetworks are fully connected in parallel to obtain a two-layer network, until the two-level cascade subnetworks are used as the N-1-layer network to construct the cascaded total network; Among them, any sub-network of the k-layer network has a sub-network in the k-1-layer network, and the output of the sub-network in the k-1-layer network is the input of the sub-network in the k-layer network.
2. The method according to claim 1, characterized in that In response to a business data graph generation request initiated by a user, a business data attribute set is obtained for hierarchical analysis to generate a business data attribute tree, including: Extracting, according to the service data attribute set, a first service data attribute to a Qth service data attribute, wherein Q is the total number of service data attributes; Based on the first business data attribute, performing association analysis on the second business data attribute up to the Qth business data attribute to obtain a first business data attribute association tree; Based on the second business data attribute, performing association analysis on the first business data attribute, the third business data attribute, and up to the Qth business data attribute to obtain a second business data attribute association tree; Until, based on the Qth business data attribute, performing association analysis on the first business data attribute to the Q-1th business data attribute, to obtain a Qth business data attribute association tree; The first service data attribute association tree, the second service data attribute association tree, and finally the Qth service data attribute association tree are integrated to generate the service data attribute tree.
3. The method according to claim 2, characterized in that Based on the first service data attribute, performing correlation analysis on the second service data attribute to the Qth service data attribute to obtain a first service data attribute correlation tree, including: According to the first service data attribute, collecting a reference data sequence; According to the second service data attribute to the Qth service data attribute, collecting and comparing data sequences; Performing grey correlation analysis on the reference data sequence and the comparison data sequence to obtain a second business data attribute correlation degree up to a Qth business data attribute correlation degree; A business data attribute set and a correlation degree set that are greater than or equal to a correlation degree threshold value are extracted from the second business data attribute correlation degree to the Qth business data attribute correlation degree, and the first business data attribute correlation tree is constructed.
4. The method according to claim 2, characterized in that The first service data attribute association tree, the second service data attribute association tree and the Qth service data attribute association tree are integrated to generate the service data attribute tree, including: Extracting the i-th business data attribute association tree and the j-th business data attribute association tree from the first business data attribute association tree, the second business data attribute association tree, and up to the Q-th business data attribute association tree; The i-th business data attribute association tree includes an i-th association attribute set and an i-th association degree set, and the j-th business data attribute association tree includes a j-th association attribute set and a j-th association degree set; Configure association tree fusion rules: When the j-th business data attribute belongs to the i-th associated attribute set, and the i-th business data attribute does not belong to the j-th associated attribute set, the i-th business data attribute is set as a first-level node, the non-intersecting associated attributes of the i-th associated attribute set and the j-th associated attribute set are set as second-level nodes, the j-th associated attribute set is set as a third-level node, and the third-level node is a child node of the j-th business data attribute, and the j-th business data attribute belongs to the non-intersecting associated attributes; When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute does not belong to the i-th associated attribute set, the j-th business data attribute is set as a first-level node, the non-intersecting associated attributes of the j-th associated attribute set and the i-th associated attribute set are set as second-level nodes, the i-th associated attribute set is set as a third-level node, and the third-level node is a child node of the i-th business data attribute, and the i-th business data attribute belongs to the non-intersecting associated attributes; When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is equal to the number of second associated attributes of the j-th associated attribute set, and when the j-th business data attribute association degree of the i-th association degree set is greater than the i-th business data attribute association degree of the j-th association degree set, the i-th business data attribute is set as a first-level node, the non-intersection associated attributes of the i-th associated attribute set and the j-th associated attribute set are set as second-level nodes, and the j-th associated attribute set is set as a third-level node; When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is equal to the number of second associated attributes of the j-th associated attribute set, and when the j-th business data attribute association degree of the i-th association degree set is less than the i-th business data attribute association degree of the j-th association degree set, set the j-th business data attribute as a first-level node, set the non-intersection associated attributes of the j-th associated attribute set and the i-th associated attribute set as a second-level node, and set the i-th associated attribute set as a third-level node; When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is greater than the number of second associated attributes of the j-th associated attribute set, the i-th business data attribute is set as a first-level node, the non-intersection associated attributes of the i-th associated attribute set and the j-th associated attribute set are set as second-level nodes, and the j-th associated attribute set is set as a third-level node; When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is less than the number of second associated attributes of the j-th associated attribute set, the j-th business data attribute is set as a first-level node, the non-intersection associated attributes of the j-th associated attribute set and the i-th associated attribute set are set as second-level nodes, and the i-th associated attribute set is set as a third-level node; When the i-th business data attribute belongs to the j-th associated attribute set, and the j-th business data attribute belongs to the i-th associated attribute set, and the number of first associated attributes of the i-th associated attribute set is equal to the number of second associated attributes of the j-th associated attribute set, and when the j-th business data attribute association degree of the i-th association degree set is equal to the i-th business data attribute association degree of the j-th association degree set, the i-th business data attribute is set as the first-level first node, the j-th business data attribute is set as the first-level second node, and the intersection association attribute of the i-th associated attribute set and the j-th associated attribute set is set as the common second-level node of the first-level first node and the second-level second node; According to the association tree fusion rule, the first business data attribute association tree, the second business data attribute association tree, and finally the Qth business data attribute association tree are fused to generate the business data attribute tree.
5. The method according to claim 1, characterized in that Based on the data lake warehouse, the cascaded total network is trained to generate a graph mapping function, including: According to the data lake warehouse, based on the business data attribute tree, cascade training data is collected; According to the cascade training data, the N-level first cascade subnetwork to the N-level M-th cascade subnetwork to the second cascade subnetwork are trained separately to generate an initial graph mapping function set; According to the cascade training data, cascade training is performed on the initial atlas mapping function set to generate the atlas mapping function.
6. The method according to claim 5, characterized in that According to the cascade training data, cascade training is performed on the initial atlas mapping function set to generate the atlas mapping function, including: Constructing a first cascade training loss function, wherein the first cascade training loss function is used to count the proportion of subnetworks whose loss values are greater than or equal to a loss threshold in the N-level first cascade subnetworks to the N-level M-th cascade subnetworks to the second cascade subnetworks; Constructing a second cascade training loss function, wherein the second cascade training loss function is used to count the mean loss values in the N-level first cascade subnetwork to the N-level M-th cascade subnetwork to the second cascade subnetwork; If during a preset number of consecutive trainings, in at least 95% of the training times, the first cascade training loss function is less than or equal to the first loss threshold, and the second cascade training loss function is less than or equal to the second loss threshold, the graph mapping function is output.
7. A business data graph generation system combined with a data lake warehouse, characterized in that: A method for generating a business data graph in combination with a data lake warehouse according to any one of claims 1 to 6 is applied to a server, and the system includes: A data attribute analysis module, which is used to respond to a business data graph generation request initiated by a user terminal, obtain a business data attribute set for hierarchical analysis, and generate a business data attribute tree; A node hierarchical extraction module, the node hierarchical extraction module is used to extract the first group of leaf nodes at level N and the first parent node at level N up to the Mth group of leaf nodes at level N and the Mth parent node at level N of the business data attribute tree; A cascade network configuration module, the cascade network configuration module is used to configure the N-level first cascade sub-network based on the N-level first group of leaf nodes and the N-level first parent node, and configure the N-level M-th cascade sub-network based on the N-level M-th group of leaf nodes and the N-level M-th parent node; A secondary relationship construction module, the secondary relationship construction module is used to extract the secondary leaf nodes and the primary nodes of the business data attribute tree and configure a secondary cascade sub-network; A network level connection module, the network level connection module is used to connect the N-level first cascade subnetwork to the N-level M-th cascade subnetwork and to the second-level cascade subnetwork based on the service data attribute tree to generate a cascaded total network; A graph function training module, which is used to perform training on the cascaded total network based on the data lake warehouse to generate a graph mapping function; A data graph construction module is used to construct a business data graph based on the graph mapping function and the business data attribute tree and feed it back to the user end.
8. The system according to claim 7, characterized in that The execution steps of the data attribute analysis module include: Extracting, according to the service data attribute set, a first service data attribute to a Qth service data attribute; Based on the first business data attribute, performing association analysis on the second business data attribute up to the Qth business data attribute to obtain a first business data attribute association tree; Based on the second business data attribute, performing association analysis on the first business data attribute, the third business data attribute, and up to the Qth business data attribute to obtain a second business data attribute association tree; Until, based on the Qth business data attribute, performing association analysis on the first business data attribute to the Q-1th business data attribute, to obtain a Qth business data attribute association tree; The first service data attribute association tree, the second service data attribute association tree, and finally the Qth service data attribute association tree are integrated to generate the service data attribute tree.
9. The system according to claim 8, characterized in that The execution steps of the data attribute analysis module also include: According to the first service data attribute, collecting a reference data sequence; According to the second service data attribute to the Qth service data attribute, collecting and comparing data sequences; Performing grey correlation analysis on the reference data sequence and the comparison data sequence to obtain a second business data attribute correlation degree up to a Qth business data attribute correlation degree; A business data attribute set and a correlation degree set that are greater than or equal to a correlation degree threshold value are extracted from the second business data attribute correlation degree to the Qth business data attribute correlation degree, and the first business data attribute correlation tree is constructed.
Citation Information
Patent Citations
Data query method and device applied to big data, equipment and product
CN113722600A
Internet of Things system
CN116368355A