Data element high-quality construction and application integrated operation platform
By building an integrated operation platform for high-quality construction and application of data elements, we have solved the format differences and redundancy problems in multi-source heterogeneous data management, achieved unified integration and efficient application of data resources, dynamically adapted to business changes, and improved the efficiency of releasing data value.
Patent Information
- Application Number
- CN202511032072.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-10-17
AI Technical Summary
In existing technologies, data element management lacks a full-process integrated management system. The multi-source heterogeneity of data leads to large format differences and redundancy. Traditional management methods cannot capture the cross-fusion characteristics of data links, the matching efficiency of data application scenarios is low, and there is a lack of dynamic iterative optimization mechanisms, which makes it difficult to release the value of data.
Build an integrated operation platform for high-quality construction and application of data elements, including a data resource base construction module, an element association map generation module, a link cross-fusion analysis module, an application scenario semantic mapping module and an operation efficiency evaluation configuration module. Through cleaning and deduplication, hierarchical labeling, cross-link screening, business scenario mapping and dynamic iterative optimization, unified management and efficient application of data resources can be achieved.
It achieves unified integration and standardization of multi-source heterogeneous data, clearly presents data transmission paths, captures cross-correlation features, improves data application matching efficiency, systematically manages data operations, dynamically adapts to business changes, and ensures the continuous release of data value.
Smart Images

Figure CN120804081A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data management, in particular to a data element high-quality construction and application integrated operation platform. BACKGROUND
[0002] With the in-depth development of digital economy, the value of data as a key production factor is increasingly prominent, and its application scenarios in government services, financial risk control, intelligent manufacturing, etc. are constantly expanding. However, the current management and application of data elements still face many challenges: on the one hand, data sources show multi-source and heterogeneous characteristics, covering structured databases, unstructured texts, streaming sensor data and other types. The data formats of different sources differ significantly, and data duplication and redundancy are common, making it difficult to establish a unified standard for data resource base construction and form an effective data sharing and reuse mechanism. On the other hand, the association between data elements is complex and dynamic, and traditional data management methods mostly use linear link analysis, which cannot capture the cross-fusion characteristics between different data links, resulting in link breakage or association failure in cross-scene data applications.
[0003] At the data application level, the mapping of data and business scenarios in the prior art mostly relies on manual configuration, lacking an automated semantic association mechanism, resulting in low matching efficiency of data application scenarios and difficulty in adapting to rapid iteration of business scenarios. At the same time, the operation process of data elements lacks systematic performance evaluation methods, making it difficult to accurately grasp the actual application effect of each data link, and the optimization and adjustment of data links mostly rely on experience, lacking a dynamic iteration mechanism based on actual feedback. In addition, the management of data elements is often limited to a single link, such as data cleaning, link analysis or scenario matching, and a full-process integrated management system from data resource construction to application operation has not been formed, resulting in limited release of the value of data elements and difficulty in meeting the demand for high-quality data applications in various industries.
[0004] In the prior art, some schemes attempt to realize centralized management of data through data warehouses or data lakes, but such schemes focus on data storage and integration, and lack sufficient mining of the association between data elements, making it difficult to build an effective data link system. Another scheme builds a data association network through knowledge graph technology, but mostly stays at the construction level of static graphs, lacks a dynamic mapping mechanism with business application scenarios, and does not involve performance evaluation and iterative optimization in the operation process of data elements, resulting in a disconnection between data graphs and actual application needs, making it difficult to support efficient operation of data elements. Therefore, how to build a technical scheme that can realize integrated management of data elements from resource base to application operation throughout the whole process has become a problem to be solved in the field of data element management. SUMMARY
[0005] The purpose of the present invention is to provide an integrated operation platform for high-quality construction and application of data elements to solve the problems raised in the above-mentioned background technology.
[0006] To achieve the above objectives, the present invention provides an integrated operation platform for high-quality construction and application of data elements, which includes: The data resource base construction module acquires multi-source heterogeneous data resources, cleans and de-duplicates data items, identifies core data fields and relationships, establishes a standard data element set, and forms a data construction basic model; The element association graph generation module constructs a basic model based on the data, performs hierarchical annotation on the data elements, establishes a data association directed link from the basic layer to the application layer based on the annotation results, collects the identification information and adjacent node relationships of all nodes in the link, and constructs the element association graph architecture; The link cross-fusion analysis module extracts the data link node sequence according to the element association graph architecture, compares the overlapping nodes and counts the frequency of the end nodes, filters the cross links, and obtains the data cross link set; The application scenario semantic mapping module collects the business scenarios associated with the end nodes in the data cross link set, sorts them by frequency of occurrence, and matches the link end scenarios to obtain a data application semantic mapping group; The operation efficiency evaluation configuration module counts the graph application nodes associated with each scenario based on the data application semantic mapping group, divides the data links into corresponding nodes, establishes the node and data link application configuration relationship structure, and generates a data element operation configuration table; The dynamic iterative optimization management module monitors the usage feedback of the data element operation configuration table, analyzes the access popularity and link abnormality rate of the configuration node, adjusts the data element annotation rules and link connection logic, and updates the data construction basic model and element association graph architecture.
[0007] Preferably, the data construction basic model includes a data element standard definition structure, a core field association fragment set, and a data item cleaning weight model; the element association graph architecture includes a data element hierarchical annotation relationship, a directed link node chain, and a node adjacency relationship set; the data cross-link set includes a link overlapping node set, an end node occurrence frequency distribution, and a cross-node screening result; the data application semantic mapping group includes a business scenario frequency ranking, a link end scenario semantic mapping, and an application scenario matching result; the data element operation configuration table includes an application node identifier, a data link grouping result, and an application configuration relationship mapping result; the dynamic iterative optimization management module includes a feedback monitoring submodule, a heat analysis submodule, and a rule adjustment submodule.
[0008] Preferably, the data resource base construction module includes: The data cleaning submodule acquires multi-source heterogeneous data resources, performs a field-level cleaning operation on the data resources, extracts an effective field set of each data item, records a position index of each field in the data item, compares a first occurrence position of a core field in a field arrangement list and a field quantity, classifies according to data items, and obtains a core field position distribution result; The standard definition submodule extracts a field segment corresponding to the core field in the data resources according to the core field position distribution result, intercepts the field based on an index interval of the core field in the data item, constructs a field segment set according to a position of each core field for intercepting the field, and recombines the field segment set in combination with a data item to which the field belongs, to obtain a core field standard segment set; The model generation submodule counts an occurrence frequency value of all core fields according to the core field standard segment set, performs position rearrangement processing on the field segment set based on an original order arrangement of the core field in the data resources, sequentially splices the standard segments of the multiple core fields in the same data item according to a first occurrence position, classifies and integrates the standard result corresponding to each data item, and acquires a data construction base model.
[0009] Preferably, the element correlation graph generation module comprises: The hierarchical labeling submodule compares and sorts all data elements according to attribute label priority values of the data elements based on the data construction base model in combination with business attribute labels of the data elements, rearranges positions of the data elements according to an order from a base layer to an application layer, establishes a reordering column index table, and obtains data element ordering index values; The link construction submodule acquires a set of adjacent nodes in a data element reordering column according to the data element ordering index values, records a connection direction of each pair of adjacent nodes, integrates structure edge information, and generates data correlation directed link graph data; The architecture extraction submodule acquires node identifiers and adjacent node pair relationships in all connection edges according to the data correlation directed link graph data, constructs a node mapping table according to adjacent structure relationships, stores upstream and downstream relationship types and connection directions between the data elements, and acquires an element correlation graph architecture.
[0010] Preferably, the link cross-fusion analysis module comprises: The link extraction submodule acquires node sequence in any two data links based on the element correlation graph architecture, sequentially extracts node identifier information under each link and establishes a data link node mapping set, marks data item identifiers and link length parameters to which the links belong, and acquires data link node identifier values; The overlap comparison submodule calls node identification sequences of any two data links according to the data link node identification value, performs an overlap comparison operation on node sets of the two links, extracts all end node identifications in the overlap, respectively counts the number of occurrences of the node identifications in different links, compares the number of occurrences with a link cross judgment reference value piece by piece, screens link pairs with a deviation value less than or equal to the reference value, and establishes a link overlap identification set satisfying the condition; The link screening submodule queries original data link identifications according to the link combination corresponding to the identification, integrates data item identifications and link cross node information, and establishes a data link relationship chain table to generate a data cross link set according to the link overlap identification set satisfying the condition.
[0011] Preferably, the application scenario semantic mapping module comprises: The scenario collection submodule collects a business scenario set associated with each end node in the data cross link set based on the end node, performs index mapping between the data link and the end scenario, and generates a link end business scenario group. The frequency statistics submodule performs a repeated number counting operation on all business scenarios based on the link end business scenario group, records the number of occurrences of each business scenario in the data link set, and performs a high-to-low ordering processing according to the number of occurrences to obtain an ordered business scenario sequence. The mapping determination submodule performs matching judgment on the scenario set corresponding to the end node in the data link according to the ordered business scenario sequence, screens the scenario item closest to the front position in the ordered sequence in each link as the application semantic category corresponding to the link, integrates the mapping scenarios of all data links, and obtains a data application semantic mapping group.
[0012] Preferably, the operation efficiency evaluation configuration module comprises: The node extraction submodule collects a graph application node corresponding to each scenario according to the data application semantic mapping group, records the number of associated link identifications and corresponding data link sets in each application node, determines the matching index between the business scenario and the graph node, and obtains a scenario attribution node identification value. The link classification submodule divides the corresponding data link into each graph application node according to the business scenario as a classification basis based on the scenario attribution node identification value, establishes a bidirectional corresponding structure between the data link identification and the node identification, extracts a link identification list attributed to each node, and obtains a node link attribution number value. The structure generation submodule integrates the graph application node and the subordinate data link identification according to the node link attribution number value, outputs an application node index, a corresponding business scenario and a total number of links, determines the attribution of the node and the data link, and generates a data element operation configuration table.
[0013] Preferably, the dynamic iterative optimization management module comprises: The feedback monitoring submodule monitors the use feedback information of the data element operation configuration table, collects the access records and link exception logs of the configuration nodes, establishes a feedback data repository, and obtains the configuration node use feedback results; The heat analysis submodule is based on the configuration node use feedback results, and the access frequency and link exception rate of each configuration node are counted, the rationality of the data element annotation rule and the stability of the link connection logic are analyzed, and the node heat analysis result is obtained; The rule adjustment submodule adjusts the business attribute tag priority value of the data element according to the node heat analysis result, optimizes the data element position rearrangement logic and the link connection direction record mode, updates the data construction basic model and the element association graph architecture.
[0014] Preferably, the feedback monitoring submodule comprises a log collection unit and a storage and arrangement unit, the log collection unit obtains the access operation logs and link exception error logs of the configuration nodes, and the storage and arrangement unit stores the log information according to the node identifier and the link identifier, and establishes a structured feedback data set.
[0015] Preferably, the rule adjustment submodule comprises a tag adjustment unit and a logic optimization unit, the tag adjustment unit corrects the priority value of the data element business attribute tag according to the node heat analysis result, and the logic optimization unit optimizes the index table establishment mode of the data element position rearrangement and the record rule of the link connection direction.
[0016] Compared with the prior art, the beneficial effects of the present application are: The platform processes the multi-source heterogeneous data through the data resource basement construction module, establishes a standard data element set, forms a unified data construction basic model, and effectively solves the problems of data format confusion and redundancy in the prior art. The module identifies the core data field and the association relationship, so that data from different sources can be integrated under a unified standard, eliminating the barriers between data, and providing a consistent and reliable basis for subsequent processing of data elements.
[0017] The element association graph generation module performs data element hierarchical annotation based on the data construction basic model, and establishes an associated directed link from the basic layer to the application layer, and then constructs the element association graph architecture. This process breaks through the limitations of traditional linear link analysis, clearly presents the transmission path of data elements from basic resources to application scenarios through hierarchical annotation and directed link establishment, and makes the association relationship between data elements more intuitive and easy to trace. At the same time, the collection of node identifier information and adjacency relationship in the link lays a foundation for subsequent link analysis and application scenario matching, and makes the flow process of data elements controllable.
[0018] The link cross-fusion analysis module can capture the cross-linking characteristics between different data links by extracting the data link node sequence, screening the cross-link and forming a data cross-link set. This function avoids the information omission that may occur in single link analysis, and makes the cross-link hidden in complex data relationships visible, which provides the possibility for the comprehensive application of data elements in multiple scenarios. Through the comparison of overlapping nodes and the frequency statistics of end nodes, the screening of cross-link is more accurate, which ensures the effectiveness of subsequent application scenario matching.
[0019] The application scenario semantic mapping module collects and sorts the associated business scenarios based on the end nodes in the cross-link set, and realizes the automatic mapping of data links and business application scenarios. This mapping mechanism does not require human intervention, and through frequency sorting and scenario matching, data links can be quickly matched to related business scenarios, reducing human error in data and scenario matching process, while improving matching efficiency. Precise mapping of business scenarios and data links enables data elements to better meet actual application needs, improving the applicability of data in business scenarios.
[0020] The operation efficiency evaluation configuration module establishes the application configuration relationship structure of nodes and data links by counting the graph application nodes associated with each scenario, and generates an operation configuration table, making the operation process of data elements more systematic. The configuration table clearly presents the correspondence between data links and application nodes, facilitating the overall management of data element applications. By dividing data links into corresponding nodes, the fine-grained configuration of data element operation is realized, so that data applications in different scenarios can find corresponding management nodes, improving the orderliness and controllability of data element operation.
[0021] The dynamic iterative optimization management module adjusts the data element labeling rules and link connection logic by monitoring the use feedback of the operation configuration table, analyzing the access frequency and link abnormal rate, and realizes the dynamic optimization of data element management. This iterative mechanism based on actual feedback enables the data construction base model and element association graph architecture to update with changes in application needs, avoiding the problem of disconnection between data elements and application needs in static management mode. Through continuous adjustment and optimization, the platform can continuously adapt to changes in business scenarios, maintain the efficiency and adaptability of data element operation, and ensure that the value of data elements can be continuously released in long-term operation. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 The working principle diagram of the data element high-quality construction and application integrated operation platform described in the present application; Figure 2 The flowchart of detailed composition of each module; Figure 3A flow chart of the data resource base construction module working process; Figure 4 A flow chart of the element correlation graph generation module working process. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0024] Please refer to Figures 1-4 The present application provides a data element high-quality construction and application integrated operation platform, which comprises: The data element high-quality construction and application integrated operation platform comprises a data resource base construction module, an element correlation graph generation module, a link cross-fusion analysis module, an application scenario semantic mapping module, an operation efficiency evaluation configuration module and a dynamic iteration optimization management module, and each module cooperatively operates to realize the construction, correlation analysis, scenario mapping, operation configuration and dynamic optimization of data elements, and the specific process is as follows: The data resource base construction module acquires multi-source heterogeneous data resources, performs cleaning and deduplication processing on data items, identifies core data fields and the correlation between these fields, and then establishes a standard data element set to form a data construction base model. The element correlation graph generation module labels data elements according to the data construction base model, constructs data correlation directed links from the base layer to the application layer according to the labeling results, collects the identification information and adjacent node relationship of all nodes in the link, and forms an element correlation graph architecture in this way. The link cross-fusion analysis module extracts data link node sequences according to the element correlation graph architecture, compares overlapping nodes and counts the frequency of end nodes, filters out cross links, and obtains a data cross-link set. The application scenario semantic mapping module collects business scenarios associated with end nodes in the data cross-link set, sorts them according to the frequency of occurrence, and matches them with link end scenarios to obtain a data application semantic mapping group. The operation efficiency evaluation configuration module counts the graph application nodes associated with each scenario based on the data application semantic mapping group, divides the data link to the corresponding node, establishes a node and data link application configuration relationship structure, and generates a data element operation configuration table. The dynamic iteration optimization management module monitors the use feedback of the data element operation configuration table, analyzes the access popularity and link abnormal rate of the configuration nodes, adjusts the data element labeling rules and link connection logic, and updates the data construction base model and the element correlation graph architecture.
[0025] Embodiment 1: The data construction base model covers data element standard definition structure, core field association fragment set, and data item cleaning weight model. The data element standard definition structure clearly defines the standardized attributes of each data element, such as name, type, length, and constraint conditions, so that data elements from different sources have a unified description specification; the core field association fragment set records the association fragment between the core fields, which reflects the internal relationship of the data fields in the business logic; and the data item cleaning weight model gives corresponding weights to the cleaning operations of different data items in the data cleaning process, which is used to measure the influence degree of the cleaning operation on the data quality. The element association graph architecture includes data element hierarchical annotation relationship, directed link node chain, and node adjacency relationship set. The data element hierarchical annotation relationship reflects the annotation of data elements at different levels such as the basic layer, the intermediate layer, and the application layer, and clearly defines the hierarchical attribution of the data elements; the directed link node chain is composed of a series of nodes connected in order, and the pointing relationship between the nodes reflects the direction of data flow; and the node adjacency relationship set records the association between each node and its adjacent nodes, including direct adjacent and indirect adjacent node information. The data cross-link set includes link overlapping node set, end node frequency distribution, and cross-node screening result. The link overlapping node set refers to the summary of nodes that appear in multiple data links; the end node frequency distribution statistics the number of times and the proportion of each end node appearing in all links; and the cross-node screening result is the node information with cross-association characteristics determined after screening. The data application semantic mapping group includes business scenario frequency ranking, link end scenario semantic mapping, and application scenario matching result. The business scenario frequency ranking is the sorting of the frequency of each business scenario appearing in the data link from high to low; the link end scenario semantic mapping establishes the semantic correspondence between the data link end node and the business scenario; and the application scenario matching result is the corresponding relationship result obtained after matching the data link with the business scenario. The data element operation configuration table includes application node identification, data link grouping result, and application configuration relationship mapping result. The application node identification is the unique marking of each application node; the data link grouping result is the result of dividing the data link into different groups according to certain rules; and the application configuration relationship mapping result records the configuration association relationship between the application node and the data link. The dynamic iteration optimization management module includes a feedback monitoring submodule, a heat analysis submodule, and a rule adjustment submodule. The feedback monitoring submodule is responsible for monitoring and collecting relevant feedback information during the platform operation; the heat analysis submodule is used to analyze the access heat and other conditions of each configuration node; and the rule adjustment submodule adjusts the relevant rules according to the analysis result.
[0026] The data cleaning submodule in the data resource base construction module first obtains data resources from multiple different sources and with different data structures. These data resources may include structured data, semi-structured data, and unstructured data. After obtaining the data resources, the data cleaning submodule performs field-level cleaning operations on the data resources. This operation involves format verification of data fields, outlier processing, missing value supplementation, etc. Through the cleaning operation, the valid field set in each data item is extracted, and the position index of each field in the corresponding data item is recorded. This index can accurately identify the specific position of the field in the data item. Afterwards, the relationship between the first appearance position of the core field in the field arrangement list and the number of fields is compared, and the data items are classified according to this relationship. Finally, the core field position distribution result is obtained, which clearly shows the position distribution of the core fields in different data items.
[0027] The standard definition submodule extracts the field fragments corresponding to the core fields in the data resource based on the core field position distribution results. During the extraction process, the fields are precisely intercepted based on the core field's index interval in the data item to ensure the accuracy of the intercepted field fragments. A set of field fragments is constructed based on the position of the intercepted field for each core field. These field fragments are then reorganized based on the data item to which the field belongs. This reorganization process ensures that the reorganized field fragments better reflect the overall characteristics of the data item. After this reorganization process, a set of standard core field fragments is obtained.
[0028] The model generation submodule counts the frequency of occurrence of all core fields in the data resource based on the core field standard fragment set. The frequency value reflects the number of times the core field appears in the data resource. Based on the original order of the core fields in the data resource, the field fragment set is rearranged to make the arrangement order of the field fragments more in line with the needs of data processing and analysis. For multiple standard fragments of core fields in the same data item, they are sequentially spliced according to their first appearance position in the data item, and the spliced results are integrated with the standard results corresponding to each data item, and classified according to the type of data item. Through such processing, the data is finally obtained to build the basic model.
[0029] The data cleaning submodule in the data resource base construction module acquires the multi-source heterogeneous data resources, performs a field-level cleaning operation on the data resources, extracts a valid field set of each data item, records a position index of each field in the data item, compares a first occurrence position of a core field in a field arrangement list and a field quantity, classifies according to the data item, and obtains a core field position distribution result. The standard definition submodule extracts a field segment corresponding to the core field in the data resources according to the core field position distribution result, intercepts the field based on an index interval of the core field in the data item, constructs a field segment set according to a position of each core field for intercepting the field, and recombines the field segment set in combination with a data item to which the field belongs, to obtain a core field standard segment set. The model generation submodule acquires a data construction base model according to the core field standard segment set, counts an occurrence frequency value of all core fields, performs position rearrangement processing on the field segment set based on an original order arrangement of the core field in the data resources, sequentially splices the standard segments of the multiple core fields in the same data item according to the first occurrence position, classifies and integrates the standard result corresponding to each data item, and obtains the data construction base model.
[0030] Embodiment 2 The hierarchical annotation submodule in the element association graph generation module takes the data construction base model as a reference, and expands processing on all data elements in combination with a business attribute label carried by the data element itself. The business attribute label of each data element corresponds to a priority value, and the hierarchical annotation submodule compares these priority values and sorts them according to the numerical value. According to the sorting result, the data elements are rearranged in the order from the base layer to the application layer, and a reordering column index table is established in this process, which records the position information of each data element in the reordered sequence, and further obtains a data element ordering index value.
[0031] After receiving the data element ordering index value, the link construction submodule extracts an adjacent node set from the data element reordering column according to the index value. For each pair of adjacent nodes, the link construction submodule records the connection direction between them, and clearly indicates that the data flows from which node to which node. Subsequently, the connection direction between the nodes, node identification and other structural edge information are integrated to form a data association directed link graph data, which completely presents the directed connection relationship between the data elements.
[0032] The architecture extraction submodule deeply processes the data correlation directed link graph data, collects node identifiers involved in all connection edges and adjacent node pair relationships. According to these adjacent structure relationships, a node mapping table is constructed, in which the upstream and downstream relationship types between various data elements, such as direct upstream and downstream relationships, indirect upstream and downstream relationships, etc., are stored in detail, and the connection direction between nodes is also recorded. Through such processing, the element correlation graph architecture is finally obtained, which clearly shows the hierarchical relationship and connection between data elements.
[0033] The link extraction submodule in the link cross-fusion analysis module collects node sequences in any two data links based on the element correlation graph architecture. For each link, the node identifier information contained therein is extracted in turn, and a data link node mapping set is established after these information is sorted. In this set, the data item identifier to which each link belongs and the link length parameter are also marked, and the link length parameter is represented by the number of nodes. Through these operations, the data link node identifier value is obtained.
[0034] The overlap comparison submodule uses the data link node identifier value to call the node identifier sequences of any two data links, and compares the node sets of the two links. In the comparison process, all end node identifiers in the overlapping part are extracted, and then the number of times each end node appears in different links is counted. The counted number is compared with the preset link cross judgment reference value, and the link pair combination with a deviation value less than or equal to the reference value is selected. After integrating the identification information of these combinations, a set of link overlap identifiers that meet the conditions is established.
[0035] The link screening submodule queries the link identifier in the original data according to the link combination corresponding to each identifier in the set of link overlap identifiers that meet the conditions. The data item identifier and the link cross node information are integrated to construct a data link relationship chain table, which clearly presents the cross correlation between links. Through the above series of processing, a data cross link set is finally generated, which contains all the filtered data link information with cross characteristics.
[0036] In the whole process, the submodules cooperate in order, the hierarchical labeling submodule provides basic sorting information for link construction, the directed link graph data generated by the link construction submodule is the direct basis for architecture extraction, and the element correlation graph architecture provides the core reference for the processing of the link cross-fusion analysis module. The node identifier value obtained by the link extraction submodule is the basis for overlap comparison, the result of overlap comparison provides a clear screening standard for link screening, and the data cross link set obtained through link screening finally presents the cross correlation characteristics between data links.
[0037] The link extraction submodule in the link cross-fusion analysis module collects node sequences in any two data links based on the element association graph architecture, sequentially extracts node identification information under each link and establishes a data link node mapping set, marks data item identification and link length parameters to which each link belongs, and obtains data link node identification values. The overlap comparison submodule compares the node identification sequences of any two data links according to the data link node identification values, performs an overlap comparison operation on the node sets of the two links, extracts all end node identifications in the overlap, respectively counts the number of times the nodes appear in different links, compares the times with link cross judgment reference values one by one, screens link pairs with deviation values less than or equal to the reference values, and establishes a link overlap identification set meeting the conditions. The link screening submodule queries original data link identifications according to the link combinations corresponding to the identifications, integrates data item identifications and link cross node information, and establishes a data link relationship linked list to generate a data cross link set according to the link overlap identification set meeting the conditions.
[0038] Embodiment 3 The scene collection submodule in the application scenario semantic mapping module takes the end nodes in the data cross link set as the starting point, collects business scenarios associated with each end node one by one, and forms a business scenario set. For each data link, an index mapping between the data link and the end scenario is established to clearly indicate the end scenario position information corresponding to the data link. Through such processing, a link end business scenario group is generated, which contains all data links and their corresponding end business scenario information.
[0039] The frequency statistics submodule processes the link end business scenario group and performs a repetition number statistics operation on all business scenarios in the group. During the statistics process, the specific number of times each business scenario appears in the data link set is recorded, and then the business scenarios are sorted from high to low according to the number of times to form a sorted business scenario sequence. The business scenarios at the front of the sequence represent higher frequencies in the data link.
[0040] After receiving the sorted business scenario sequence, the mapping determination submodule performs matching judgment on the scenario set corresponding to the end nodes in each data link. The scenario item at the most front position in the sorted business scenario sequence is selected and compared with the scenario set of the end nodes of the link to screen out the scenario consistent with the scenario item as the application semantic category corresponding to the link. Such operation is performed on all data links to integrate the mapping scenarios of all links together to obtain a data application semantic mapping group. The group clearly presents the association relationship between each data link and the corresponding application semantic category.
[0041] The node extraction submodule in the operation performance evaluation configuration module extracts the graph application nodes corresponding to each scene based on the data application semantic mapping group. In the extraction process, the associated link identifiers in each application node are recorded, as well as the number of data link sets corresponding to each link identifier. Through this information, the matching index between the business scene and the graph node is determined, which is used to determine which graph application node a specific business scene corresponds to, and then the scene attribution node identifier value is obtained.
[0042] The link classification submodule divides the corresponding data link into each graph application node based on the scene attribution node identifier value and the business scene as the classification basis. In the division process, a bidirectional corresponding structure between the data link identifier and the node identifier is established, that is, the corresponding node identifier can be found through the data link identifier, and the data link identifier contained can also be found through the node identifier. The link identifier list attributed to each node is extracted from the bidirectional corresponding structure, the number of link identifiers in each list is counted, and the node link attribution number value is obtained.
[0043] The structure generation submodule integrates the structure mapping of the graph application node and the subordinate data link identifier according to the node link attribution number value. In the integration process, the application node index is output, which uniquely identifies each application node; at the same time, the business scene corresponding to each application node and the total number of links contained under the node are output. Through these information, the attribution relationship between the node and the data link is determined, and finally the data element operation configuration table is generated.
[0044] In the processing process, the calculation of the scene matching degree is involved, which can adopt the formula: Wherein, represents the scene matching degree, represents the number of scenes in the link end scene set that match the scene item with the highest priority, represents the total number of scenes in the link end scene set. Through this formula, the degree of scene matching can be quantified, which helps the mapping determination submodule to more accurately screen out the application semantic category corresponding to the link.
[0045] The submodules work cooperatively, the scene collection submodule provides basic data for frequency statistics, the results of frequency statistics provide ranking basis for mapping determination, the data application semantic mapping group obtained by mapping determination provides key information for the node extraction of the operation performance evaluation configuration module, and the results of node extraction support link classification and structure generation, and finally the data element operation configuration table formed presents the associated configuration relationship between the application node, the data link and the business scene.
[0046] The node extraction submodule in the operation performance evaluation configuration module extracts the graph application nodes corresponding to each scene according to the data application semantic mapping group, records the associated link identifiers and the number of corresponding data link sets in each application node, determines the matching index between the business scene and the graph node, and obtains the scene attribution node identifier value. The link classification submodule divides the corresponding data link into each graph application node according to the business scene as the classification basis based on the scene attribution node identifier value, establishes a bidirectional corresponding structure between the data link identifier and the node identifier, extracts the link identifier list attributed to each node, and obtains the node link attribution number value. The structure generation submodule integrates the graph application node and the subordinate data link identifier according to the node link attribution number value, outputs the application node index, the corresponding business scene and the total number of links, determines the attribution of the node and the data link, and generates the data element operation configuration table.
[0047] Embodiment 4: The feedback monitoring submodule in the dynamic iterative optimization management module continuously pays attention to the use feedback information of the data element operation configuration table, which covers the operation feedback of the user to the configuration table, the running state information automatically recorded by the system, and the like. The feedback monitoring submodule actively collects the access records of the configuration node, including the time of each access, the access source, the access duration and the like, and collects the link exception logs, such as the specific records of the link interruption, the data transmission error, the node response timeout and the like. The access records and the exception logs collected are stored centrally, a special feedback data storage library is established, the use feedback result of the configuration node is obtained through the arrangement and extraction of the data in the storage library, and the result comprehensively reflects various situations of the configuration node in actual use.
[0048] The heat analysis submodule takes the use feedback result of the configuration node as the input, and counts the access frequency of each configuration node, that is, calculates the total number of times that each node is accessed within a certain time. At the same time, the link exception rate is counted, that is, the proportion of the number of links with exceptions to the total number of links. Based on the access frequency and the link exception rate, the rationality of the data element labeling rule is analyzed, for example, whether the current labeling rule can accurately reflect the business attribute of the data element, whether it leads to deviation in node association; the stability of the link connection logic is analyzed, that is, whether the link connection mode is prone to interruption, error and other unstable situations. Through these analyses, the node heat analysis result is obtained, which contains the access heat level of each node, the main types and distribution of link exceptions and the like.
[0049] The rule adjustment submodule adjusts the business attribute tag priority value of the data element according to the node heat analysis result. For the data element associated with the node with high access heat and low link abnormal rate, the priority value of the business attribute tag of the data element is appropriately increased; for the data element associated with the node with low access heat or high link abnormal rate, the priority value of the data element is appropriately decreased. Meanwhile, the data element position rearrangement logic is optimized, for example, the factor weight referred to in rearrangement is adjusted, so that the sequence after rearrangement is more in line with actual access demand; the link connection direction recording mode is optimized, a clearer and more maintainable recording format is adopted, and link abnormality caused by improper recording mode is reduced. According to these adjustments and optimizations, the data construction basic model and the element association graph architecture are updated, so that they can better adapt to actual operation scenarios.
[0050] The log collection unit in the feedback monitoring submodule captures the access operation log of the configured node in real time, including detailed records of user login and subsequent query, modification, deletion and other operations on the node, and captures link abnormal error logs, such as error codes, error descriptions, occurrence times and other information generated by the system when link connection fails. The storage and arrangement unit receives the log information obtained by the log collection unit, classifies the access operation log according to the node identifier, classifies the abnormal error log according to the link identifier, stores the classified log information in the corresponding database table, forms a structured feedback data set, and facilitates subsequent query and analysis.
[0051] The label adjustment unit in the rule adjustment submodule receives the node heat analysis result, and for the business attribute tag of each data element, corrects the corresponding priority value according to the access heat and link abnormality of the associated node. For example, if the node associated with a certain data element has a significant increase in recent access frequency and fewer abnormalities, the label adjustment unit will increase the priority value of the business attribute tag of the data element. The logic optimization unit optimizes the index table establishment mode of data element position rearrangement, for example, increases the update frequency of the index table, or adjusts the sorting mode of the fields in the index table; at the same time, the recording rules of the link connection direction are optimized, the format requirements and storage location of the records are clarified, and the accuracy and integrity of the connection direction information are ensured.
[0052] Example 5: The data resource base construction module starts with the acquisition of multi-source heterogeneous data resources from different systems, platforms or databases, including structured tables, semi-structured documents and unstructured texts, etc. The module performs cleaning and deduplication processing on these data, removes redundant fields and duplicate records in data items, processes missing values and outliers, and extracts valid information in each data item. In this process, core data fields that play a key role in business logic are identified, and the association between these core fields is analyzed, such as inclusion relationship and mapping relationship, etc. Based on these analyses, a standard data element set is established, each data element has a clear definition and attributes, and finally a data construction base model is formed, which integrates various types of processed data information and provides a unified data base for subsequent modules.
[0053] The element association graph generation module takes the data construction base model as input and labels the data elements in it at different levels. According to the business attributes and importance of the data elements, they are divided into different levels such as basic layer, intermediate layer and application layer. According to the labeling results, data association directed links are constructed from the basic layer to the application layer, and the direction of the links reflects the flow path and dependency relationship of the data. Collect the identification information of all nodes in the link, that is, the unique identifier of each data element, and the relationship between each node and its adjacent nodes, such as direct connection and indirect connection. Through these information, the element association graph architecture is constructed, which intuitively displays the hierarchical structure and association path between data elements in the form of a graph.
[0054] The link cross-fusion analysis module extracts the node sequence in the data link based on the element association graph architecture, and each link is represented as a series of nodes connected in order. By comparing the node sequences of different links, overlapping nodes are found, that is, nodes that exist in multiple links, and the frequency of end nodes appearing in all links is counted. According to the distribution of overlapping nodes and the frequency of end nodes, cross-link is selected, that is, links that have node overlap and certain association, and the data cross-link set is obtained by integrating these cross-links, which reflects the cross-association characteristics between data links.
[0055] The application scenario semantic mapping module collects the business scenarios associated with each end node in the data cross-link set. These business scenarios are specific scenarios that data serves in actual applications, such as user analysis scenarios and risk assessment scenarios. The collected business scenarios are sorted by frequency, with the scenarios with higher frequency ranked first. The sorted scenarios are matched with the end link scenarios, that is, whether the scenarios corresponding to the end nodes of the link are consistent or associated with the scenarios ranked first, and the data application semantic mapping group is obtained through matching, which establishes the semantic association between data links and business scenarios.
[0056] The operation performance evaluation configuration module maps the data application semantics group, counts the graph application nodes associated with each business scenario, and these nodes are the nodes in the graph that directly serve the scenario. The data link is divided into the corresponding application node according to the association relationship, and it is clear which data link each node contains. The application configuration relationship structure between the node and the data link is established, that is, the function of the node, the role of the data link and their cooperation mode, etc. Based on the structure, the data element operation configuration table is generated, which records the configuration relationship of the application node, the data link and the business scenario in detail.
[0057] The dynamic iterative optimization management module continuously monitors the use feedback of the data element operation configuration table, including user evaluation of the configuration table, problems occurring in system operation, etc. The access heat of each configuration node is analyzed, that is, the frequency of node access, and the link abnormal rate, that is, the proportion of link failure. According to the analysis result, the data element marking rule is adjusted, for example, the standard of hierarchical marking is modified to make the marking more accurate; the link connection logic is adjusted, for example, the connection mode between nodes is changed to reduce the occurrence of abnormalities.
[0058] It should be noted that in this paper, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment.
[0059] Although the embodiments of the present application have been shown and described, it can be understood by those of ordinary skill in the art that various changes, modifications, replacements and variations of the embodiments can be made without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. An integrated operation platform for high-quality construction and application of data elements, characterized by: The platform includes: The data resource base construction module acquires multi-source heterogeneous data resources, cleans and de-duplicates data items, identifies core data fields and relationships, establishes a standard data element set, and forms a data construction basic model; The element association graph generation module constructs a basic model based on the data, performs hierarchical annotation on the data elements, establishes a data association directed link from the basic layer to the application layer based on the annotation results, collects the identification information and adjacent node relationships of all nodes in the link, and constructs the element association graph architecture; The link cross-fusion analysis module extracts the data link node sequence according to the element association graph architecture, compares the overlapping nodes and counts the frequency of the end nodes, filters the cross links, and obtains the data cross link set; The application scenario semantic mapping module collects the business scenarios associated with the end nodes in the data cross link set, sorts them by frequency of occurrence, and matches the link end scenarios to obtain a data application semantic mapping group; The operation efficiency evaluation configuration module counts the graph application nodes associated with each scenario based on the data application semantic mapping group, divides the data links into corresponding nodes, establishes the node and data link application configuration relationship structure, and generates a data element operation configuration table; The dynamic iterative optimization management module monitors the usage feedback of the data element operation configuration table, analyzes the access popularity and link abnormality rate of the configuration node, adjusts the data element annotation rules and link connection logic, and updates the data construction basic model and element association graph architecture.
2. The integrated operation platform for high-quality construction and application of data elements according to claim 1 is characterized in that: The data construction basic model includes a data element standard definition structure, a core field association fragment set, and a data item cleaning weight model; the element association graph architecture includes a data element hierarchical annotation relationship, a directed link node chain, and a node adjacency relationship set; the data cross-link set includes a link overlapping node set, an end node occurrence frequency distribution, and a cross-node screening result; the data application semantic mapping group includes a business scenario frequency ranking, a link end scenario semantic mapping, and an application scenario matching result; the data element operation configuration table includes an application node identifier, a data link grouping result, and an application configuration relationship mapping result; the dynamic iterative optimization management module includes a feedback monitoring submodule, a heat analysis submodule, and a rule adjustment submodule.
3. The integrated operation platform for high-quality construction and application of data elements according to claim 1 is characterized in that: The data resource base construction module includes: The data cleaning submodule acquires multi-source heterogeneous data resources, performs field-level cleaning operations on the data resources, extracts the valid field set of each data item, records the position index of each field in the data item, compares the relationship between the first appearance position of the core field in the field arrangement list and the number of fields, classifies them by data item, and obtains the core field position distribution result; The standard definition submodule extracts the field segments corresponding to the core fields in the data resources based on the core field position distribution results, intercepts the fields based on the index interval of the core fields in the data items, constructs a field segment set based on the position of each core field interception field, and reorganizes them in combination with the data items to which the fields belong to obtain a core field standard segment set; The model generation submodule counts the occurrence frequency values of all core fields based on the core field standard fragment set, performs position rearrangement processing on the field fragment set based on the original order of the core fields in the data resource, and sequentially splices the standard fragments of multiple core fields in the same data item according to the first appearance position. The standard results corresponding to each data item are integrated and classified to obtain data to build a basic model.
4. The integrated operation platform for high-quality construction and application of data elements according to claim 1 is characterized in that: The element association graph generation module includes: The hierarchical annotation submodule builds a basic model based on the data, combines the business attribute labels of the data elements, compares and sorts all the data elements according to the priority values of their attribute labels, rearranges the positions of the data elements in the order from the basic layer to the application layer, establishes a rearrangement sequence index table, and obtains the data element sorting index value; The link construction submodule obtains the set of adjacent nodes in the data element rearrangement sequence according to the data element sorting index value, records the connection direction of each pair of adjacent nodes, integrates the structural edge information, and generates data association directed link graph data; The architecture extraction submodule associates the directed link graph data according to the data, collects the node identifiers and adjacent node pair relationships in all connection edges, constructs a node mapping table based on the adjacent structure relationship, stores the upstream and downstream relationship types and connection directions between each data element, and obtains the element association graph architecture.
5. The integrated operation platform for high-quality construction and application of data elements according to claim 1 is characterized in that: The link cross-fusion analysis module includes: The link extraction submodule collects node sequences from any two data links based on the element association graph architecture, extracts node identification information under each link in turn and establishes a data link node mapping set, marks the data item identification and link length parameters of each link, and obtains the data link node identification value; The overlap comparison submodule calls the node identification sequences of any two data links based on the data link node identification values, performs an overlap comparison operation on the node sets of the two links, extracts all the end node identifications in the overlap, counts the number of times such nodes appear in different links, compares them one by one with the link intersection judgment reference value, selects the link pair combinations with deviation values less than or equal to the reference value, and establishes a link overlap identification set that meets the conditions; The link screening submodule queries the original data link identifier according to the link combination corresponding to the identifier based on the link overlap identifier set that meets the conditions, integrates the data item identifier and the link cross node information, establishes a data link relationship list, and generates a data cross link set.
6. The integrated operation platform for high-quality construction and application of data elements according to claim 1 is characterized in that: The application scenario semantic mapping module includes: The scenario collection submodule collects the business scenario set associated with each end node based on the end nodes in the data cross link set, performs index mapping between the data link and its end scenario, and generates a link end business scenario group; The frequency statistics submodule performs a repetition count operation on all business scenarios based on the link end business scenario group, records the number of times each business scenario appears in the data link set, and sorts them from high to low according to the number of appearances to obtain a sorted business scenario sequence; The mapping judgment submodule performs matching judgment on the scene set corresponding to the end node in the data link according to the sorted business scenario sequence, selects the scene item in each link that is at the front of the sorted sequence as the application semantic category corresponding to the link, integrates the mapping scenarios of all data links, and obtains the data application semantic mapping group.
7. The integrated operation platform for high-quality construction and application of data elements according to claim 1 is characterized in that: The operational effectiveness evaluation configuration module includes: The node extraction submodule collects the graph application nodes corresponding to each scenario based on the data application semantic mapping group, records the number of associated link identifiers and their corresponding data link sets in each application node, determines the matching index between the business scenario and the graph node, and obtains the scenario-attribution node identifier value; The link classification submodule divides the corresponding data links into various graph application nodes based on the scene-based node identification value and the business scenario, establishes a bidirectional correspondence structure between the data link identification and the node identification, extracts the link identification list of each node, and obtains the node link identification quantity value; The structure generation submodule integrates the structure mapping of the graph application nodes and the subordinate data link identifiers according to the node link ownership quantity value, outputs the application node index, the corresponding business scenario and the total number of links, determines the ownership of the nodes and data links, and generates a data element operation configuration table.
8. The integrated operation platform for high-quality construction and application of data elements according to claim 1 is characterized in that: The dynamic iterative optimization management module includes: The feedback monitoring submodule monitors the usage feedback information of the data element operation configuration table, collects the access records and link abnormality logs of the configuration nodes, establishes a feedback data repository, and obtains the usage feedback results of the configuration nodes; The heat analysis submodule calculates the access frequency and link abnormality rate of each configuration node based on the configuration node usage feedback result, analyzes the rationality of the data element annotation rules and the stability of the link connection logic, and obtains the node heat analysis result; The rule adjustment submodule adjusts the business attribute label priority value of the data element according to the node heat analysis results, optimizes the data element position rearrangement logic and link connection direction recording method, and updates the data construction basic model and element association graph architecture.
9. The integrated operation platform for high-quality construction and application of data elements according to claim 8 is characterized in that: The feedback monitoring submodule includes a log collection unit and a storage and sorting unit. The log collection unit obtains the access operation log and link abnormality error log of the configuration node. The storage and sorting unit classifies and stores the log information according to the node identifier and link identifier to establish a structured feedback data set.
10. The high-quality construction and application integrated operation platform for data elements according to claim 8 is characterized in that: The rule adjustment submodule includes a label adjustment unit and a logic optimization unit. The label adjustment unit corrects the priority value of the data element business attribute label according to the node heat analysis result. The logic optimization unit optimizes the index table establishment method for rearranging the data element positions and the recording rules of the link connection direction.
Citation Information
Patent Citations
Foreign trade data classification management system based on knowledge graph
CN120045712A
Data exchange system of nested tag structure
CN120162466A
Multi-source knowledge processing and querying method and device, equipment and medium
CN120197681A
Method and apparatus for processing electronic data
US20130290338A1