Multi-source heterogeneous data fusion governance method based on data base

By assessing the governance difficulty and volatility coefficient of multi-source heterogeneous data at the end of the data collection cycle and dynamically adjusting the governance strategy, the problem of difficulty in capturing the dynamic volatility characteristics of data sources in existing technologies is solved, thereby improving the accuracy and adaptability of data governance.

CN121765660AInactive Publication Date: 2026-03-31HUNAN TIANXIANGHE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-04
Publication Date
2026-03-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing technologies, multi-source heterogeneous data fusion governance methods cannot accurately perceive the dynamic fluctuation characteristics of data sources, resulting in low data governance quality and failing to meet enterprises' high requirements for the accuracy and timeliness of data assets.

Method used

By collecting data points from multiple heterogeneous data sources at the end of the current collection cycle, assessing the governance difficulty and volatility coefficient of each data source, and combining historical data for dynamic evaluation and differentiated governance, the dynamic changes of data sources can be captured and accurately quantified.

Benefits of technology

It improves the accuracy and adaptability of multi-source heterogeneous data fusion governance, ensures data governance quality, adapts to the dynamic changes of data sources, and improves the timeliness and accuracy of data assets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765660A_ABST
    Figure CN121765660A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data governance, in particular to a multi-source heterogeneous data fusion governance method based on a data base, and solves the technical problem of governance quality in the prior art. The method comprises the following steps: when a current acquisition period is ended, respectively acquiring data points updated by a plurality of data sources in multi-source heterogeneous data in the current acquisition period; for each data source, performing governance difficulty assessment according to the data distribution of each data point updated by the data source in the current acquisition period to obtain the corresponding governance difficulty of the data source in the current acquisition period; determining a data source fluctuation coefficient corresponding to the data source in the current acquisition period according to the treatment difficulty corresponding to the data source in the current acquisition period and the treatment difficulty corresponding to the plurality of historical acquisition periods; and according to the data source fluctuation coefficients corresponding to the plurality of data sources in the current acquisition period, carrying out data fusion treatment evaluation on the updated data points of the plurality of data sources in the multi-source heterogeneous data in the current acquisition period.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data governance technology, specifically to a multi-source heterogeneous data fusion governance method based on a data foundation. Background Technology

[0002] Against the backdrop of rapid development in the digital economy, enterprise business scenarios are becoming increasingly complex, and data is exhibiting significant characteristics of multi-source and heterogeneity. As a unified data infrastructure and architecture for enterprises, the data foundation bears the core responsibility of aggregating, processing, and integrating dispersed data. The resulting trusted data asset system is a crucial support for enterprise decision-making and business operations. Multi-source heterogeneous data fusion and governance, as a key aspect of data foundation construction, directly impacts the quality and availability of data assets, thereby determining the efficiency of data-driven business response and the accuracy of decision-making for enterprises.

[0003] In existing technologies, data governance processes often employ static result analysis methods, conducting governance work by pre-labeling the importance of different data source types. However, the fluctuation characteristics of different data sources vary significantly, and static labeling models struggle to accurately perceive the dynamic fluctuation characteristics of data sources. This results in insufficient accuracy in data description, an inability to adapt to changes in the difficulty of data source governance, and ultimately, low-quality data governance that fails to meet enterprises' high requirements for the accuracy and timeliness of data assets. Summary of the Invention

[0004] To address the technical problem of low data governance quality in existing technologies, the present invention aims to provide a multi-source heterogeneous data fusion and governance method based on a data foundation. The specific technical solution adopted is as follows: This application provides a multi-source heterogeneous data fusion and governance method based on a data foundation, including: At the end of the current collection period, collect each data point updated by multiple data sources in the multi-source heterogeneous data during the current collection period; For each data source, the governance difficulty is assessed based on the data distribution of each data point updated within the current collection period, thus obtaining the governance difficulty of the data source in the current collection period; the governance difficulty is used to characterize the complexity of the contradiction between data dispersion and feature concentration within the data source. For each data source, the data source fluctuation coefficient is determined based on the governance difficulty corresponding to the data source in the current collection period and the governance difficulty corresponding to multiple historical collection periods. The data source fluctuation coefficient is used to characterize the degree of fluctuation of the governance difficulty of the data source in the historical time series. Based on the data source fluctuation coefficients corresponding to multiple data sources in the current collection period, a data fusion governance assessment is conducted on each data point updated by multiple data sources in the current collection period in multi-source heterogeneous data.

[0005] The present invention has the following beneficial effects: Based on the above technical solution, at the end of the current collection cycle, this application collects data points updated by multiple data sources from the multi-source heterogeneous data during the current collection cycle. For each data source, the governance difficulty is assessed based on its data distribution to obtain the governance difficulty corresponding to the current collection cycle. Then, the data source fluctuation coefficient is determined by combining the historical governance difficulty. Finally, data fusion governance assessment is performed on each data point based on the fluctuation coefficients of multiple data sources. In this way, by introducing a governance difficulty assessment mechanism, this application can accurately quantify the data structure complexity within each data source. By introducing a data source fluctuation coefficient, it can effectively capture the dynamic change characteristics of the data source in historical time series, thereby solving the problem that existing static assessment methods cannot adapt to dynamic changes in data sources. Through dynamic assessment and differentiated governance, this application improves the accuracy and adaptability of multi-source heterogeneous data fusion governance and ensures the data governance quality of multi-source heterogeneous data. Attached Figure Description

[0006] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0007] Figure 1 This is a flowchart illustrating a multi-source heterogeneous data fusion and governance method based on a data foundation, provided in one embodiment of the present invention. Figure 2 This is a schematic diagram of the hardware structure of a multi-source heterogeneous data fusion and governance device based on a data base, provided in one embodiment of the present invention. Detailed Implementation

[0008] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of the multi-source heterogeneous data fusion governance method based on a data foundation proposed by the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0009] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0010] In all division and logarithmic operations covered in this application, a smoothing mechanism is employed to prevent computer program crashes or invalid values ​​from being generated due to a zero denominator or a zero input. Specifically, a positive correction factor is superimposed on the denominator term of the division operation or the argument term of the logarithmic function. For example, the value is This ensures the robustness and feasibility of the algorithm under extreme conditions.

[0011] The normalization function mentioned in this application Unless otherwise specified, all values ​​are normalized using maximum and minimum values. The maximum and minimum values ​​are preset empirical extreme values ​​derived from a large amount of historical experimental data. If the calculated result exceeds the [0,1] interval, it is restricted to the [0,1] range by a truncation function (i.e., if the result is less than 0, it is taken as 0, and if it is greater than 1, it is taken as 1) to eliminate the influence of outliers on the evaluation index.

[0012] The specific solution of the multi-source heterogeneous data fusion and governance method based on a data foundation provided by the present invention will be described in detail below with reference to the accompanying drawings.

[0013] Please see Figure 1 The diagram illustrates a flowchart of a multi-source heterogeneous data fusion governance method based on a data foundation according to an embodiment of the present invention. The method includes the following steps: Step 101: At the end of the current collection period, collect the data points updated by multiple data sources in the multi-source heterogeneous data during the current collection period.

[0014] Multi-source heterogeneous data refers to datasets originating from different business systems, with different data formats (such as structured data and semi-structured data), and varying data types (such as time-series data and discrete data points). For example, data sources may include enterprise business databases (such as heating system databases and e-commerce transaction databases), sensor acquisition devices, and third-party data interfaces. The current acquisition cycle refers to a pre-set data collection time interval. For example, the acquisition cycle can be set to 4 hours, meaning that a data acquisition and processing process is completed every 4 hours. At the end of each acquisition cycle, this application proactively obtains newly generated or updated data points from each data source within that acquisition cycle, forming the updated data source dataset for the current acquisition cycle.

[0015] In some embodiments, data acquisition is achieved through a database device deployed in a data infrastructure. This database device provides a standardized SQL interface to transmit updated data from various data sources to the computing units of the data infrastructure. The acquired data points must include both numerical information about the data itself and a timestamp of the acquisition, for subsequent time-series analysis.

[0016] For example, taking a heating system scenario, the updated data points collected from the heating database may include "number of heat exchange stations in operation", "pipeline temperature (degrees Celsius)", "boiler energy consumption (kilowatts)", etc., and each data point is associated with the corresponding collection time.

[0017] The above solution can obtain the latest dynamic data from various data sources, providing a basis for subsequent governance difficulty assessment and fluctuation analysis, and avoiding the disconnect between governance results and actual data status due to data lag.

[0018] Step 102: For each data source, assess the governance difficulty based on the data distribution of each data point updated within the current collection period to obtain the governance difficulty corresponding to the data source in the current collection period.

[0019] The governance difficulty characterizes the complexity of the contradiction between data dispersion and feature concentration within a data source. Data dispersion refers to the relatively scattered distribution of data points, lacking obvious clustering characteristics. Feature concentration refers to the tendency of data points to cluster in specific areas. When data exhibits both discrete points and multiple clustering areas, the complexity of data governance is high, requiring more refined governance strategies. The greater the governance difficulty, the more prominent the contradiction between data dispersion and feature concentration in the data source, the more difficult it is to extract effective data features, and the more complex the governance process.

[0020] Data distribution refers to the clustering state of data points in the sample space, including the degree of data concentration, dispersion, and clustering characteristics.

[0021] In some embodiments, for each data source, this application first obtains all data points within the current collection period, then analyzes the distribution characteristics of these data points, identifies clustering patterns in the data through clustering algorithms, and quantifies the unevenness of data distribution and clustering differences based on the clustering results, thereby obtaining a governance difficulty index that can reflect the inherent structural complexity of the data source.

[0022] For example, for the "user click volume" data source in the e-commerce transaction database, if the data points in the current period are distributed very unevenly in different time periods, with dense data in some time periods and sparse data in others, and the numerical differences between data points are large, it indicates that the data source has a high degree of data dispersion, the difficulty of feature concentration is high, and the corresponding governance difficulty value is high.

[0023] The above-mentioned approach assesses the difficulty of governance by focusing on the characteristics of data distribution, which breaks through the limitations of existing technologies that statically label the importance of data sources. It can dynamically reflect the real-time governance complexity of each data source and provide a basis for subsequent targeted governance.

[0024] Step 103: For each data source, determine the data source fluctuation coefficient corresponding to the current collection period based on the governance difficulty of the data source in the current collection period and the governance difficulty corresponding to multiple historical collection periods.

[0025] The data source volatility coefficient characterizes the degree of fluctuation in the governance difficulty of a data source over historical time. A higher volatility coefficient indicates more drastic changes in the governance difficulty of the data source across different periods, a more unstable data state, and a greater need for attention during the governance process.

[0026] The historical collection period refers to multiple consecutive collection periods preceding the current collection period. For example, if the current collection period is the nth period, the governance difficulty data of the (n-1), (n-2), (n-3), (n-4), and (n-5)th historical collection periods can be selected.

[0027] It should be noted that this application not only focuses on the governance difficulty of the current period, but also analyzes the temporal trend of governance difficulty changes in the data source across multiple historical collection periods, particularly identifying a deteriorating trend of continuously increasing governance difficulty, thereby obtaining a fluctuation coefficient that reflects the dynamic characteristics of the data source. For example, if the governance difficulty of a data source was 1.2, 2.5, and 1.8 in three consecutive historical periods, and the governance difficulty of the current period is 3.1, the fluctuation coefficient of the data source can be obtained by calculating the magnitude and dispersion of these values, thus determining whether the fluctuation of its governance difficulty is within a reasonable range.

[0028] The above solution introduces time-series governance difficulty data to accurately capture the dynamic fluctuation characteristics of data sources, solving the problem that existing technologies cannot detect data source fluctuations and providing support for subsequent differentiated governance.

[0029] Step 104: Based on the data source fluctuation coefficients corresponding to multiple data sources in the current collection period, conduct data fusion governance assessment on each data point updated by multiple data sources in the current collection period in the multi-source heterogeneous data.

[0030] Among them, data fusion governance assessment refers to comprehensively judging the reliability and importance of all updated data points by combining the fluctuation characteristics of each data source, and carrying out targeted governance operations such as data cleaning, feature extraction, and anomaly handling.

[0031] After obtaining the data source fluctuation coefficients of each data source, this application comprehensively considers the fluctuation situation of all data sources, determines the overall data source fluctuation characteristics of multi-source heterogeneous data, and performs fusion governance assessment on the data points within the current collection period based on the fluctuation coefficients of each data source, identifies the data that needs to be focused on, thereby realizing differentiated and dynamic governance of multi-source heterogeneous data.

[0032] For example, for data sources with larger fluctuation coefficients, the probability of anomalies in their data points is higher, and the sensitivity of anomaly detection needs to be increased during the governance assessment process. For data sources with smaller fluctuation coefficients, their data status is more stable, and the detection intensity can be appropriately reduced to improve governance efficiency.

[0033] Based on the above technical solution, at the end of the current collection cycle, this application collects data points updated by multiple data sources from the multi-source heterogeneous data during the current collection cycle. For each data source, the governance difficulty is assessed based on its data distribution to obtain the governance difficulty corresponding to the current collection cycle. Then, the data source fluctuation coefficient is determined by combining the historical governance difficulty. Finally, data fusion governance assessment is performed on each data point based on the fluctuation coefficients of multiple data sources. In this way, by introducing a governance difficulty assessment mechanism, this application can accurately quantify the data structure complexity within each data source. By introducing a data source fluctuation coefficient, it can effectively capture the dynamic change characteristics of the data source in historical time series, thereby solving the problem that existing static assessment methods cannot adapt to dynamic changes in data sources. Through dynamic assessment and differentiated governance, this application improves the accuracy and adaptability of multi-source heterogeneous data fusion governance and ensures the data governance quality of multi-source heterogeneous data.

[0034] As a possible embodiment of this application, step 102 above can be implemented through the following steps: Step 201: For each data source, perform cluster analysis based on the data points updated by the data source within the current collection period to obtain multiple data clusters.

[0035] Cluster analysis is a data analysis method that divides data points into multiple sets (i.e., data clusters) based on their similarity. Similarity can be measured by the distance between data points (such as Euclidean distance or Manhattan distance). A data cluster is a local set of data formed after clustering. Data points within the same data cluster have high similarity, while data points between different data clusters have large differences.

[0036] For example, this embodiment uses the Ordering Points To Identify the Clustering Structure (OPTICS) algorithm for cluster analysis. This algorithm is suitable for clustering heterogeneous data and can effectively handle situations with uneven data density. Specifically, a sample space is established with the acquisition time as the x-axis and the acquired data value as the y-axis. All updated data points from the current data source are mapped to this sample space. The reachability distance and core distance between data points are calculated using the OPTICS algorithm. Based on a preset clustering threshold (for example, the clustering threshold is set to 1.5 times the average distance of data points), data points with high similarity are divided into the same data cluster, ultimately resulting in multiple independent data clusters.

[0037] Step 201: Determine the governance difficulty of the data source based on the distribution characteristics of multiple data clusters.

[0038] The distribution characteristics of data clusters include the number of data clusters, the size of each data cluster, the degree of concentration of data within a cluster, and the degree of difference between data between clusters. For example, if a data source yields a large number of data clusters after clustering, and the size of each data cluster varies greatly (some clusters contain a large number of data points, while some clusters contain only a small number of data points), and the dispersion of data within a cluster is high, it indicates that the data distribution of this data source is chaotic, feature extraction is difficult, and the corresponding governance is difficult.

[0039] In one possible implementation, this application can determine the data distance between data points within each data cluster and the number of data points in each data cluster for each data cluster. Then, the governance difficulty of the data source is determined based on the data distance between data points within each data cluster and the number of data points in each data cluster.

[0040] Data distance, measured using methods such as Euclidean distance and Manhattan distance, characterizes the similarity between data points. The number of data points reflects the size of the data cluster. By analyzing the distance characteristics within a data cluster, the concentration of data can be assessed; by comparing the size differences between different data clusters, the imbalance in data distribution can be evaluated. Data sources with higher concentration and more significant imbalance are more difficult to manage.

[0041] For example, the difficulty of governance satisfies the following formula: in, For data source The difficulty of governance For data source The number of data clusters, For data source The The maximum value of the data distance between any two data points in a data cluster. For data source The The average data distance between any two data points in a data cluster. To find the maximum value function, The difference between the number of data points in any two clusters. The maximum value of the difference in the number of data points in any two clusters. For data source The The number of data points in each data cluster This represents the maximum number of data points across all clusters.

[0042] This ratio represents the average degree of unevenness within each data cluster. The larger the ratio, the more uneven the data distribution within the cluster, indicating a clear central clustering and peripheral dispersion. This ratio represents the relative difference in size between data clusters; a larger ratio indicates a more significant difference in size between data clusters. (Governance difficulty) This comprehensively reflects the complexity of the data source's internal structure. The larger the value, the more complex the internal data structure of the data source, the more prominent the contradiction between data dispersion and feature concentration, and the greater the difficulty of data governance. The smaller the value, the simpler the data structure, the more even the data distribution, and the easier the governance.

[0043] Based on the above technical solution, this application divides the data points of the data source into multiple data clusters through cluster analysis, and determines the governance difficulty based on the distribution characteristics of the data clusters. In this way, this application can quantify the inherent complexity of the data source from the perspective of data structure, and quantitatively express the contradiction between data dispersion and feature concentration through clustering results. This provides an accurate basic assessment indicator for subsequent dynamic governance, improving the objectivity and accuracy of governance difficulty assessment.

[0044] As a possible embodiment of this application, step 103 above can be implemented through the following steps: Step 301: For each data source, determine the range of increasing governance difficulty based on the governance difficulty corresponding to the data source in the current collection period and the governance difficulty corresponding to multiple historical collection periods.

[0045] The "increasing trend range of governance difficulty" is used to characterize the time span during which the governance difficulty of the data source has continuously increased up to the current collection period. Specifically, the "increasing trend range of governance difficulty" refers to the length of the period (in units of collection periods) during which the governance difficulty has shown a continuous upward trend from the historical collection period to the current collection period.

[0046] In one possible implementation, this application can determine the governance difficulty change sequence of each data source based on the governance difficulty corresponding to the current collection cycle and the governance difficulty corresponding to multiple historical collection cycles. Then, the upward trend range of governance difficulty corresponding to the data source can be determined based on the governance difficulty change sequence.

[0047] For example, if the current collection period is the nth period, and the governance difficulty of a certain data source is 1.5 in the n-3th period, 1.8 in the n-2nd period, 2.2 in the n-1st period, and 2.7 in the nth period, and the governance difficulty of these four periods all show a trend of increasing period by period, then the length of the upward trend interval of governance difficulty is 4; if the governance difficulty of the n-4th period is 2.0 (greater than 1.5 in the n-3rd period), then the n-4th period is not included in the upward trend interval.

[0048] In other words, the upward trend in governance difficulty must meet the condition of continuous growth, meaning the governance difficulty in the later period must be greater than that in the earlier period (i.e., as long as the governance difficulty in the later period is greater than that in the earlier period, it is considered an increase; if a stricter determination is required, the growth threshold can be set at 5% of the governance difficulty in the earlier period, i.e., the governance difficulty in the later period must be greater than 1.05 times the governance difficulty in the earlier period to be considered an increase). By tracing back through historical periods from the current period until the first period that does not meet the growth condition is found, the tracing stops. The number of consecutive periods that meet the condition during the tracing process is the length of the upward trend in governance difficulty.

[0049] Step 302: Determine the data source fluctuation coefficient for each data source in the current collection period based on the upward trend range of governance difficulty for each data source and the governance difficulty for each data source in the current collection period.

[0050] For example, the data source volatility coefficient satisfies the following formula: in, For data source The data source fluctuation coefficient corresponding to the current collection period This represents the minimum interval length of the increasing governance difficulty range corresponding to all data sources. For data source The length of the corresponding range of increasing governance difficulty. For data source Given the governance difficulty corresponding to the current data collection cycle. This represents the average governance difficulty across all data sources during the current collection period. This represents the standard deviation of the governance difficulty for all data sources during the current collection period.

[0051] Characterizing the rate of deterioration, when The smaller the time frame (i.e., the more rapidly the difficulty of governance increases in the short term), the larger this ratio, indicating that the abnormal fluctuations in the data source are more severe. The relative significance factor indicates that the higher the value, the more significant the governance difficulty of this data source is relative to the overall level. Multiplying the two values ​​yields the processing volatility coefficient. This comprehensively reflects the degree of dynamic anomalies in the data source. The larger the value, the more unstable the data source is, the faster the governance difficulty deteriorates and the further it deviates from the overall level, requiring greater attention to governance. The smaller the value, the more stable the data source status.

[0052] Based on the above technical solution, this application determines the range of increasing governance difficulty by analyzing the temporal trend of governance difficulty, and then determines the data source fluctuation coefficient by combining the range length and the relative level of governance difficulty. In this way, this application can timely capture the dynamic deterioration trend of data source governance difficulty, identify data sources that need to be focused on by quantifying the deterioration rate and relative significance, and provide accurate dynamic indicators for subsequent differentiated sampling and governance, thereby improving the timeliness and targeting of data governance.

[0053] As a possible embodiment of this application, step 104 above can be implemented through the following steps: Step 401: Determine the global fluctuation coefficient of multi-source heterogeneous data based on the data source fluctuation coefficients corresponding to multiple data sources in the current collection period.

[0054] The global volatility coefficient is used to characterize the overall volatility of the governance difficulty of all data sources. The larger the global volatility coefficient, the more drastic the overall fluctuation in the governance difficulty of all data sources, and the more complex the governance environment of the data foundation; conversely, the smaller the global volatility coefficient, the more stable the overall data state, and the lower the governance difficulty.

[0055] For example, the global volatility coefficient satisfies the following formula: in, The global fluctuation coefficient for multi-source heterogeneous data. The number of data sources in a multi-source heterogeneous dataset. For data source The data source fluctuation coefficient corresponding to the current collection period To find the maximum value function, This represents the maximum value of the data source fluctuation coefficient for all data sources in the current collection period.

[0056] The volatility coefficients of each data source are normalized to between 0 and 1. A cube operation is used to amplify the weight of larger values ​​and suppress the influence of smaller values, making the contribution of more volatile data sources to the global volatility coefficient more significant. The global volatility coefficient is obtained by averaging the weighted values ​​of all data sources. . The larger the value, the more drastic the overall fluctuation of multi-source heterogeneous data, the more unstable the data environment, and the more proactive governance measures are needed. The smaller the value, the more stable the overall data environment.

[0057] Step 402: For each data source, based on the global fluctuation coefficient of the multi-source heterogeneous data and the data source fluctuation coefficient corresponding to the current collection period, perform data fusion governance evaluation on each data point updated by the data source within the current collection period.

[0058] Data fusion governance assessment can combine the overall volatility status with the volatility characteristics of individual data sources to classify the reliability and importance of data points and carry out targeted governance operations. For example, when the global volatility coefficient is large, it indicates that the overall data is fluctuating drastically. In this case, data sources with larger absolute values ​​of volatility coefficients have a higher risk of data point anomalies, requiring the use of more stringent anomaly detection algorithms (such as increasing the sensitivity of anomaly thresholds) and increasing the frequency of data cleaning. For data sources with smaller absolute values ​​of volatility coefficients, their data points are relatively stable, and conventional governance strategies can be used to balance governance effectiveness and efficiency.

[0059] For example, data governance operations include: data cleaning (removing outlier data points), data completion (interpolating missing data points appropriately), feature extraction (prioritizing the extraction of core features from stable data sources), and data weight allocation (assigning higher weights to data points from data sources with less fluctuation for subsequent data fusion).

[0060] Based on the above technical solution, this application can determine the global fluctuation coefficient of multi-source heterogeneous data according to the data source fluctuation coefficients corresponding to the current collection period from multiple data sources. This achieves a precise grasp of the overall fluctuation state of multi-source heterogeneous data. Subsequently, for each data source, based on the global fluctuation coefficient of the multi-source heterogeneous data and the data source fluctuation coefficient corresponding to the current collection period, data fusion governance assessment is performed on each data point updated within the current collection period. This takes into account both overall fluctuation and individual differences, making the governance assessment more comprehensive and targeted. The above technical solution can effectively adapt to data sources with different fluctuation states, strengthen governance when overall data fluctuations are severe, and optimize governance efficiency when the data state is stable. This further enhances the flexibility and effectiveness of data governance, providing a stronger guarantee for the output of high-quality data assets from the data foundation.

[0061] As a possible embodiment of this application, step 402 above can be implemented through the following steps: Step 501: For each data source, determine the sampling weight of the data source in the current collection period based on the global fluctuation coefficient of the multi-source heterogeneous data and the data source fluctuation coefficient of the data source in the current collection period.

[0062] The sampling weight is used to characterize the importance of data points in the current collection period.

[0063] For example, the sampling weights satisfy the following formula: in, For data source In the current collection cycle The corresponding sampling weights, For data source In the current collection cycle The corresponding data source volatility coefficient, For data source In the previous collection cycle The corresponding data source volatility coefficient, For data source The variance of the data source fluctuation coefficient corresponding to each sampling period, For multi-source heterogeneous data in the current acquisition cycle The global volatility coefficient, For multi-source heterogeneous data in the previous acquisition cycle The global fluctuation coefficient.

[0064] This reflects the change in the degree of fluctuation of the data source between two adjacent periods. This reflects the historical variation of the volatility coefficient of the data source and is used to normalize current changes. It reflects the changing trends of the overall data environment. It can be positive or negative; the larger the absolute value, the more noteworthy the changes in the data source are in the current period.

[0065] Step 502: For each data source, perform data fusion governance evaluation on each data point updated by the data source in the current collection period according to the sampling weight corresponding to the data source in the current collection period.

[0066] In one possible implementation, this application can weight each data point updated by the data source in the current collection period according to the sampling weight corresponding to the data source in the current collection period, so as to obtain the attention of each data point.

[0067] Among them, attention level is used to characterize the degree of attention a data point receives in a multi-source heterogeneous data fusion governance scenario. The larger the absolute value of attention level, the more important the data point needs to be in the governance process. The positive or negative value of attention level reflects the distribution bias of the data point (positive value is higher, negative value is lower).

[0068] For example, attention levels satisfy the following formula: in, For data points attention, For data points The collected values, For data points Data source In the current collection cycle The corresponding sampling weights. This is a normalization function used to normalize the calculation results to a range between -1 and 1.

[0069] This application was first approved For the raw collected values Perform weighted adjustments when When the value is magnified, Reduce the value at time; then through The function linearly maps the adjusted values ​​to -1 to 1, making the monitoring values ​​of different dimensions comparable. The absolute value indicates the intensity of attention; the larger the value, the more attention is needed for the collected value. The sign indicates the bias of the collected value in the overall distribution, with negative values ​​indicating a bias towards the lower side and positive values ​​indicating a bias towards the higher side.

[0070] Subsequently, the data fusion governance assessment results are obtained based on the attention given to each data point updated by each data source within the current collection period.

[0071] In some embodiments, this application can map the attention of each data point updated by each data source in the current collection period to visual color information, and generate a data governance visualization view based on the visual color information.

[0072] The visualized color information satisfies the following mapping rules: the color of the visualized color information is used to represent the distribution bias of the attention of the corresponding data point, and the color depth of the visualized color information is used to represent the attention intensity of the corresponding data point.

[0073] For example, a red-blue false color mapping rule is used: when When the value is close to -1, it is mapped to dark blue, indicating a low level that requires close attention; when... When the value is close to 0, it is mapped to white or light green, indicating a normal or stable state; when... When the value approaches +1, it is mapped to dark red, indicating a high level requiring close monitoring. The color depth is determined by... The color tone was determined by The positive or negative nature of the business rules determines this.

[0074] The visualization view can be displayed on the data dashboard, allowing users to intuitively observe the attention status of data points from various data sources at different times: dark blue and dark red areas represent highly concerned data points that require urgent attention, light blue and light red areas represent moderately concerned data points that require regular attention, and white / light green areas represent data points with stable status.

[0075] Based on the above technical solution, this application calculates a global volatility coefficient to reflect the instability of the overall data environment, then combines the stability coefficients of each data source to determine the sampling weights, and finally obtains the attention level of each data point through weighted processing. In this way, this application can assess the relative importance of each data source within the context of the global data environment, dynamically adjust the sampling weights to give higher attention to data sources with drastic fluctuations and significant changes, and make data of different dimensions comparable through normalization processing. This achieves accurate, dynamic, and differentiated integrated governance assessment of multi-source heterogeneous data, improving the precision and effectiveness of data governance.

[0076] It should be noted that the various embodiments of this application can be referenced or learned from each other. For example, the same or similar steps, method embodiments, system embodiments and device embodiments can be referenced from each other without limitation.

[0077] This application also provides a hardware structure diagram of a multi-source heterogeneous data fusion and governance device based on a data foundation (hereinafter referred to as a multi-source heterogeneous data fusion and governance device 20 based on a data foundation), see [link to diagram]. Figure 2 The multi-source heterogeneous data fusion governance device 20 based on the data base includes a processor 21, and optionally, a memory 22 connected to the processor 21.

[0078] In the first possible implementation, see Figure 2 The multi-source heterogeneous data fusion and governance device 20 based on the data foundation also includes a communication interface 23. The processor 21, memory 22, and communication interface 23 are connected via a bus. The communication interface 23 is used to communicate with other devices or communication networks. Optionally, the communication interface 23 may include a transmitter and a receiver. The device in the communication interface 23 used to implement the receiving function can be considered as a receiver, which is used to perform the receiving steps in the embodiments of this application. The device in the communication interface 23 used to implement the transmitting function can be considered as a transmitter, which is used to perform the transmitting steps in the embodiments of this application.

[0079] Based on the first possible implementation method Figure 2 The structural diagram shown can be used to illustrate the structure of the multi-source heterogeneous data fusion governance device based on the data base involved in the above embodiments.

[0080] in, Figure 2The diagram can also illustrate the system chip in a multi-source heterogeneous data fusion and governance device based on a data platform. In this case, the actions performed by the aforementioned multi-source heterogeneous data fusion and governance device based on a data platform can be implemented by this system chip. The specific actions performed can be found above and will not be repeated here.

[0081] In implementation, each step of the method provided in this embodiment can be completed by integrated logic circuits in the processor or by instructions in software form. The steps of the method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.

[0082] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0083] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A multi-source heterogeneous data fusion and governance method based on a data foundation, characterized in that, include: At the end of the current collection period, each data point updated by multiple data sources in the multi-source heterogeneous data during the current collection period is collected. For each data source, the governance difficulty is assessed based on the data distribution of each data point updated by the data source within the current collection period, thus obtaining the governance difficulty of the data source in the current collection period; the governance difficulty is used to characterize the complexity of resolving the contradiction between data dispersion and feature concentration within the data source. For each data source, the data source fluctuation coefficient corresponding to the current collection period is determined based on the governance difficulty of the data source in the current collection period and the governance difficulty corresponding to multiple historical collection periods. The data source volatility coefficient is used to characterize the degree of volatility in the governance difficulty of the data source over historical time. Based on the data source fluctuation coefficients of the multiple data sources corresponding to the current collection period, a data fusion governance assessment is performed on each data point updated by the multiple data sources in the multi-source heterogeneous data within the current collection period.

2. The multi-source heterogeneous data fusion governance method based on a data foundation according to claim 1, characterized in that, For each data source, the governance difficulty is assessed based on the data distribution of each data point updated by the data source within the current collection period, resulting in the governance difficulty of the data source corresponding to the current collection period, including: For each data source, cluster analysis is performed on the data points updated by the data source within the current collection period to obtain multiple data clusters; The governance difficulty of the data source is determined based on the distribution characteristics of the multiple data clusters.

3. The multi-source heterogeneous data fusion governance method based on a data foundation according to claim 2, characterized in that, Based on the distribution characteristics of the multiple data clusters, the governance difficulty of the data source is determined, including: For each data cluster, determine the data distance between each data point within the data cluster, and determine the number of data points in each data cluster; The governance difficulty of the data source is determined based on the data distance between data points within each data cluster and the number of data points in each data cluster.

4. The multi-source heterogeneous data fusion governance method based on a data foundation according to claim 1, characterized in that, For each data source, based on the governance difficulty corresponding to the data source in the current collection period and the governance difficulty corresponding to multiple historical collection periods, the data source fluctuation coefficient corresponding to the data source in the current collection period is determined, including: For each data source, based on the governance difficulty of the data source in the current collection period and the governance difficulty in multiple historical collection periods, the upward trend range of governance difficulty for the data source is determined; the upward trend range of governance difficulty is used to characterize the time span of continuous increase in the governance difficulty of the data source up to the current collection period. The data source fluctuation coefficient for each data source in the current collection period is determined based on the range of increasing governance difficulty for each data source and the governance difficulty for each data source in the current collection period.

5. The multi-source heterogeneous data fusion governance method based on a data foundation according to claim 4, characterized in that, For each data source, based on the governance difficulty corresponding to the current collection period and the governance difficulty corresponding to multiple historical collection periods, the upward trend range of the governance difficulty corresponding to the data source is determined, including: For each data source, the governance difficulty change sequence of the data source is determined based on the governance difficulty corresponding to the current collection period and the governance difficulty corresponding to multiple historical collection periods. The range of increasing governance difficulty corresponding to the data source is determined based on the sequence of changes in governance difficulty.

6. The multi-source heterogeneous data fusion governance method based on a data foundation according to claim 1, characterized in that, Based on the data source fluctuation coefficients corresponding to the current acquisition period, a data fusion governance assessment is performed on each data point updated by multiple data sources within the current acquisition period in the multi-source heterogeneous data, including: Based on the data source fluctuation coefficients of the multiple data sources corresponding to the current collection period, the global fluctuation coefficient of the multi-source heterogeneous data is determined; the global fluctuation coefficient is used to characterize the overall fluctuation degree of the governance difficulty of all data sources. For each data source, based on the global fluctuation coefficient of the multi-source heterogeneous data and the data source fluctuation coefficient corresponding to the current acquisition period, a data fusion governance evaluation is performed on each data point updated by the data source within the current acquisition period.

7. The multi-source heterogeneous data fusion governance method based on a data foundation according to claim 6, characterized in that, For each data source, based on the global fluctuation coefficient of the multi-source heterogeneous data and the data source fluctuation coefficient corresponding to the current acquisition period, a data fusion governance evaluation is performed on each data point updated by the data source within the current acquisition period, including: For each data source, the sampling weight of the data source in the current acquisition period is determined based on the global fluctuation coefficient of the multi-source heterogeneous data and the data source fluctuation coefficient of the data source in the current acquisition period; the sampling weight is used to characterize the importance of the data points of the data source in the current acquisition period. For each data source, a data fusion governance evaluation is performed on each data point updated by the data source within the current collection period, based on the sampling weight corresponding to the data source in the current collection period.

8. The multi-source heterogeneous data fusion governance method based on a data foundation according to claim 7, characterized in that, For each data source, a data fusion governance evaluation is performed on each data point updated by the data source within the current collection period, based on the sampling weight corresponding to the data source in the current collection period, including: For each data source, the data points updated by the data source in the current collection period are weighted according to the sampling weight corresponding to the data source in the current collection period to obtain the attention level of each data point; the attention level is used to characterize the degree of attention a data point receives in the multi-source heterogeneous data fusion governance scenario; The data fusion governance assessment results are obtained based on the attention given to each data point updated by each data source within the current collection period.

9. The multi-source heterogeneous data fusion governance method based on a data foundation according to claim 8, characterized in that, The data fusion governance evaluation results are derived based on the level of attention given to each data point updated within the current collection period for each data source. These include: The attention of each data point updated by each data source within the current collection period is mapped to visual color information, and a data governance visualization view is generated based on the visual color information.

10. The multi-source heterogeneous data fusion governance method based on a data foundation according to claim 9, characterized in that, The visualized color information satisfies the following mapping rules: the color of the visualized color information is used to characterize the distribution bias of the attention of the corresponding data point, and the color depth of the visualized color information is used to characterize the attention intensity of the corresponding data point.