A multi-source data integration method based on data space

Through the multi-source data integration method of data space, combined with qualification scoring and verification records, multi-dimensional scoring is performed, which solves the problem of invalid data accumulation in multi-source data integration and realizes efficient data management and storage.

CN120578657BActive Publication Date: 2025-10-24福州城投新基建集团有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511081858.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-10-24
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

Existing multi-source data integration technologies lack a dynamic verification mechanism for effectiveness, resulting in the accumulation of a large amount of invalid data in the system, occupying storage space and interfering with data processing efficiency.

Method used

Through the multi-source data integration method of the data space, combined with the qualification score and verification record of the data source, frequent analysis is carried out, the same type of difference analysis and accuracy scoring of the data are carried out, a multi-dimensional scoring mechanism is established, high-quality data is screened and allocated to high-priority data space.

Benefits of technology

Ensure the accuracy and credibility of data, accurately identify and clean up invalid data, rationally utilize storage resources, meet the application priorities of different data in the business, and improve the accuracy and management efficiency of data storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578657B_ABST
    Figure CN120578657B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of multi-source data integration, and relates to a multi-source data integration method based on a data space. The method comprises the following steps: S1, obtaining a data source for storing data, and performing retention verification in combination with the stored data in the data source, and setting a verification time for the stored data in the data source according to a verification record, so that the stored data is divided into passed verification time and failed verification time; and S2, obtaining qualification data of the data source, and performing initial qualification scoring on the data source according to the qualification data; the application controls data quality basis from the data source by performing qualification scoring on the data source and combining frequent analysis of the verification record, and simultaneously evaluates data accuracy scoring by analyzing and evaluating differences between the data source and standard data sources and distributed data sources, so that data source comprehensive scoring is obtained by comprehensively scoring the qualification scoring and the data accuracy scoring, and the accuracy of the data is further ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multi-source data integration, in particular to a multi-source data integration method based on data space. BACKGROUND

[0002] In the data-driven information age, multi-source data integration is a key technology in the field of data processing, and its core role is to integrate information from different data sources, realize centralized management and efficient use of data resources;

[0003] Existing multi-source data integration technology is mostly applied to enterprise data middle platform, industry data sharing platform and other scenarios, and gathers data from various data sources through database connection, interface calling and other ways, lacks dynamic verification mechanism for data effectiveness, and is easy to accumulate a large amount of invalid data in the system, which not only occupies storage space, but also interferes with data processing efficiency, therefore, a multi-source data integration method based on data space is proposed. SUMMARY

[0004] The purpose of the present application is to provide a multi-source data integration method based on data space to solve the problems raised in the background art.

[0005] To achieve the above purpose, a multi-source data integration method based on data space is provided, comprising the following steps:

[0006] S1, obtain a data source for storing data, and perform retention verification in combination with the stored data in the data source, and set the verification time of the stored data in the data source according to the verification record, and divide the stored data into passed verification time and failed verification time;

[0007] S2, obtain the qualification data of the data source, score the initial qualification of the data source according to the qualification data, then analyze the data source frequently according to the verification record, and evaluate the analysis result in combination with the initial qualification score to obtain the final qualification score;

[0008] S3, select a standard data source, perform same-type difference analysis on the stored data of the data source according to the stored data in the standard data source, select a distribution data source according to the difference analysis result and the final qualification score, and perform same-type difference analysis again, then score the accuracy of the stored data according to the difference analysis result of the data source;

[0009] S4, combine the accuracy score of the stored data with the final qualification score to score the data source comprehensively, establish a corresponding data space for different comprehensive scores, and then assign a data space to the stored data according to the comprehensive score of the data source;

[0010] S5, correlating the storage data of the data source that does not pass the verification time with the storage data of the standard data source, when the storage data of the data source that does not pass the verification time is not correlated with the storage data of the standard data source, combining the storage data of the distribution data source with the storage data of the data source that does not pass the verification time to perform time period score evaluation, and allocating data space according to the time period score.

[0011] As a further improvement of the technical solution, the S1 extracts the storage data collected by the system database by establishing a connection with the system database, and determines the data source of the storage data according to the collection location of the storage data.

[0012] As a further improvement of the technical solution, the S1 has the following steps:

[0013] S1.1, obtaining the storage data of the data source, combining the storage data of the system database with the storage data of the data source to perform retention verification, when the system database does not match the same storage data in the data source, deleting the unmatched storage data in the system database, otherwise, if the same storage data is matched, continuing to monitor;

[0014] S1.2, calculating the difference time between the publishing time and the deletion time of the deleted storage data in the data source, taking the difference time as the retention time of the deleted storage data in the data source, and then extracting the longest retention time of the data source as the verification time of the data source.

[0015] As a further improvement of the technical solution, the S1 further includes the following steps:

[0016] S1.3, after setting the verification time of each data source in S1.2, when new storage data is collected through the data source, comparing the retention time of the storage data with the verification time;

[0017] S1.4, when the retention time of the storage data is less than the verification time, saving the storage data in the data space by using the time period score of S5;

[0018] When the storage data does not pass the verification time, the data space is saved according to the time period score;

[0019] S1.5, when the retention time of the storage data is greater than the verification time, saving the storage data in the data space by using the comprehensive score of the data source of S4;

[0020] After the storage data passes the verification time, the comprehensive score of the data source to which the storage data belongs is extracted, and the data space is saved according to the extracted comprehensive score.

[0021] As a further improvement of the technical solution, the S2 has the following steps:

[0022] S2.1, obtain qualification data of the data source, perform initial qualification scoring on the data source according to the qualification data, and obtain an initial qualification score of the data source;

[0023] S2.2, obtain a retention verification record of the data source, perform frequent analysis according to the retention verification record, then evaluate the frequent analysis result in combination with the initial qualification score, and obtain a final qualification score of the data source;

[0024] The retention verification record includes a verification time and a deletion record of the stored data of the data source;

[0025] The longer the verification time in the verification record is and the more the deletion records are, the greater the influence of the frequent analysis result on the initial qualification score is.

[0026] As a further improvement of the technical solution, the steps of S3 are as follows:

[0027] S3.1, select a standard data source, perform same-type difference analysis on the stored data of other data sources according to the stored data corresponding to the standard data source, and obtain a difference analysis result of the same-type data of the data source and the standard data source;

[0028] S3.2, set a qualification selection threshold, then select a data source with a final qualification score greater than the qualification selection threshold and having same-type data with the standard data source as a distribution data source according to the qualification selection threshold;

[0029] S3.3, perform same-type difference analysis on the stored data of the data source again according to the stored data in the distribution data source, and obtain a difference analysis result of the same-type data of the data source and the distribution data source;

[0030] S3.4, evaluate the data accuracy score of the data source by combining the difference analysis results corresponding to S3.1 and S3.3, and obtain an accuracy score corresponding to each data source.

[0031] As a further improvement of the technical solution, in the process of evaluating the accuracy score, the more similar the stored data of a data source is to the stored data of the standard data source and the distribution data source, the smaller the influence of the difference analysis result is, and the higher the corresponding data accuracy score is;

[0032] The less similar the stored data of a data source is to the stored data of the standard data source and the distribution data source, the greater the influence of the difference analysis result is, and the lower the corresponding data accuracy score is.

[0033] As a further improvement of the technical solution, the steps of S4 are as follows:

[0034] S4.1, combine the accurate score of the data source with the qualification score to obtain a comprehensive score;

[0035] S4.2, establish a plurality of data spaces, each data space corresponding to a different comprehensive score, and the higher the comprehensive score, the higher the priority of the data space;

[0036] S4.3, distribute the stored data of each data source to the corresponding data space according to the comprehensive score of the data source.

[0037] As a further improvement of the technical solution, the steps of S5 are as follows:

[0038] S5.1, associate the stored data with a verification time less than the retention time with the stored data of the standard data source, and when the stored data has an association with the standard data source, evaluate the time period score according to the association analysis result, otherwise, when the stored data does not have an association with the standard data source, go to S5.2;

[0039] S5.2, associate the stored data with a verification time less than the retention time with the stored data of the distributed data source, and evaluate the time period score according to the association analysis result to determine the time period score of the stored data.

[0040] Compared with the prior art, the beneficial effects of the present application are:

[0041] 1. In the multi-source data integration method based on data space, the data quality is controlled from the data source by giving a qualification score to the data source and combining the frequent analysis of the verification record to obtain a final qualification score, and the data accuracy score is evaluated by difference analysis with the standard data source and the distributed data source, and the data source comprehensive score is obtained by combining the qualification score and the data accuracy score, which further ensures the accuracy of the data. This multi-dimensional scoring mechanism can filter out high-quality data and allocate them to high-priority data spaces, effectively ensuring the reliability and credibility of the data in the system, and providing a high-quality data basis for subsequent data application.

[0042] 2. In the multi-source data integration method based on data space, the retention verification mechanism of S1 can accurately identify the stored data in the system database that no longer exists in the data source and clean it up in time, avoiding invalid data occupying storage space. At the same time, the verification time is set based on the retention time of the deleted data in the data source, so that the data retention strategy is more in line with the actual situation of the data source. When new data is collected, the appropriate storage method is selected by combining the verification time and the retention time, ensuring that the stored data is valid and meets the retention rules, greatly improving the accuracy and effectiveness of data storage.

[0043] 3、The multi-source data integration method based on data space, by establishing different priority data spaces according to the comprehensive score of the data source, the data is allocated to the corresponding space according to the score, which ensures the priority storage and use of high-value data, for the data with shorter retention time, the time period score evaluation is carried out through S5 step combined with the standard data source and the distributed data source, and the evaluation result is stored, so that the space allocation of short-term data is more in line with its actual value and business scenario demand, this flexible allocation method can not only reasonably utilize the storage resources, but also meet the application priority of different data in business, and improves the efficiency and rationality of data space management. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 The overall flowchart of the present application is shown in the figure;

[0045] Figure 2 The flowchart of obtaining existing data of the data source of the present application is shown in the figure;

[0046] Figure 3 The flowchart of obtaining qualification data of the data source of the present application is shown in the figure;

[0047] Figure 4 The flowchart of selecting the standard data source of the present application is shown in the figure;

[0048] Figure 5 The flowchart of accurately scoring the data of the data source and combining the qualification score for comprehensive scoring of the present application is shown in the figure;

[0049] Figure 6 The flowchart of determining the time period score of the stored data of the present application is shown in the figure. DETAILED DESCRIPTION

[0050] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0051] Please refer to Figures 1-6 The present embodiment aims to provide a multi-source data integration method based on data space, comprising the following steps:

[0052] S1, obtaining a data source for storing data, and combining the stored data in the data source to perform retention verification, and setting the verification time of the stored data in the data source according to the verification record, and dividing the stored data into passed verification time and failed verification time;

[0053] S1 establishes a connection with the system database to extract the storage data collected by the system database, and determines the data source of the storage data according to the collection site of the storage data.

[0054] A connection with the system database is established through a database interface provided by the system, an SQL query statement is executed, and storage data in a specified time period is extracted from the database, including data records, collection timestamps, and metadata.

[0055] Collection site information is parsed from the extracted storage data to establish a mapping relationship between the collection site and the data source, and each collection site is associated with the corresponding data source through a preset configuration table or a rule engine.

[0056] The steps of S1 are as follows:

[0057] S1.1, obtain the storage data of the data source, combine the storage data of the system database with the storage data of the data source for retention verification, when the system database does not match the same storage data in the data source, delete the storage data in the system database that does not match, otherwise, if the same storage data is matched, continue to monitor, the specific steps are as follows:

[0058] Establish a data connection: establish a real-time or timed connection with each data source (external database) through a system interface, obtain the current state and data update timestamp of the data source, then extract the storage data to be verified from the system database, and synchronously obtain the existing data of the corresponding data source;

[0059] Data matching and verification: take the data unique identifier as the key, match the system database data with the existing data of the data source, for each storage data, check whether it exists in the data source;

[0060] If it does not exist (matching fails), it is marked as “to be deleted”;

[0061] If it exists (matching succeeds), it is marked as “valid” and continues to be monitored;

[0062] Data cleaning and synchronization: batch delete the data records in the system database marked as “to be deleted”, and synchronously record the timestamp and deletion reason (the data source no longer exists) of the deletion operation.

[0063] S1.2, calculate the difference time of the publishing time and the deletion time of the deleted storage data in the data source, take the difference time as the retention time of the deleted storage data in the data source, and then extract the longest retention time of the data source as the verification time of the data source.

[0064] For each deleted data, calculate its retention duration in the data source: the time interval from the publishing time to the deletion time;

[0065] Grouping the deletion data according to the data source ID, each group corresponds to an independent data source, and the longest retention time is selected as the verification time threshold for each data source, and the verification time is stored in the data source metadata for subsequent data retention strategy decision.

[0066] S1 further comprises the following steps:

[0067] S1.3, after setting the verification time for each data source in S1.2, when new storage data is collected through the data source, the retention time of the storage data is compared with the verification time;

[0068] S1.4, when the retention time of the storage data is less than the verification time, the storage data is saved in the data space using the time period score of S5;

[0069] When the storage data does not pass the verification time, the data space is saved according to the time period score;

[0070] S1.5, when the retention time of the storage data is greater than the verification time, the storage data is saved in the data space using the comprehensive score of the data source of S4;

[0071] After the storage data passes the verification time, the comprehensive score of the data source to which the storage data belongs is extracted, and the data space is saved according to the extracted comprehensive score.

[0072] After the data is saved, the data state is continuously monitored, and if the subsequent data source, data itself, etc. changes (such as data source verification time adjustment, data retention time change), the comparison and data space adjustment can be performed again.

[0073] S2, obtaining the qualification data of the data source, performing initial qualification scoring on the data source according to the qualification data, then performing frequent analysis on the verification record of the data source, and evaluating the analysis result combined with the initial qualification score to obtain the final qualification score;

[0074] The steps of S2 are as follows:

[0075] S2.1, obtaining the qualification data of the data source, performing initial qualification scoring on the data source according to the qualification data, and obtaining the initial qualification score of the data source;

[0076] The basic qualification data of the data source is obtained from the data source registration information, the management system or the third party certification agency, including the authority of the data provider (such as industry association, commercial company), data update frequency (real-time, daily, weekly, etc.), data integrity (field coverage, historical data retention time), data accuracy history record (comparison result with standard data set), compliance certification (such as ISO27001 information security certification);

[0077] The collected qualification data is quantitatively processed, and a preset scoring standard (such as a weight table) is used to calculate the initial qualification score of each data source.

[0078] S2.2, obtaining the retention verification record of the data source, and performing frequent analysis according to the retention verification record, and then evaluating the frequent analysis result in combination with the initial qualification score, so as to obtain the final qualification score of the data source;

[0079] The retention verification record includes the verification time and the deletion record of the storage data of the data source;

[0080] When the verification time in the verification record is longer and the deletion record is more, the influence of the frequent analysis result on the initial qualification score is greater, and the formula is as follows:

[0081] ;

[0082] Wherein, T avgverify is the average verification time interval, m is the verification times, T verify,j is the timestamp of the jth verification;

[0083] ;

[0084] Wherein, F adjust is a score adjustment factor (the greater the value, the more the score is reduced), a is an adjustment coefficient (configurable, such as 0.3), T is the average verification time of all data sources, β is a deletion record influence coefficient, R delete is the deletion record rate (calculated by combining the deleted storage data with the total storage data to calculate the proportion);

[0085] ;

[0086] Wherein, S final is the final qualification score.

[0087] S3, selecting a standard data source, performing same-type difference analysis on the storage data of the data source according to the storage data in the standard data source, simultaneously selecting a distribution data source according to the difference analysis result and the final qualification score and performing same-type difference analysis again, and then performing accuracy scoring on the storage data according to the difference analysis result of the data source;

[0088] The steps of S3 are as follows:

[0089] S3.1, selecting a standard data source, performing same-type difference analysis on the storage data corresponding to other data sources according to the storage data corresponding to the standard data source, and obtaining the difference analysis result of the same-type data of the data source and the standard data source;

[0090] According to business needs or industry standards, select authoritative, stable and high credibility data sources as standard data sources, or managers can establish a data source as a standard data source, and then store the data into the storage data through manual audit authority, so as to improve the accuracy of all data;

[0091] The stored data of other data sources is classified by type, and compared with the same type data of the standard data source one by one, and the difference (such as numerical deviation, field missing, format inconsistency, etc.) is analyzed, and the difference degree and difference type are recorded, and the formula is as follows:

[0092] ;

[0093] Among them, D std (i, l) is the difference degree of the ith data source and the lth data in the standard data source (the smaller the value, the smaller the difference), V i,l,k is the kth record value of the lth data in the ith data source, V std,l,k is the kth record value of the lth data in the standard data source, and n is the total number of records of the lth data;

[0094] S3.2, set the qualification selection threshold, and then select the data source with final qualification score greater than the qualification selection threshold and with the same type data as the standard data source as the distribution data source;

[0095] Set the qualification selection threshold (such as 80 points, which can be adjusted according to business scenarios), and filter out the data sources with final qualification score higher than the threshold, and further select the data sources containing the same type data as the standard data source from the filtered high-quality data sources, and determine them as distribution data sources (used to assist verification of data accuracy).

[0096] S3.3, according to the stored data in the distribution data source, the stored data in the data source is analyzed again, and the difference analysis result of the data source and the distribution data source is obtained; according to the same type data of the distribution data source, the difference analysis of the stored data of other data sources is carried out again, and the difference degree with the distribution data source is recorded, which is consistent with the method of S3.1;

[0097] S3.4, the data source combines the difference analysis results of S3.1 and S3.3 to evaluate the data accuracy score, and obtains the corresponding accuracy score of each data source.

[0098] Periodically (such as every week), repeat the above analysis process, and when the data of the standard data source or the distribution data source is updated, trigger reevaluation to ensure the timeliness of the data accuracy score.

[0099] S3.4 In the process of evaluating the accuracy of data scoring, the results of the difference analysis are converted into a calculable impact value: the greater the difference, the higher the impact value; the smaller the difference, the lower the impact value;

[0100] The difference impact value of a perfect match is 0, and the difference impact value of a perfect mismatch is 1;

[0101] At the same time, some data sources with known accuracy are selected for scoring verification.

[0102] The more similar the stored data of a data source is to the stored data of the standard data source and the distributed data source, the smaller the impact of the difference analysis results will be, and the higher the corresponding data accuracy score will be;

[0103] The less similar the stored data of a data source is to the stored data of the standard data source and the distributed data source, the greater the impact of the difference analysis results and the lower the corresponding data accuracy score.

[0104] S4. Combine the accurate score of the stored data with the final qualification score to perform a comprehensive score on the data source, establish corresponding data spaces for different comprehensive scores, and then allocate data spaces for the stored data based on the comprehensive score of the data source;

[0105] The steps for S4 are as follows:

[0106] S4.1. Combine the data source's accuracy score with the qualification score to create a comprehensive score;

[0107] The weights of the two scores are set according to business needs (for example, the data accuracy score accounts for 60% and the qualification score accounts for 40%), and the comprehensive score of each data source is obtained through weighted calculation.

[0108] S4.2. Establish multiple data spaces, each corresponding to a different comprehensive score. The data space with a higher comprehensive score has a higher priority.

[0109] Design multiple data spaces, divide them into levels according to the comprehensive score range (for example, high-priority space corresponds to 90-100 points, medium-priority space corresponds to 60-89 points, and low-priority space corresponds to 0-59 points), and configure storage policies for each data space.

[0110] S4.3. Allocate the stored data of each data source to the corresponding data space for storage according to the comprehensive score of the data source.

[0111] For the stored data of each data source, the corresponding scoring interval is matched according to its comprehensive score, and the data is allocated to the data space of the corresponding level. The comprehensive score of the data source is recalculated regularly (such as every quarter). If the score change causes the interval to change, the data will be migrated to the new corresponding data space.

[0112] S5, associate the storage data of the unverified time with the storage data of the standard data source for analysis, when the storage data of the unverified time is not associated with the storage data of the standard data source, combine the storage data of the distribution data source with the storage data of the unverified time for time period score evaluation, and allocate data space according to the time period score.

[0113] The steps of S5 are as follows:

[0114] S5.1, associate the storage data of the retention time less than the verification time with the storage data of the standard data source for analysis, when the storage data is associated with the standard data source, evaluate the time period score according to the association analysis result, otherwise, when the storage data is not associated with the standard data source, go to S5.2;

[0115] S5.2, associate the storage data of the retention time less than the verification time with the storage data of the distribution data source for analysis, evaluate the time period score according to the association analysis result, and determine the time period score of the storage data, the specific steps are as follows:

[0116] Standard data source association analysis: extract all storage data of the retention time less than the verification time (i.e. short-term data) from the system, for each short-term storage data, compare it with the data of the standard data source based on the preset association rule (such as timestamp matching, business primary key association, data feature similarity), and judge whether there is an association: if there is at least one standard data satisfying the association rule, it is considered to have an association; otherwise, it is considered to have no association;

[0117] Time period score evaluation based on standard data source: for the short-term data associated with the standard data source, calculate the initial time period score according to the association strength (such as the number of matching fields, similarity score), and modify the initial score according to the relationship between data collection time and business peak / valley period (such as assigning higher score to data in peak period);

[0118] Distribution data source association analysis (when there is no standard association): for short-term data without association with the standard data source, sequentially associate with the data of each distribution data source, and select the distribution data source with the highest association strength as the reference for time period score evaluation;

[0119] Time period score evaluation based on distribution data source: calculate the basic score according to the association result with the distribution data source, and consider the authority weight of the distribution data source (the higher the qualification score, the greater the weight), and also modify it according to the collection period to obtain the final time period score;

[0120] Result application and storage: store the time period score of each short-term data into the metadata for subsequent data space allocation.

[0121] The above shows and describes the basic principles, main features and advantages of the present application. Those skilled in the art should understand that the present application is not limited to the above-mentioned embodiments, and the above-mentioned embodiments and descriptions in the specification are only preferred examples of the present application and are not intended to limit the present application. Various changes and improvements can be made to the present application without departing from the spirit and scope of the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.

Claims

1. A method for multi-source data integration based on data space, characterized in that: The method comprises the following steps: S1, obtaining a data source for storing data, and performing retention verification on the stored data in the data source, and setting the verification time of the stored data in the data source according to the verification record, and dividing the stored data into passed verification time and failed verification time; The steps of S1 are as follows: S1.1, obtaining the stored data of the data source, combining the stored data of the data source with the stored data of the system database for retention verification, and deleting the unmatched stored data in the system database when the same stored data is not matched in the data source, otherwise, if the same stored data is matched, the monitoring is continued; S1.2, calculating the difference time of the publishing time and the deletion time of the deleted stored data in the data source, taking the difference time as the retention time of the deleted stored data in the data source, and then setting the longest retention time of the data source as the verification time of the data source; The S1 further comprises the following steps: S1.3, after setting the verification time of each data source in S1.2, when new stored data is collected through the data source, the retention time of the stored data is compared with the verification time; S1.4, when the retention time of the stored data is less than the verification time, the stored data is saved in the data space by using the time period score of S5; When the stored data fails to pass the verification time, the data space is saved according to the time period score; S1.5, when the retention time of the stored data is greater than the verification time, the stored data is saved in the data space by using the comprehensive score of the data source of S4; After the stored data passes the verification time, the comprehensive score of the data source to which the stored data belongs is extracted, and the data space is saved according to the extracted comprehensive score; S2, obtaining the qualification data of the data source, performing initial qualification scoring on the data source according to the qualification data, then performing frequent analysis on the verification record of the data source, and evaluating the analysis result combined with the initial qualification score to obtain the final qualification score; The steps of S2 are as follows: S2.1, obtaining the qualification data of the data source, performing initial qualification scoring on the data source according to the qualification data, and obtaining the initial qualification score of the data source; S2.2, obtaining the retention verification record of the data source, performing frequent analysis according to the retention verification record, and then evaluating the frequent analysis result combined with the initial qualification score to obtain the final qualification score of the data source; The retention verification record comprises the verification time and the deletion record of the stored data of the data source; The longer the verification time in the verification record is, the more the deletion record is, and the greater the influence of the frequent analysis result on the initial qualification score is; S3, selecting a standard data source, performing same-type difference analysis on the stored data of the data source according to the stored data in the standard data source, selecting a distribution data source according to the difference analysis result and the final qualification score, and performing same-type difference analysis again, and then performing accuracy scoring on the stored data according to the difference analysis result of the data source; The steps of S3 are as follows: S3.1, select a standard data source, and perform same-type difference analysis on the storage data corresponding to other data sources according to the storage data corresponding to the standard data source, to obtain difference analysis results of the same-type data of the data source and the standard data source; S3.2, set a qualification selection threshold, and then select a data source with a final qualification score greater than the qualification selection threshold and having same-type data with the standard data source as a distribution data source according to the qualification selection threshold; S3.3, perform same-type difference analysis on the storage data in the data source again according to the storage data in the distribution data source, to obtain difference analysis results of the same-type data of the data source and the distribution data source; S3.4, the data source evaluates the data accuracy score by combining the difference analysis results corresponding to S3.1 and S3.3, to obtain an accuracy score corresponding to each data source; S4, combine the accuracy score of the storage data with the final qualification score of the data source to perform comprehensive scoring, establish a corresponding data space for different comprehensive scores, and then assign a data space to the storage data according to the comprehensive score of the data source; S5, perform correlation analysis on the storage data with a verification time in the data source and the storage data of the standard data source, when the storage data with the verification time is not correlated with the storage data of the standard data source, perform time period score evaluation on the storage data of the distribution data source combined with the storage data with the verification time, and assign a data space according to the time period score. 2.The method of claim 1, wherein: The S1 extracts the storage data collected by the system database by establishing a connection with the system database, and determines the data source of the storage data according to the collection location of the storage data. 3.The method of claim 1, wherein: In the process of evaluating the accuracy score, the more similar the results between the storage data of a data source and the storage data of the standard data source and the distribution data source, the smaller the influence of the difference analysis result, and the higher the corresponding data accuracy score; The more dissimilar the results between the storage data of a data source and the storage data of the standard data source and the distribution data source, the greater the influence of the difference analysis result, and the lower the corresponding data accuracy score.

4. The method of claim 1, wherein: The steps of S4 are as follows: S4.1, combine the accuracy score of the data source with the qualification score to perform comprehensive scoring; S4.2, establish a plurality of data spaces, each data space corresponding to a different comprehensive score, and the higher the corresponding comprehensive score, the higher the priority; S4.3, store the storage data of each data source in the corresponding data space according to the comprehensive score of the data source.

5. The method of claim 1, wherein: The steps of S5 are as follows: S5.1, perform correlation analysis on the storage data with a verification time and the storage data of the standard data source, when the storage data is correlated with the standard data source, perform time period score evaluation according to the correlation analysis result, otherwise, when the storage data is not correlated with the standard data source, enter S5.2; S5.2, perform correlation analysis on the storage data with a verification time and the storage data of the distribution data source, and perform time period score evaluation according to the correlation analysis result to determine the time period score of the storage data.

Citation Information

Patent Citations

  • Multi-source heterogeneous data integration method and system of data integration middleware

    CN118964469A

  • Data processing method and processing equipment thereof

    CN119938653A