Integrated data platform management method based on cloud computing
By standardizing the format of diverse data and detecting anomalies, combined with a manual review process, the problem of accidental or missed data deletion during data processing was solved, improving the accuracy and reliability of data management and enhancing data processing efficiency in a cloud computing environment.
Patent Information
- Application Number
- CN202511097299.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-21
AI Technical Summary
In data processing and analysis, traditional methods suffer from data congestion, which may lead to the accidental deletion or omission of correct data during rapid verification, affecting the reliability of integrated data.
The cloud computing platform performs unified format processing on multi-dimensional data, obtains anomaly information and evaluates the anomaly coefficient of multi-dimensional data, determines whether to trigger rapid verification, performs secondary verification on the first abnormal data and enters the manual review process to ensure data accuracy.
It significantly improved the accuracy and reliability of data management, solved the problem of accidental or missed data deletion, and improved the overall quality and efficiency of data processing.
Smart Images

Figure CN120994458A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data management, and more particularly to a cloud computing-based integrated data platform management method. BACKGROUND
[0002] With the deepening of digital transformation, the demand for data management in various industries has shown explosive growth. Traditional data management methods face serious problems such as data island, low processing efficiency, and insufficient resource utilization. In emerging fields such as intelligent manufacturing and smart cities, real-time collection, storage, and analysis of integrated data have become a key challenge. The maturity of cloud computing technology provides a new way to solve these problems. Cloud computing integrated data management platform realizes the unified management of integrated data through resource pooling, elastic expansion, and intelligent scheduling.
[0003] However, in the process of implementing the technical scheme of the embodiments of the present application, it is found that the above-mentioned technology at least has the following technical problems: In data processing and analysis, there may be a situation of triggering rapid verification due to data congestion, which may skip the verification step and mistakenly delete or miss correct data, affecting the reliability of integrated data. SUMMARY
[0004] In order to overcome the above-mentioned defects of the prior art, the present application provides a cloud computing-based integrated data platform management method to solve the problems existing in the background art.
[0005] To achieve the above object, the present application provides the following technical scheme:
[0006] A cloud computing-based integrated data platform management method, the system comprises: collecting multi-element data through a cloud computing platform, performing format unification processing on the multi-element data, and obtaining multi-element data in a unified format;
[0007] The unified format multi-element data is preprocessed to obtain preprocessed multi-element data, the abnormal information of the preprocessed multi-element data is obtained, the abnormal information includes field information, source abnormal information and time distribution information, and the multi-element data abnormal coefficient is obtained according to the abnormal information; Determine whether to trigger rapid verification according to the multi-element data abnormal coefficient; If it is determined to trigger rapid verification, the data marked as "abnormal" in rapid verification is obtained, which is recorded as the first abnormal data; The first abnormal data is verified again, and if the second verification result is normal data, the data is temporarily stored in the "suspicious data buffer area", and the manual review process is triggered; The second determination abnormal data is manually reviewed, the deterministic processing is executed, and the actual effective data set is obtained; According to the actual effective data set, the data management operation is executed in the cloud storage platform.
[0008] Preferably, the multi-source data anomaly coefficient obtaining step is: obtaining field information of the multi-source data, the field information including missing field number, mandatory field number, valid associated record number, and field value number meeting value range requirement, and obtaining data integrity according to the field information; obtaining source abnormal information of the multi-source data, the source abnormal information including historical abnormal rate, equipment health rate, transmission interruption rate, and same-source data conflict rate, and obtaining data source abnormal rate according to the source abnormal information; obtaining time distribution information of the multi-source data, the time distribution information including timestamp outlier index, arrival interval abnormal rate, periodic pattern deviation rate, time density abnormal rate, and time sequence conflict rate, and obtaining time distribution deviation rate according to the time distribution information; and performing normalization processing on the data integrity, the data source abnormal rate, and the time distribution deviation rate, and calculating the multi-source data anomaly coefficient according to the normalized data integrity, the normalized data source abnormal rate, and the normalized time distribution deviation rate.
[0009] Preferably, the data integrity obtaining step is: obtaining missing field number and total field number of the multi-source data, performing ratio calculation on the missing field number and the total field number to obtain a missing rate; obtaining filled mandatory field number of the multi-source data, performing ratio calculation on the filled mandatory field number and the total field number to obtain a mandatory rate; obtaining valid associated record of the multi-source data, performing ratio calculation on the valid associated record number and total associated record number to obtain an associated integrity rate; obtaining field value number meeting value range of the multi-source data, performing ratio calculation on the field value number meeting value range and total field value number to obtain a value range integrity rate; and performing normalization processing on the missing rate, the mandatory rate, the associated integrity rate, and the value range integrity rate, and calculating the data integrity according to the normalized missing rate, the normalized mandatory rate, the normalized associated integrity rate, and the normalized value range integrity rate.
[0010] Preferably, the data source abnormal rate obtaining step is: obtaining historical abnormal data number and historical total data number of the multi-source data, performing ratio calculation on the historical abnormal data number and the historical total data number to obtain a historical abnormal rate; obtaining past thirty-day fault number and total running days of the multi-source data, performing ratio calculation on the past thirty-day fault number and the total running days to obtain an equipment health rate; obtaining transmission interruption number and expected transmission number of the multi-source data in the last 24 hours, performing ratio calculation on the transmission interruption number and the expected transmission number to obtain a transmission interruption rate; obtaining logical conflict field number and total comparable character number of the multi-source data, performing ratio calculation on the logical conflict field number and the total comparable character number to obtain a same-source data conflict rate; and performing normalization processing on the historical abnormal rate, the equipment health rate, the transmission interruption rate, and the same-source data conflict rate, and calculating the data source abnormal rate according to the normalized historical abnormal rate, the normalized equipment health rate, the normalized transmission interruption rate, and the normalized same-source data conflict rate.
[0011] Preferably, the step of obtaining the time distribution deviation rate comprises: obtaining the actual arrival time, the expected time and the allowable time deviation threshold of each data in the multi-element data, and calculating the timestamp outlier index of each data by the actual arrival time, the expected time and the allowable time deviation threshold; obtaining the current period sequence and the reference period sequence, and calculating the minimum cumulative distance of the optimal matching path of the current period sequence and the reference period sequence to obtain the dynamic time warping similarity; calculating the ratio of the dynamic time warping similarity and the reference period similarity according to the dynamic time warping similarity and the reference period similarity to obtain the period pattern deviation rate; obtaining the data amount in the current time window of the multi-element data, the mean of the historical same period window data amount and the standard deviation of the historical same period window data amount, and obtaining the time density anomaly rate by calculating the data amount in the current time window of the multi-element data, the mean of the historical same period window data amount and the standard deviation of the historical same period window data amount; and performing normalization processing on the timestamp outlier index, the period pattern deviation rate and the time density anomaly rate, and calculating the time distribution deviation rate by the normalized timestamp outlier index, the period pattern deviation rate and the time density anomaly rate.
[0012] Preferably, the step of performing secondary verification on the first abnormal data comprises: obtaining the credibility data of the first abnormal data, the credibility data comprising time distribution data, data source information and context information; obtaining the time distribution consistency coefficient by evaluating the time distribution data of the first abnormal data; obtaining the data source credibility coefficient by evaluating the data source information of the first abnormal data; obtaining the context similarity coefficient by evaluating the context information of the first abnormal data; performing normalization processing on the time distribution consistency coefficient, the data source credibility coefficient and the context similarity coefficient, and calculating the second abnormal data credibility index by the normalized time distribution consistency coefficient, the data source credibility coefficient and the context similarity coefficient; in the last stage of the data processing flow, the system calculates the second abnormal data credibility index by evaluating the data source credibility coefficient, the context similarity coefficient and the time distribution consistency coefficient, and compares the index with a preset credibility threshold; when the second abnormal data credibility index is greater than or equal to the preset credibility threshold, the data is judged as normal data and does not need to be deleted; when the second abnormal data credibility index is less than the preset credibility threshold, the data is marked as "suspicious data buffer" and enters the manual review process; the credibility threshold is obtained by a multi-index joint stability analysis method, which is a method of introducing multiple characteristic indexes reflecting data stability or credibility, analyzing the minimum credible value of each index, calculating the comprehensive stability boundary, and taking the boundary value as the threshold.
[0013] Preferably, the time distribution consistency coefficient obtaining step is: obtaining the value of the current timestamp of the multi-element data, the standard deviation of the historical timestamp and the mean value of the historical timestamp, and obtaining the time distribution consistency coefficient by calculating the value of the current timestamp of the multi-element data, the standard deviation of the historical timestamp and the mean value of the historical timestamp.
[0014] Preferably, the data source reliability coefficient obtaining step is: obtaining the historical abnormal data amount and the historical total data amount, calculating the historical abnormal rate by ratio calculation of the historical abnormal data amount and the historical total data amount, obtaining the data fault times in the past thirty days and the total data running times, calculating the equipment health index by ratio calculation of the past thirty days of fault days and the total running days, and calculating the source reliability by the historical abnormal rate and the equipment health index in the data to obtain the data source reliability coefficient.
[0015] Preferably, the context similarity coefficient obtaining step is: obtaining the category set by extracting unique values through the category field in the current data, obtaining the reference set by extracting high-frequency values of the category field in the historical data, obtaining the category type coefficient by calculating the category set and the reference set, obtaining the space type by calculating the actual spherical surface and the distance threshold between the geographic position of the current data point and the reference position, and obtaining the time series coefficient by calculating the attenuation coefficient and the dynamic time warping according to the attenuation coefficient and the dynamic time warping.
[0016] Technical effects and advantages of the present application:
[0017] The present application significantly improves the accuracy and reliability of data management by the multi-level abnormality detection and verification mechanism combined with the artificial auditing process, effectively solves the problem of misdeletion or omission of data in the traditional method, and thus improves the overall quality and efficiency of data processing in the cloud computing environment. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 A flow chart of a cloud computing-based integrated data platform management method is provided for the embodiments of the present application. DETAILED DESCRIPTION
[0019] The technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application, and additionally, the forms of each structure described in the following embodiments are only examples, and the cloud computing-based integrated data platform management method involved in the present application is not limited to each structure described in the following embodiments, and all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope of protection of the present application.
[0020] The application provides a cloud computing-based integrated data platform management method, which comprises the following steps: Figure 1
[0021] Step 1: Collecting multi-element data through a cloud computing platform, and performing format unification processing on the multi-element data to obtain multi-element data in a unified format.
[0022] Step 2: Preprocessing the multi-element data in the unified format to obtain preprocessed multi-element data, obtaining abnormal information of the preprocessed multi-element data, the abnormal information comprising field information, source abnormal information and time distribution information, and obtaining a multi-element data abnormal coefficient according to the abnormal information; and determining whether to trigger rapid verification according to the multi-element data abnormal coefficient.
[0023] In this embodiment, it needs to be specifically explained that the multi-element data abnormal coefficient obtaining step is as follows:
[0024] Obtaining field information of the multi-source data, the field information comprising a missing field number, a required field number, a valid associated record number and a field value number meeting a value range requirement, and obtaining data integrity according to the field information.
[0025] Obtaining source abnormal information of the multi-source data, the source abnormal information comprising a historical abnormal rate, a device health rate, a transmission interruption rate and a same-source data conflict rate, and obtaining a data source abnormal rate according to the source abnormal information.
[0026] Obtaining time distribution information of the multi-source data, the time distribution information comprising a time stamp outlier index, an arrival interval abnormal rate, a periodic pattern deviation rate, a time density abnormal rate and a time sequence conflict rate, and obtaining a time distribution deviation rate according to the time distribution information.
[0027] Performing normalization processing on the data integrity, the data source abnormal rate and the time distribution deviation rate, and calculating the multi-element data abnormal coefficient according to the normalized integrity, the data source abnormal rate and the time distribution deviation rate, and the specific steps are as follows:
[0028]
[0029] In the formula, AS is the multi-element data abnormal coefficient, DI is the integrity, SR is the data source abnormal rate, and TD is the time distribution deviation rate, the multiplication and square root operation balances the contributions of the four indexes, avoids the single index dominating the result, and quantifies the overall data reliability.
[0030] In this embodiment, it needs to be specifically explained that the specific obtaining step of the data integrity is as follows:
[0031] Obtaining the number of missing fields of the multi-element data and the total number of fields, calculating the missing rate by ratio of the number of missing fields to the total number of fields;
[0032] Obtaining the number of filled mandatory fields of the multi-element data, calculating the mandatory rate by ratio of the number of filled mandatory fields to the total number of fields;
[0033] Obtaining the number of valid associated records of the multi-element data, calculating the associated completeness rate by ratio of the number of valid associated records to the total number of associated records;
[0034] Obtaining the number of field values meeting the value range of the multi-element data, calculating the value range completeness rate by ratio of the number of field values meeting the value range to the total number of field values;
[0035] The missing rate, the mandatory rate, the associated completeness rate and the value range completeness rate are normalized, and the data integrity is calculated by the normalized missing rate, the mandatory rate, the associated completeness rate and the value range completeness rate, and the specific obtaining steps are:
[0036]
[0037] In the formula, The data integrity is, The missing rate is, The mandatory rate is, The associated completeness rate is, The value range completeness rate is, and the four indexes are multiplied and raised to the fourth power to realize balanced influence of each index and avoid covering of extreme values by arithmetic mean.
[0038] In the embodiment, it needs to be specifically explained that the step of obtaining the data source abnormality rate is:
[0039] Obtaining the number of historical abnormal data of the multi-element data and the total number of historical data, calculating the historical abnormality rate by ratio of the number of historical abnormal data to the total number of historical data;
[0040] Obtaining the number of past thirty days of faults and the total number of running days of the multi-element data, calculating the device health rate by ratio of the number of past thirty days of faults to the total number of running days;
[0041] Obtaining the number of transmission interruptions in the last 24 hours and the expected number of transmissions of the multi-element data, calculating the transmission interruption rate by ratio of the number of transmission interruptions in the last 24 hours to the expected number of transmissions;
[0042] Obtaining the number of logical conflict fields and the total number of comparable words of the multi-element data, calculating the homologous data conflict rate by ratio of the number of logical conflict fields to the total number of comparable words;
[0043] The historical abnormal rate, the equipment health rate, the transmission interruption rate and the homologous data conflict rate are normalized, and the data source abnormal rate is calculated based on the normalized historical abnormal rate, the equipment health rate, the transmission interruption rate and the homologous data conflict rate, and the specific acquisition steps are as follows:
[0044]
[0045] In the formula, The data source abnormal rate is The number of indicators participating in the calculation is The historical abnormal rate is The equipment health rate is CR, the transmission interruption rate is HC, and the homologous data conflict rate is HC. The harmonic average formula method is used to give greater weight to smaller abnormal values, so as to quickly expose the weakest abnormal data source. The number of indicators refers to the total number of statistical indicators used to calculate the data source abnormal rate, which specifically includes four indicators of historical abnormal rate, equipment health rate, transmission interruption rate and homologous data conflict rate.
[0046] In this embodiment, it needs to be specifically pointed out that the step of obtaining the time distribution offset rate is as follows:
[0047] The actual arrival time, the expected time and the allowed time deviation threshold of each data in the multi-element data are obtained, and the timestamp outlier index of each data is obtained by calculating the actual arrival time, the expected time and the allowed time deviation threshold. The specific acquisition steps are as follows:
[0048]
[0049] In the formula, The timestamp outlier index is ta s The expected time is Δmt, and the allowed deviation threshold is
[0050] The minimum cumulative distance of the optimal matching path of the current cycle sequence and the reference cycle sequence is calculated based on the current cycle sequence and the reference cycle sequence to obtain the dynamic time warping similarity, and the reference cycle similarity is obtained by ratio calculation of the current cycle sequence and the historical reference sequence.
[0051] According to the dynamic time warping similarity and the reference cycle similarity, the dynamic time warping similarity and the reference cycle similarity are calculated by ratio calculation to obtain the cycle mode deviation rate.
[0052] The data amount in the current time window of the multi-element data, the mean of the historical synchronous window data amount and the standard deviation of the historical synchronous window data amount are obtained, and the time density abnormal rate is obtained by calculating the data amount in the current time window of the multi-element data, the mean of the historical synchronous window data amount and the standard deviation of the historical synchronous window data amount. The specific acquisition steps are as follows:
[0053]
[0054] In the formula, is the time density anomaly rate, N t is the data amount in the current time window, is the mean of the data amount in the historical same period window, is the standard deviation of the data amount in the historical same period window;
[0055] The timestamp outlier index, the cycle pattern deviation rate, and the time density anomaly rate are normalized, and the time distribution deviation rate is calculated based on the normalized timestamp outlier index, cycle pattern deviation rate, and time density anomaly rate. The specific acquisition steps are as follows:
[0056]
[0057] In the formula, is the time distribution deviation rate, is the timestamp outlier index, and CD is the cycle pattern deviation rate, is the time density anomaly rate. The average root formula is used to square the operation to amplify the contribution of significant anomaly indicators, highlight major problems, and be more sensitive to time problems that significantly deviate from the normal range. t is the reference unit of time distribution, which is used to unify the calculation of the timestamp outlier index, cycle pattern deviation rate, and time density anomaly rate to ensure the consistency of the time dimension.
[0058] In this embodiment, it needs to be specifically explained that the step of determining whether to trigger the rapid check based on the multi-element data anomaly coefficient is as follows: comparing the multi-element data anomaly coefficient with the anomaly threshold value. If the multi-element data anomaly coefficient is greater than or equal to the preset anomaly threshold value, the data will trigger the rapid check. If the multi-element data anomaly coefficient is less than the preset anomaly threshold value, the rapid check will not be triggered. The anomaly threshold value is obtained by the static threshold value acquisition method. The static threshold value is obtained by analyzing historical data to calculate the mean, standard deviation, or percentile and a fixed value preset based on business experience.
[0059] Step 3: If it is determined to trigger the rapid check, the data marked as “abnormal” in the rapid check is obtained, which is recorded as the first abnormal data.
[0060] Step 4: The first abnormal data is subjected to secondary check. If the secondary check result is abnormal, the first abnormal data is deleted. If the secondary check result is normal, the data is temporarily stored in the “suspicious data buffer area” to trigger the manual review process.
[0061] In this embodiment, it needs to be specifically explained that the step of performing secondary check on the first abnormal data is as follows:
[0062] obtain the credibility data of the second abnormal data, the credibility data including time distribution data, data source information and context information;
[0063] evaluate the time distribution consistency coefficient according to the time distribution data of the first abnormal data;
[0064] evaluate the data source credibility coefficient according to the data source information of the first abnormal data;
[0065] evaluate the context similarity coefficient according to the context information of the first abnormal data;
[0066] normalize the time distribution consistency coefficient, the data source credibility coefficient and the context similarity coefficient, and calculate the second abnormal data credibility index according to the normalized time distribution consistency coefficient, the normalized data source credibility coefficient and the normalized context similarity coefficient, the specific obtaining steps being:
[0067] DT = w1 x TD + w2 x SC + w3 x CS;
[0068] In the formula, DT is the second abnormal data credibility index, TD is the time distribution consistency coefficient, SC is the data source credibility coefficient, CS is the context credibility coefficient, w1, w2 and w3 are the time distribution consistency weight coefficient, the data source credibility weight coefficient and the context credibility weight coefficient.
[0069] In the last stage of the data processing process, the system calculates the second abnormal data credibility index by evaluating the data source credibility coefficient, the context similarity coefficient and the time distribution consistency coefficient, compares the index with the preset credibility threshold, and thus obtains the final data quality determination conclusion. When the second abnormal data credibility index is greater than or equal to the preset credibility threshold, the data is determined to be normal data and does not need to be deleted. When the second abnormal data credibility index is less than the preset credibility threshold, the data is marked as "suspicious data buffer" and enters the manual review process.
[0070] The credibility threshold is obtained by the multi-index joint stability analysis method. The multi-index joint stability analysis method is to analyze the minimum credibility value of each index by introducing multiple characteristic indexes reflecting data stability or credibility, i.e. the time distribution consistency coefficient, the data source credibility coefficient and the context similarity coefficient, to calculate the comprehensive stability boundary, and to take the boundary value as the threshold.
[0071] In this embodiment, it is specifically noted that the specific obtaining steps of the time distribution consistency coefficient are:
[0072] The value of the current timestamp of the multi-element data, the standard deviation of the historical timestamp and the mean value of the historical timestamp are obtained, the time distribution consistency coefficient is obtained by calculating the value of the current timestamp of the multi-element data, the standard deviation of the historical timestamp and the mean value of the historical timestamp, and the specific obtaining steps are as follows:
[0073]
[0074] In the formula, The time distribution consistency coefficient is X, the value of the current timestamp, σ is the standard deviation of the historical timestamp, and μ is the mean value of the historical timestamp.
[0075] In this embodiment, it needs to be specifically explained that the specific obtaining steps of the data source credibility coefficient are as follows:
[0076] The historical abnormal data amount and the historical total data amount are obtained, the historical abnormal data amount and the historical total data amount are calculated by ratio, and the historical abnormal rate is obtained.
[0077] The data fault times and the total running times in the past thirty days are obtained, the fault days and the total running days in the past thirty days are calculated by ratio, and the equipment health index is obtained.
[0078] According to the historical abnormal rate and the equipment health index in the data, the source credibility is calculated, and the specific obtaining steps are as follows:
[0079]
[0080] In the formula, The data source credibility coefficient is The historical abnormal rate is The equipment health index is.
[0081] In this embodiment, it needs to be specifically explained that the specific obtaining steps of the context similarity coefficient according to the context information of the first abnormal data are as follows:
[0082] The category set is obtained by extracting unique values by counting the category field in the current data, and the reference set is obtained by counting the high frequency values of the category field in the historical data.
[0083] The category set and the reference set are obtained, and the category type coefficient is obtained by calculating the category set and the reference set, and the specific obtaining steps are as follows:
[0084]
[0085] In the formula, The category type coefficient is A, the current category set, and B is the reference set.
[0086] The actual spherical surface and distance threshold between the geographical position of the current data point and the reference position is obtained, and the spatial type coefficient is obtained by calculating the actual spherical surface and distance threshold between the geographical position of the current data point and the reference position, and the specific obtaining steps are as follows:
[0087]
[0088] In the formula, is the spatial type coefficient, max is to ensure that the spatial type is non-negative, d is the actual spherical surface between the geographical position of the current data point and the reference position, dt is the distance threshold, and the geographical position of the data point is the latitude and longitude coordinates recorded when the data point is collected. The reference position refers to a known reference coordinate point, such as a certain business base station, a monitoring station, a warehouse center, a task area center point, etc.
[0089] The dynamic time warping is obtained, which is obtained by calculating the similarity between the current period sequence and the historical reference sequence through ratio;
[0090] The decay coefficient is obtained, which is a preset parameter for controlling the influence degree of time series similarity on context similarity;
[0091] The time series coefficient is obtained by calculating the decay coefficient and the dynamic time warping, and the specific obtaining steps are as follows:
[0092] ST = e -λ·DTW ;
[0093] In the formula, ST is the time series coefficient, λ is the decay coefficient, and DTW is the dynamic time warping.
[0094] The context similarity coefficient is calculated according to the category type coefficient, the spatial type coefficient and the time series coefficient in the data, and the specific obtaining steps are as follows:
[0095]
[0096] In the formula, CS is the context similarity coefficient, is the number of indicators participating in the calculation, Cg is the category type coefficient, S is the spatial type coefficient, and ST is the time series coefficient. The square root formula is used to multiply and then take the cube root, so that the influence of the three types of coefficients on the result is balanced. The matching degree of the category type coefficient, the spatial type coefficient and the time series coefficient is considered, and the number of indicators is used to evaluate the total number of the three types of core indicators, namely the data category type coefficient, the spatial type coefficient and the time series coefficient.
[0097] Step 5: manually auditing the data with normal secondary verification result, performing deterministic processing, and obtaining an actual effective data set;
[0098] In this embodiment, it needs to be specifically pointed out that the actual effective data set step is:
[0099] If the artificial audit determines that the secondary determination abnormal data is normal data, it is moved from the "suspicious data buffer" to the "valid data set", and marked as "artificially confirmed normal";
[0100] If the artificial audit determines that the secondary determination abnormal data is abnormal data: delete it from the "suspicious data buffer";
[0101] Traverse all secondary determination abnormal data, and record the current valid data set as the actual valid data set.
[0102] Need to be specified, the valid data set is a high-quality data set selected through the format unification, abnormality detection, secondary verification and artificial audit process, finally stored and backed up by the cloud storage platform, ensuring data integrity, traceability and high availability.
[0103] Step 6: According to the actual valid data set, execute data management operation in the cloud storage platform, the data management operation includes indexing storage of the actual valid data set according to data type, source identification and timestamp information, establishing version control relationship and recording change information, at the same time, calling the redundant backup mechanism of the cloud storage platform to perform multi-copy disaster recovery backup, so as to ensure the integrity, traceability and high availability of the data.
[0104] Finally: The above only describes the preferred embodiments of the present application and is not used to limit the present application, any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application should be included in the protection scope of the present application.
[0105] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited to this, any modification, equivalent replacement, improvement, etc. within the technical range disclosed by the present application can be easily thought by those skilled in the art, and should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A cloud computing-based integrated data platform management method, characterized in that, The system includes; Step 1: Collect multi-source data through a cloud computing platform, process the multi-source data to unify the format, and obtain multi-source data in a unified format; Step 2: Preprocess the multi-data in a unified format to obtain preprocessed multi-data. Obtain the anomaly information of the preprocessed multi-data, which includes field information, source anomaly information, and time distribution information. Evaluate the multi-data anomaly coefficient based on the anomaly information. Determine whether to trigger fast verification based on the multi-data anomaly coefficient. Step 3: If a quick check is triggered, retrieve the data marked as "abnormal" in the quick check and record it as the first abnormal data; Step 4: Perform a second check on the first abnormal data. If the second check result is abnormal, delete the first abnormal data. If the second check result is normal, temporarily store the data in the "suspicious data buffer" and trigger the manual review process. Step 5: Manually review the data whose secondary verification results are normal, perform deterministic processing, and obtain the actual valid dataset; Step 6: Perform data management operations on the cloud storage platform based on the actual valid dataset.
2. The integrated data platform management method based on cloud computing according to claim 1, characterized in that, The steps for obtaining anomaly coefficients in multivariate data are as follows: Obtain field information from multi-source data, including the number of missing fields, the number of required fields, the number of valid associated records, and the number of field values that meet the value range requirements. Evaluate data integrity based on the field information. Obtain source anomaly information for multi-source data, including historical anomaly rate, device health rate, transmission interruption rate, and same-source data conflict rate. Evaluate the data source anomaly rate based on the source anomaly information. Obtain the time distribution information of multi-source data, including timestamp outlier index, arrival interval anomaly rate, periodic pattern deviation rate, time density anomaly rate, and time sequence conflict rate. Based on the time distribution information, obtain the time distribution offset rate. The data integrity, data source anomaly rate, and time distribution offset rate are normalized, and the multivariate data anomaly coefficient is calculated from the normalized integrity, data source anomaly rate, and time distribution offset rate.
3. The integrated data platform management method based on cloud computing according to claim 2, characterized in that, The steps to ensure data integrity are as follows: Obtain the number of missing fields and the total number of fields in the multivariate data, and calculate the missing rate by the ratio of the number of missing fields to the total number of fields. Obtain the number of filled and required fields in the multivariate data, and calculate the ratio of the number of filled and required fields to the total number of fields to obtain the required field rate; Obtain valid related records from multi-source data, and calculate the ratio of the number of valid related records to the total number of related records to obtain the association completeness rate; Obtain the number of field values that conform to the value range in the multivariate data, and calculate the ratio of the number of field values that conform to the value range to the total number of field values to obtain the value range completeness rate; The missing rate, required fill rate, association completeness rate, and value range completeness rate are normalized, and the data integrity is calculated from the normalized missing rate, required fill rate, association completeness rate, and value range completeness rate.
4. The integrated data platform management method based on cloud computing according to claim 2, characterized in that, The steps to obtain the data source anomaly rate are as follows: Obtain the number of historical outlier data entries and the total number of historical data entries from the multivariate data set. Calculate the ratio of the number of historical outlier data entries to the total number of historical data entries to obtain the historical outlier rate. The equipment health rate is obtained by obtaining the number of failures and the total number of operating days in the past 30 days from multi-dimensional data, and calculating the ratio of the number of failures in the past 30 days to the total number of operating days. Obtain the number of transmission interruptions and the expected number of transmissions in the past 24 hours from multi-dimensional data, and calculate the ratio of the number of transmission interruptions to the expected number of transmissions to obtain the transmission interruption rate; Obtain the number of logically conflicting fields and the total number of comparable words from the multi-source data. Calculate the ratio of the number of logically conflicting fields to the total number of comparable words to obtain the conflict rate of the source data. The historical anomaly rate, equipment health rate, transmission interruption rate, and same-source data conflict rate are normalized, and the data source anomaly rate is calculated from the normalized historical anomaly rate, equipment health rate, transmission interruption rate, and same-source data conflict rate.
5. The integrated data platform management method based on cloud computing according to claim 2, characterized in that: The steps to obtain the time distribution offset are as follows: Obtain the actual arrival time, expected time, and allowable time deviation threshold for each data point in the multivariate data set. Calculate the timestamp outlier index for each data point using the actual arrival time, expected time, and allowable time deviation threshold. Obtain the current periodic sequence and the reference periodic sequence, calculate the minimum cumulative distance of the optimal matching path between the current periodic sequence and the reference periodic sequence to obtain the dynamic time warping similarity, and calculate the reference period similarity between the current periodic sequence and the historical reference sequence by the ratio; The ratio of the dynamic event regularization similarity to the baseline periodic similarity is calculated to obtain the periodic pattern deviation rate; The mean of the data volume in the current time window of the multivariate data, the mean of the data volume in the same window in the same period of the history, and the standard deviation of the data volume in the same window in the history are obtained. The time density anomaly rate is obtained by calculating the mean of the data volume in the current time window of the multivariate data, the mean of the data volume in the same window in the history, and the standard deviation of the data volume in the same window in the history. The timestamp outlier index, periodic pattern deviation rate, and time density anomaly rate are normalized, and the time distribution offset rate is calculated from the normalized timestamp outlier index, periodic pattern deviation rate, and time density anomaly rate.
6. The integrated data platform management method based on cloud computing according to claim 1, characterized in that, The steps for performing a second verification on the first abnormal data are as follows: Obtain the credibility data of the first abnormal data. The credibility data includes time distribution data, data source information, and contextual information. The time distribution consistency coefficient is obtained by evaluating the time distribution data of the first abnormal data. The data source credibility coefficient is obtained by evaluating the data source information of the first abnormal data. The context similarity coefficient is obtained by evaluating the context information of the first abnormal data; The time distribution consistency coefficient, data source credibility coefficient, and context similarity coefficient are normalized, and the normalized time distribution consistency coefficient, data source credibility coefficient, and context similarity coefficient are used to calculate the second abnormal data credibility index. The credibility index of the second abnormal data is compared with the preset credibility threshold. If the credibility index of the second abnormal data is greater than or equal to the preset credibility threshold, the data is judged to be normal data and is not deleted. If the credibility index of the second abnormal data is less than the preset credibility threshold, the data is stored in the "suspicious data buffer" and enters the manual review process.
7. The integrated data platform management method based on cloud computing according to claim 6, characterized in that, The steps to obtain the time distribution consistency coefficient are as follows: Obtain the current timestamp value, the standard deviation of historical timestamps, and the mean of historical timestamps from the multivariate data. Calculate the time distribution consistency coefficient by taking the current timestamp value, the standard deviation of historical timestamps, and the mean of historical timestamps from the metadata.
8. The integrated data platform management method based on cloud computing according to claim 6, characterized in that, The steps to obtain the data source credibility coefficient are as follows: Obtain the historical abnormal data volume and the historical total data volume, and calculate the ratio of the historical abnormal data volume to the historical total data volume to obtain the historical abnormality rate; Obtain the number of data failures and the total number of data runs in the past 30 days, and calculate the ratio of the number of failure days to the total number of running days in the past 30 days to obtain the equipment health index; The source credibility is calculated by obtaining the historical anomaly rate and equipment health index from the data. The data source credibility coefficient is obtained by calculating the source credibility of the data source based on the historical anomaly rate and equipment health index.
9. The integrated data platform management method based on cloud computing according to claim 6, characterized in that, The steps to obtain the context similarity coefficient are as follows: By statistically analyzing the category field in the current data, unique values are extracted to obtain a category set, and a benchmark set is obtained by statistically analyzing the high-frequency values of the category field in historical data. Obtain the category set and the benchmark set, and calculate the category coefficients by combining the category set and the benchmark set; Obtain the actual sphere and distance threshold between the current data point's geographic location and the reference location, and calculate the spatial coefficient by calculating the actual sphere and distance threshold between the current data point's geographic location and the reference location; Based on the attenuation coefficient and dynamic time warping, the timing coefficient is obtained by calculating the attenuation coefficient and dynamic time warping. The context similarity coefficient is obtained by calculating the categorical coefficient, spatial coefficient, and temporal coefficient.