Quality governance and real-time data verification method based on data sandbox

CN120596369BActive Publication Date: 2026-08-11GUANGZHOU YITUO SOFTWARE DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]然而,尽管基于数据沙箱的质量治理与实时数据验证方法有许多优点,但在实际应用中,仍然存在一些潜在的缺陷和挑战,传统的数据治理方法多侧重于数据采集后期的批处理与验证,存在较大的数据延迟,且人工干预较多,处理效率较低,数据沙箱技术为数据验证提供了一个隔离的测试环境,能够在不影响实际业务的前提下进行数据质量验证,但现有技术中的数据沙箱环境下的质量治理方法多存在同步性能不足、实时性差及智能化程度低等问题

Benefits of technology

[0200]通过分布式存储技术和流式数据传输,生产环境的数据能够被实时同步到沙箱环境中;通过多层次质量治理规则引擎,系统根据数据的特性和实时变化动态调整验证规则;通过集成异常检测模型,系统能够自动识别数据中的异常模式,及时发现潜在的数据质量问题;根据历史验证结果和实时检测的数据质量问题,系统能够自动调整验证策略;在沙箱环境中采用数据脱敏和加密技术,确保即使数据在沙箱环境中进行验证,也不会泄露敏感信息;系统实时监控数据验证过程,确保数据处理和验证的每个环节都符合预期;系统通过生成报告和追踪记录,提供了全面的历史数据验证和异常事件记录;通过多层次的数据验证与异常检测,数据质量得到确保;系统能够根据实时数据流的变化动态调整数据验证策略,保证始终根据最新的数据情况做出最优化的决策。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure QLYQS_6
    Figure QLYQS_6
  • Figure QLYQS_12
    Figure QLYQS_12
Patent Text Reader

Abstract

This invention discloses a quality governance and real-time data verification method based on a data sandbox, comprising: real-time synchronization of production data to the sandbox environment using distributed storage and streaming data transmission technologies; automatic adjustment of verification rules using a multi-level quality governance rule engine; real-time monitoring of each data item by the system; identification and alarm using an anomaly detection model; automatic adjustment of verification strategies; security assurance through de-identification and encryption technologies during data verification; and the introduction of an automated monitoring and reporting module to periodically generate data quality reports and provide anomaly tracking and improvement suggestions. This invention's data sandbox-based quality governance and real-time data verification method improves data accuracy and intelligence through multi-level data verification, automated synchronization and monitoring mechanisms, and intelligent anomaly detection, reducing reliance on manual intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of quality governance of data sandboxes, specifically involving a method for quality governance and real-time data verification based on data sandboxes. Background Technology

[0002] Data sandbox-based quality governance and real-time data validation are techniques that isolate datasets in a controlled environment for quality governance and validation. This approach effectively ensures data accuracy, integrity, and security, and guarantees that data meets predetermined quality standards during processing. A data sandbox is a restricted, isolated testing environment that allows data scientists, analysts, and others to process and analyze data without impacting the production environment. Quality governance involves ensuring data meets certain quality standards, such as accuracy, consistency, completeness, timeliness, and reliability. Sandbox-based quality governance helps organizations perform data quality checks and remediation without directly affecting the production environment. Real-time data validation means immediate inspection and verification of data streams to ensure data conforms to predetermined rules and standards. By performing real-time validation before data enters the main database or analytics platform, organizations can avoid the impact of erroneous data and reduce the cost and time of subsequent remediation. Data processing within a data sandbox avoids potential risks and data contamination to the production environment. Multiple data processing, model testing, and validation can be performed in the sandbox environment to flexibly address various business needs. Rigorous quality governance processes ensure data meets business requirements and reduce quality issues. Real-time validation ensures that problems are detected and corrected promptly before data enters the production environment.

[0003] However, despite the many advantages of data sandbox-based quality governance and real-time data verification methods, there are still some potential defects and challenges in practical applications. Traditional data governance methods often focus on batch processing and verification in the later stages of data collection, resulting in significant data latency, excessive manual intervention, and low processing efficiency. Data sandbox technology provides an isolated testing environment for data verification, enabling data quality verification without affecting actual business operations. However, existing data sandbox-based quality governance methods often suffer from insufficient synchronization performance, poor real-time performance, and low level of intelligence. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the purpose of this invention is to provide a quality governance and real-time data verification method based on a data sandbox. Through a multi-layered data verification mechanism, an automated synchronization and monitoring mechanism, and intelligent anomaly detection methods, the real-time performance, accuracy, and intelligence of data verification are improved, while reducing the reliance on manual intervention.

[0005] The technical solution adopted by this invention to solve its technical problem is:

[0006] The quality governance and real-time data verification method based on data sandbox includes the following steps:

[0007] A data sandbox environment is built, employing distributed storage technology and a real-time data synchronization mechanism to synchronize data from the production environment to the sandbox environment in real time. This synchronization process is achieved through streaming data transmission technology, ensuring high efficiency and low latency in the data synchronization process.

[0008] In the sandbox environment, a multi-level quality governance rule engine is adopted. This engine automatically assigns different verification rules according to the characteristics of the data, and the rule engine dynamically adjusts the verification rules according to real-time data changes.

[0009] The system monitors each piece of data synchronized to the sandbox in real time, automatically identifies abnormal patterns in the data based on the anomaly detection model, detects potential data quality problems, and automatically sends an alarm to the administrator after an anomaly is detected, prompting relevant personnel to handle the situation.

[0010] Based on changes in real-time data streams, anomaly detection results, and historical verification results, the system automatically adjusts its data verification strategy. When the system detects that certain data quality issues occur frequently, the verification strategy will increase the verification intensity; conversely, the system will automatically reduce the verification burden.

[0011] The system introduces an automated monitoring and report generation module to monitor the data verification process in real time and generate data quality reports for managers to review regularly. The report content includes statistical information on data quality issues, tracking records of abnormal events, and suggestions for targeted improvement measures.

[0012] When processing data, the system employs data anonymization and encryption technologies to ensure that data in the sandbox environment is not leaked during the verification process. In addition, the system includes inspection tools to automatically verify whether the use of data in the sandbox environment complies with relevant data protection regulations.

[0013] As a preferred approach, a data sandbox environment is constructed, employing distributed storage technology and a real-time data synchronization mechanism to synchronize data from the production environment to the sandbox environment in real time. This synchronization process is achieved through streaming data transmission technology. The method to ensure the high efficiency and low latency of the data synchronization process is as follows:

[0014] Data transmission latency is expressed as the time required for data to travel from the source system to the target system, and is defined as follows:

[0015] T sync This represents the total delay time of the data synchronization process.

[0016] T nctwork This refers to network transmission time.

[0017] Tprocessing For data processing time;

[0018] T queue The message queue wait time;

[0019] The formula is:

[0020] T sync =T network +T proccssing +T queuc

[0021] Among them, T network Calculated using network bandwidth and data volume:

[0022]

[0023] Where D is the data size and B is the network bandwidth;

[0024] T processing For the time required for data processing;

[0025] T qucuc Related to the queue length and throughput of streaming data, assuming a queue of Q messages and a system processing rate of R messages / second, then:

[0026]

[0027] Throughput is the amount of data successfully transmitted per unit of time, measured in data per second, and is defined as follows:

[0028] T throughput For throughput;

[0029] D is the size of each data block;

[0030] R is the number of data blocks processed per second;

[0031] The formula is:

[0032] T throughput =D×R

[0033] Where D is the size of the data block and R is the processing speed of the data block per second;

[0034] Streaming data synchronization efficiency is measured as the ratio of actual data transmission efficiency to processing efficiency, defined as:

[0035] E sync To improve data synchronization efficiency;

[0036] T sync This represents the total time for data synchronization.

[0037] T idealThis represents the ideal data synchronization time.

[0038] The formula is:

[0039]

[0040] When synchronizing data in real time, ensure the consistency and integrity of the data in the sandbox environment. Establish a data verification process using a checksum method, with the following formula:

[0041] H prod For checksums of data in the production environment;

[0042] H sandbox This is a checksum for the synchronized data in the sandbox environment.

[0043] The formulas for consistency and integrity verification are:

[0044] Valid Date = (H prod ==H sandbox )

[0045] When H prod and H sandbox When they are equal, the data is synchronized and kept consistent in the sandbox environment.

[0046] As a preferred approach, a multi-level quality governance rule engine is employed in the sandbox environment. This engine automatically assigns different verification rules based on the characteristics of the data. The method by which the rule engine dynamically adjusts the verification rules based on real-time data changes is as follows:

[0047] The characteristics of the data are used to assign corresponding validation rules based on a predefined set of rules. The following variables are set:

[0048] D is the data characteristic vector, representing the various characteristics of the data, and is represented by a multi-dimensional vector:

[0049] D = [D1, D2, D3, ..., D] n ]

[0050] Among them, D i This represents the i-th dimension in the data characteristics;

[0051] R is the validation rule set, which contains various rules:

[0052] R = {R1, R2, R3, ..., R} m}

[0053] Among them, R i This is the i-th validation rule;

[0054] C(D i R j ) represents data characteristic D iWith rule R j Fit, representing the degree of matching between a feature and a rule, is measured using a scoring function:

[0055] C(D i R j )=f(D i R j )

[0056] Where f is set according to the relationship between characteristics and rules;

[0057] The verification rule assignment process is completed using the following formula:

[0058] R assigncd ={R j |C(D i R j )>θ}

[0059] Where θ is a threshold, when the data characteristics fit the rule C(D) i R j When the value is higher than the threshold, rule R j Assigned to data characteristic D i ;

[0060] As data changes in real time, the rule engine dynamically adjusts validation rules, involving real-time monitoring of data changes and adjustments based on new data characteristics or quality standards, establishing:

[0061] ΔD represents the amount of change in data over a specific time period;

[0062] ΔR represents the amount of change in the validation rules caused by changes in the data;

[0063] The formula for dynamically adjusting the verification rules is:

[0064] ΔR = g(ΔD, α, β)

[0065] Where g(·) is a function used to dynamically adjust the rule set based on the amount of data change and system settings;

[0066] α is a sensitivity parameter, representing how sensitive the rule adjustment is to changes in the data, and β is an adjustment threshold used to determine when to change the rule;

[0067] During the rule adjustment process, the following was established:

[0068] R new =R assigncd ∪ΔR

[0069] Among them, R new For the new set of verification rules;

[0070] The rule engine adjusts future rule assignments based on the validation results, establishing:

[0071] V rcsult (R j ) is rule R j The verification results;

[0072] ω(R j ) is rule R j The weight represents the importance of the rule in the overall verification process;

[0073] The adjustment formula for the rule feedback mechanism is:

[0074] ω new (R j )=ω(R j )×(1+λ×V result (R j ))

[0075] Where λ is the sensitivity factor for feedback adjustment;

[0076] In a multi-layered rule engine, rules are assigned and adjusted through a hierarchical decision tree, with the following variables set:

[0077] L represents the rule hierarchy;

[0078] P(L) represents the priority of level L;

[0079] The decision formula is:

[0080]

[0081] Where P(L) represents the priority of the rule hierarchy, and τ is the priority threshold, ensuring that only rules with high priority are selected into the final rule set R. final .

[0082] As a preferred approach, the system monitors each piece of data synchronized to the sandbox in real time, automatically identifies abnormal patterns in the data based on an anomaly detection model, detects potential data quality issues, and automatically sends an alarm to the administrator upon detecting an anomaly, prompting relevant personnel to take appropriate action.

[0083] The system monitoring data stream is set as a time series or multidimensional dataset, where each data point D t It will undergo anomaly detection;

[0084] D t =[D t1 D t2 , ..., D tn ] represents the data collected at time t;

[0085] This represents an anomaly detection model that outputs whether an anomaly exists. It can be built using a statistical model, a machine learning model, or a deep learning model.

[0086] The anomaly detection method used is distance-based, employing either an isolated forest or a probability-based model, such as the Gaussian distribution assumption model. The anomaly detection formula is expressed as:

[0087]

[0088] Among them, A t For data D at time t t If the abnormal detection result of A t =1 indicates that there is an anomaly in the data. If A t =0 indicates that the data is normal;

[0089] If based on a statistical model:

[0090]

[0091] Among them, D t This is the current data point;

[0092] μ is the mean vector of the data;

[0093] v is the covariance matrix of the data;

[0094] P(D t |μ,∑) represents the data D t The probability density;

[0095] For multidimensional or sequential data, anomaly patterns are sudden fluctuations in a certain feature or changes in the relationship between multiple features. Anomaly patterns can be identified by calculating the data's deviation or fluctuation.

[0096]

[0097] Among them, D t This refers to the observation data at the current moment;

[0098] These are predicted values ​​based on historical data.

[0099] σ(D t ) for data D t Standard deviation;

[0100] Data quality issues include missing values, duplicate values, and incorrect formatting. The detection of data quality issues is modeled using the following methods:

[0101] Missing value detection:

[0102]

[0103] Among them, M t Indicates whether there are missing values ​​at time t, 1 is the indicator function, if D ti If the value is missing, the value is 1; otherwise, the value is 0.

[0104] Duplicate value detection:

[0105]

[0106] Among them, R t This indicates whether there is duplicate data at time t. If the data items are equal, the value is 1; otherwise, it is 0.

[0107] Once an abnormal pattern or data quality issue is detected, the system will automatically issue an alarm. The alarm conditions are set as follows:

[0108] Alarm t =1(A t =1 or M t >0 or R t >00

[0109] If A t =1 or M t >0 or R t If the value is greater than 0, an alarm will be triggered;

[0110] Alarm priority is calculated based on the severity of the anomaly, with a weight w defined for each anomaly type. type :

[0111] Priority t =w type ×Severity t

[0112] Among them, w type Severity is the weight for the abnormal type. t The severity score is used to indicate the abnormality.

[0113] The formula for generating alarm notifications is:

[0114] Notification t =GenerateNotification(ALarm) t Priority t )

[0115] This formula generates a notification sent to the relevant administrator, indicating the abnormal data and its priority level.

[0116] Preferably, according to the changes in real-time data streams, anomaly detection results, and historical verification results, the system automatically adjusts the data verification strategy. When the system detects that certain data quality problems occur frequently, the verification strategy will increase the verification intensity. Conversely, the method for the system to automatically reduce the verification burden is as follows:

[0117] Define model variables:

[0118] D t : represents the data stream collected at time t;

[0119] A t : represents the anomaly detection result at time t;

[0120] M t : represents the frequency of data quality problems at time t;

[0121] V t : represents the data verification intensity at time t;

[0122] H t : represents the historical verification result at time t;

[0123] The system should calculate the frequency of anomalies based on the anomaly detection results and the frequency of data quality problems, and define the anomaly frequency F at each time point t as:

[0124]

[0125] A t is the result of anomaly detection, M t is the frequency of data quality problems. Through this formula, the system comprehensively considers the frequencies based on anomalies and data quality problems;

[0126] Based on the historical verification result H t and the current anomaly frequency F t , the system automatically adjusts the data verification strategy. Assume the system adopts a non-linear adjustment strategy and adjusts the verification intensity V according to the frequency and historical results t ;

[0127] The formula for adjusting the verification intensity is expressed as:

[0128] V t = V t1 + α × (F t - F threshold ) + β × H t

[0129] where, V t 1 is the verification intensity at the previous time point;

[0130] α is the anomaly frequency adjustment coefficient, which controls the impact of the anomaly frequency on the verification intensity;

[0131] F threshold This is the threshold for the frequency of anomalies, representing the minimum standard for frequent anomalies.

[0132] β is the adjustment factor for historical validation results;

[0133] H t This is a result verified historically;

[0134] In the system, the verification burden B t To verify the strength V t Measured, higher V t This indicates a high verification burden, and the system dynamically adjusts its approach by calculating the verification burden:

[0135] B t =γ×V t

[0136] Where γ is a constant coefficient.

[0137] As a preferred option, the system incorporates an automated monitoring and report generation module to monitor the data verification process in real time and periodically generate data quality reports for administrator review. The reports include statistical information on data quality issues, tracking records of abnormal events, and recommendations for targeted improvement measures.

[0138] Data quality issues include missing data, duplicate data, and format errors. The quantity and ratio of these issues are defined by the following formula.

[0139] Q t : Represents the total number of quality problems present in the data collected at time t;

[0140] Q total : Indicates the total amount of data in the dataset;

[0141] Data quality issue frequency P t Represented as:

[0142]

[0143] For each abnormal event, the system records the occurrence time, event type, and number of records affected;

[0144] E t : Indicates the number of abnormal events detected at time t;

[0145] E types : Indicates the type of abnormal event;

[0146] For each abnormal event e i Statistically analyze its occurrence frequency F i :

[0147]

[0148] The system should generate targeted improvement measures to address data quality issues and anomalies.

[0149] S t : Suggestions for improvement measures generated at time t

[0150] Each question type e i The corresponding suggested weight W i ;

[0151] The proposed improvement measures are generated using the following formula:

[0152]

[0153] n is the number of types of abnormal events, W i For each exception event type e i The weight of relevant improvement measures;

[0154] Monitor the real-time status of the data verification process, define real-time feedback indicators, and set up the system to generate a real-time monitoring report (R). t It includes statistical information on data quality issues and detailed tracking records of abnormal events;

[0155] Monitoring Report R t It consists of the following parts:

[0156] R t ={P t F1, F2, ..., F n S t}

[0157] Among them, P t F1, F2, ..., F are the ratios of data quality problems. n The frequency of occurrence of various abnormal events, S t Improvement suggestions based on abnormal events;

[0158] Regular reports should include a summary of data quality issues, trend analysis of abnormal events, and recommendations for targeted improvement measures. Regular reports should be generated periodically using the following formula:

[0159] The summary of data quality issues in periodic reports is based on the average data quality issue rate over a past period:

[0160]

[0161] Where T is the period length;

[0162] For each type of abnormal event, the system should track its changing trend and use a moving average method to smooth the frequency changes of abnormal events:

[0163]

[0164] Regular reports should also include a summary of improvement measures for each type of anomaly, calculated based on suggested weights and frequencies:

[0165]

[0166] As a preferred approach, the system employs data anonymization and encryption technologies when processing data to ensure that data in the sandbox environment is not leaked during the verification process. Furthermore, the system includes a verification tool that automatically verifies whether data usage in the sandbox environment complies with relevant data protection regulations.

[0167] Data anonymization involves modifying or hiding sensitive information in data. An original dataset D = {d1, d2, ..., d...} is established. n}, where each data point d i Contains sensitive information, which is de-identified through the operation Δ(d) i ), to obtain a de-identified dataset D′={d′1,d′2,...,d′ n}, where d′ i =Δ(d) i () indicates a desensitized version;

[0168] The de-identification operation Δ employs various methods for data field f. i Desensitization, the desensitization rules are defined as follows:

[0169] Δ(f i ) = mask(f i )

[0170] Encryption transforms sensitive data into unreadable ciphertext using encryption algorithms, ensuring that even if the data is leaked, it cannot be maliciously used. Let the original data be D = {d1, d2, ..., d...}. n The encrypted data is D′={e1, e2, ..., e}. n}, where e i =Encrypt(d i );

[0171] Encryption algorithms include symmetric encryption and asymmetric encryption, and the encryption formula is:

[0172] e i =Encrypt k (d i )

[0173] Where k is the encryption key, Encrypt k For encryption operations, output e i The data is encrypted;

[0174] Ensure that data usage in the sandbox environment complies with relevant data protection regulations, and verify compliance through automated inspection tools. Establish an inspection function to verify whether dataset D′ complies with relevant regulations.

[0175] The specific inspection items include:

[0176] Data access control: Confirms whether data access is restricted to authorized users;

[0177] Data processing purpose: To ensure that data is used only for legitimate purposes;

[0178] Data minimization: Ensure that the amount of data processed complies with regulations and avoid over-collection;

[0179] Define a compliance verification formula as follows:

[0180]

[0181] To further ensure data security, the system monitors and records all data access behaviors, assuming A = {a1, a2, ..., a...} m} represents all access records, where a i =(u i , t i d i ) represents user u i At time t i For data d i Access behavior;

[0182] The system should generate an access audit report. t And conduct compliance checks on data access:

[0183] L t ={a1, a2, ..., a m}

[0184] Verify compliance with access control policies by checking access logs:

[0185]

[0186] The system introduces a function to detect potential data breaches. By monitoring anonymized and encrypted data, it identifies data breach risks. The data breach warning trigger formula is as follows:

[0187]

[0188] Before entering the sandbox environment, the data undergoes desensitization and encryption processing, using the following formula:

[0189] D′={Δd(d1),Δ(d2),...,Δ(d n )}

[0190] Or encrypt:

[0191] D′={Encrypt k (d1), Encrypt k (d2), ..., Encrypt k (d n )}

[0192] Ensure data usage complies with relevant regulations through automated inspection tools:

[0193] CheckCompliant(D′)→True / False

[0194] All data access is logged and subject to compliance checks:

[0195] CheckAccess(L t → True / False and monitor potential risks through a leak warning mechanism:

[0196] LeakPrevention(D′)→True / False.

[0197] Another technical problem to be solved by the present invention is to provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the quality governance and real-time data verification method based on data sandbox as described above.

[0198] Another technical problem to be solved by the present invention is to provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method such as a data sandbox-based quality governance and real-time data verification method.

[0199] The beneficial effects of this invention are:

[0200] Through distributed storage technology and streaming data transmission, production environment data can be synchronized to the sandbox environment in real time. A multi-layered quality governance rule engine dynamically adjusts verification rules based on data characteristics and real-time changes. An integrated anomaly detection model automatically identifies anomaly patterns in the data, promptly detecting potential data quality issues. Based on historical verification results and real-time detected data quality problems, the system automatically adjusts verification strategies. Data anonymization and encryption technologies are used in the sandbox environment to ensure that sensitive information is not leaked even when data is verified there. The system monitors the data verification process in real time, ensuring that every step of data processing and verification meets expectations. The system provides comprehensive historical data verification and anomaly event records through report generation and tracking. Through multi-layered data verification and anomaly detection, data quality is ensured. The system dynamically adjusts data verification strategies based on changes in the real-time data stream, ensuring that optimal decisions are always made based on the latest data. Detailed Implementation

[0201] The principles and features of the present invention are described below. The examples given are for illustrative purposes only and are not intended to limit the scope of the invention. The invention is described more specifically by way of example in the following paragraphs. The advantages and features of the invention will become clearer from the following description and claims.

[0202] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0203] Example

[0204] The technical solution adopted by this invention to solve its technical problem is:

[0205] The quality governance and real-time data verification method based on data sandbox includes the following steps:

[0206] A data sandbox environment is built, employing distributed storage technology and a real-time data synchronization mechanism to synchronize data from the production environment to the sandbox environment in real time. This synchronization process is achieved through streaming data transmission technology, ensuring high efficiency and low latency in the data synchronization process.

[0207] In the sandbox environment, a multi-level quality governance rule engine is adopted. This engine automatically assigns different verification rules according to the characteristics of the data, and the rule engine dynamically adjusts the verification rules according to real-time data changes.

[0208] The system monitors each piece of data synchronized to the sandbox in real time, automatically identifies abnormal patterns in the data based on the anomaly detection model, detects potential data quality problems, and automatically sends an alarm to the administrator after an anomaly is detected, prompting relevant personnel to handle the situation.

[0209] Based on changes in real-time data streams, anomaly detection results, and historical verification results, the system automatically adjusts its data verification strategy. When the system detects that certain data quality issues occur frequently, the verification strategy will increase the verification intensity; conversely, the system will automatically reduce the verification burden.

[0210] The system introduces an automated monitoring and report generation module to monitor the data verification process in real time and generate data quality reports for managers to review regularly. The report content includes statistical information on data quality issues, tracking records of abnormal events, and suggestions for targeted improvement measures.

[0211] When processing data, the system employs data anonymization and encryption technologies to ensure that data in the sandbox environment is not leaked during the verification process. In addition, the system includes inspection tools to automatically verify whether the use of data in the sandbox environment complies with relevant data protection regulations.

[0212] By verifying every piece of data synchronized to the sandbox in real time, the system can immediately identify potential data quality issues, such as missing values, format errors, or data deviations. Through dynamic adjustment of verification rules and strategies, the system ensures that verification efforts are focused on high-risk data, avoiding over-verification of data without anomalies, thus saving computing resources and time. The system uses anonymization and encryption technologies during data verification to ensure that sensitive information is not leaked even when data testing and verification are conducted in the sandbox environment. Multi-layered quality governance and real-time anomaly detection ensure the quality of the data ultimately entering the decision-making system. The system monitors the data synchronization and verification process in real time, issuing alarms upon detecting anomalies. Every anomaly event is recorded in detail for subsequent analysis and tracing. This provides strong support for data governance, helping managers locate problems and analyze their causes.

[0213] A data sandbox environment is constructed, employing distributed storage technology and a real-time data synchronization mechanism to synchronize data from the production environment to the sandbox environment in real time. This synchronization process is achieved through streaming data transmission technology. The methods to ensure the high efficiency and low latency of the data synchronization process are as follows:

[0214] Data transmission latency is expressed as the time required for data to travel from the source system to the target system, and is defined as follows:

[0215] T sync This represents the total delay time of the data synchronization process.

[0216] T network This refers to network transmission time.

[0217] T proccssing For data processing time;

[0218] T queuc The message queue wait time;

[0219] The formula is:

[0220] T sync =T network +T proccssing +T qucuc

[0221] Among them, T network Calculated using network bandwidth and data volume:

[0222]

[0223] Where D is the data size and B is the network bandwidth;

[0224] T proccssing For the time required for data processing;

[0225] T qucuc Related to the queue length and throughput of streaming data, assuming a queue of Q messages and a system processing rate of R messages / second, then:

[0226]

[0227] Throughput is the amount of data successfully transmitted per unit of time, measured in data per second, and is defined as follows:

[0228] T throughput For throughput;

[0229] D is the size of each data block;

[0230] R is the number of data blocks processed per second;

[0231] The formula is:

[0232] T throughput =D×B

[0233] Where D is the size of the data block and R is the processing speed of the data block per second;

[0234] Streaming data synchronization efficiency is measured as the ratio of actual data transmission efficiency to processing efficiency, defined as:

[0235] E sync To improve data synchronization efficiency;

[0236] T sync This represents the total time for data synchronization.

[0237] T ideal This represents the ideal data synchronization time.

[0238] The formula is:

[0239]

[0240] When synchronizing data in real time, ensure the consistency and integrity of the data in the sandbox environment. Establish a data verification process using a checksum method, with the following formula:

[0241] H prod For checksums of data in the production environment;

[0242] H sandbox This is a checksum for the synchronized data in the sandbox environment.

[0243] The formulas for consistency and integrity verification are:

[0244] Valid Data = (H prod ==H sandbox )

[0245] When H prod and H sandbox When they are equal, the data is synchronized and kept consistent in the sandbox environment.

[0246] By employing streaming data transmission technology, data from the production environment can be synchronized to the sandbox environment almost in real time. To ensure the consistency and integrity of data between the production and sandbox environments, a checksum method is used to verify the accuracy of the synchronized data. By dynamically adjusting the data processing rate based on the relationship between queue length and message throughput, efficient and stable data synchronization to the sandbox environment is guaranteed. Distributed storage technology allows for flexible expansion of data processing capabilities; as the data volume increases, storage nodes can be dynamically added to maintain high availability and scalability of the system. Real-time checksum verification is performed during data synchronization to ensure consistency between the data in the sandbox and production environments. By optimizing various technical parameters during the data synchronization process, data transmission efficiency can be maximized, while reducing system resource consumption and operating costs.

[0247] In the sandbox environment, a multi-level quality governance rule engine is adopted. This engine automatically assigns different verification rules based on the characteristics of the data. The method by which the rule engine dynamically adjusts the verification rules according to real-time data changes is as follows:

[0248] The characteristics of the data are used to assign corresponding validation rules based on a predefined set of rules. The following variables are set:

[0249] D is the data characteristic vector, representing the various characteristics of the data, and is represented by a multi-dimensional vector:

[0250] D = [D1, D2, D3, ..., D] n ]

[0251] Among them, D i This represents the i-th dimension in the data characteristics;

[0252] R is the validation rule set, which contains various rules:

[0253] R = {R1, R2, R3, ..., R} m}

[0254] Among them, R i This is the i-th validation rule;

[0255] C(D i R j ) represents data characteristic D i With rule R j Fit, representing the degree of matching between a feature and a rule, is measured using a scoring function:

[0256] C(D i R j )=f(D i R j )

[0257] Where f is set according to the relationship between characteristics and rules;

[0258] The verification rule assignment process is completed using the following formula:

[0259] R assigned ={R j |C(D i R j )>θ}

[0260] Where θ is a threshold, when the data characteristics fit the rule C(D) i R j When the value is higher than the threshold, rule R j Assigned to data characteristic D i ;

[0261] As data changes in real time, the rule engine dynamically adjusts validation rules, involving real-time monitoring of data changes and adjustments based on new data characteristics or quality standards, establishing:

[0262] ΔD represents the amount of change in data over a specific time period;

[0263] ΔR represents the amount of change in the validation rules caused by changes in the data;

[0264] The formula for dynamically adjusting the verification rules is:

[0265] ΔR = g(ΔD, α, β)

[0266] Where g(·) is a function used to dynamically adjust the rule set based on the amount of data change and system settings;

[0267] α is a sensitivity parameter, representing how sensitive the rule adjustment is to changes in the data, and β is an adjustment threshold used to determine when to change the rule;

[0268] During the rule adjustment process, the following was established:

[0269] R new =R assigned ∪ΔR

[0270] Among them, R new For the new set of verification rules;

[0271] The rule engine adjusts future rule assignments based on the validation results, establishing:

[0272] V result (R j ) is rule R j The verification results;

[0273] ω(R j ) is rule R j The weight represents the importance of the rule in the overall verification process;

[0274] The adjustment formula for the rule feedback mechanism is:

[0275] ω new (R j )=ω(R j )×(1+λ×V result (R j ))

[0276] Where λ is the sensitivity factor for feedback adjustment;

[0277] In a multi-layered rule engine, rules are assigned and adjusted through a hierarchical decision tree, with the following variables set:

[0278] L represents the rule hierarchy;

[0279] P(L) represents the priority of level L;

[0280] The decision formula is:

[0281]

[0282] Where P(L) represents the priority of the rule hierarchy, and τ is the priority threshold, ensuring that only rules with high priority are selected into the final rule set R. final.

[0283] By automatically assigning validation rules based on data characteristics, the system can accurately apply the most suitable rules to each data type, reducing human intervention. The rule engine allocates the most suitable validation rules based on data characteristic vectors, improving the accuracy of data validation and avoiding unnecessary error checking and inefficient processing. The system's rule adjustment mechanism can dynamically change the rule set to adapt to different data quality requirements based on data variation and sensitivity parameters. The multi-level rule engine can make decisions based on rule priority settings, ensuring that high-priority rules always dominate validation, which helps ensure that critical data is always rigorously validated. By automatically assigning and adjusting validation rules based on data characteristics, the waste of computing resources can be significantly reduced, because each rule is adjusted and optimized according to the actual needs of the data. By dynamically adjusting validation rules, unnecessary validation operations are avoided, reducing system burden and latency, optimizing performance, and thus reducing operating costs.

[0284] The system monitors each piece of data synchronized to the sandbox in real time, automatically identifies abnormal patterns in the data based on an anomaly detection model, and detects potential data quality issues. Upon detecting an anomaly, the system automatically sends an alarm to the administrator, prompting relevant personnel to take appropriate action.

[0285] The system monitoring data stream is set as a time series or multidimensional dataset, where each data point D t It will undergo anomaly detection;

[0286] D t =[D t1 D t2 , ..., D tn ] represents the data collected at time t;

[0287] This represents an anomaly detection model that outputs whether an anomaly exists. It can be built using a statistical model, a machine learning model, or a deep learning model.

[0288] The anomaly detection method used is distance-based, employing either an isolated forest or a probability-based model, such as the Gaussian distribution assumption model. The anomaly detection formula is expressed as:

[0289]

[0290] Among them, A t For data D at time t t If the abnormal detection result of A t =1 indicates that there is an anomaly in the data. If A t =0 indicates that the data is normal;

[0291] If based on a statistical model:

[0292]

[0293] Among them, D t This is the current data point;

[0294] μ is the mean vector of the data;

[0295] Σ is the covariance matrix of the data;

[0296] P(D t |μ,Σ) represents the data D t The probability density;

[0297] For multidimensional or sequential data, anomaly patterns are sudden fluctuations in a certain feature or changes in the relationship between multiple features. Anomaly patterns can be identified by calculating the data's deviation or fluctuation.

[0298]

[0299] Among them, D t This refers to the observation data at the current moment;

[0300] These are predicted values ​​based on historical data.

[0301] σ(D t ) for data D t Standard deviation;

[0302] Data quality issues include missing values, duplicate values, and incorrect formatting. The detection of data quality issues is modeled using the following methods:

[0303] Missing value detection:

[0304]

[0305] Among them, M t Indicates whether there are missing values ​​at time t, 1 is the indicator function, if D ti If the value is missing, the value is 1; otherwise, the value is 0.

[0306] Duplicate value detection:

[0307]

[0308] Among them, R t This indicates whether there is duplicate data at time t. If the data items are equal, the value is 1; otherwise, it is 0.

[0309] Once an abnormal pattern or data quality issue is detected, the system will automatically issue an alarm. The alarm conditions are set as follows:

[0310] Alarmt =1(A t =1or M t >0 or R t >0)

[0311] If A t =1 or M t >0 or R t If the value is greater than 0, an alarm will be triggered;

[0312] Alarm priority is calculated based on the severity of the anomaly, with a weight w defined for each anomaly type. typc :

[0313] Priority t =w type ×Severity t

[0314] Among them, w type Weights for exception types, Severity t The severity score is used to indicate the abnormality.

[0315] The formula for generating alarm notifications is:

[0316] Notification t =GenerateNotification(Alarm) t Priority t This formula generates a notification sent to the relevant administrator, indicating the abnormal data and its priority level.

[0317] By monitoring every piece of data synchronized to the sandbox in real time, the system can promptly detect abnormal patterns and potential data quality issues. This solution supports multiple anomaly detection models, such as the distance-based isolated forest algorithm and the probability-based Gaussian distribution hypothesis model, enabling the system to handle different data types and anomaly patterns, enhancing its flexibility and robustness. The system not only detects abnormal patterns but also effectively identifies common data quality problems such as missing values, duplicate values, and incorrect formats. Based on the severity scores of different anomaly types, the system automatically adjusts alarm priorities to ensure that the most urgent anomalies are addressed promptly. Through automated monitoring and anomaly detection, the system can automatically complete data quality control tasks, reducing the need for manual intervention and inspection, thereby saving significant labor costs.

[0318] Based on changes in real-time data streams, anomaly detection results, and historical verification results, the system automatically adjusts its data verification strategy. When the system detects frequent occurrences of certain data quality issues, the verification strategy will increase its intensity; conversely, the system will automatically reduce the verification burden.

[0319] Define model variables:

[0320] D t : Represents the data stream collected at time t;

[0321] A t : Represents the anomaly detection result at time t;

[0322] M t : Indicates the frequency of data quality problems at time t;

[0323] V t : Indicates the data verification strength at time t;

[0324] H t : Represents the historical verification result at time t;

[0325] The system should calculate the frequency of anomalies based on anomaly detection results and the frequency of data quality issues, defining the anomaly frequency F at each time point. t for:

[0326]

[0327] A t M is the result of anomaly detection. t This formula represents the frequency of data quality issues. The system integrates the frequency of anomalies and data quality problems using this formula.

[0328] Based on historical verification results H t and the current abnormal frequency F t The system automatically adjusts its data validation strategy, employing a non-linear adjustment strategy that adjusts the validation strength V based on frequency and historical results. t ;

[0329] The formula for adjusting the verification strength is expressed as:

[0330] V t =V t 1 +α×(F t -F threshold )+β×H t

[0331] Among them, V t 1 The verification strength at the previous time step;

[0332] α is the abnormal frequency adjustment coefficient, which controls the impact of abnormal frequency on the verification strength;

[0333] F thrcshold This is the threshold for the frequency of anomalies, representing the minimum standard for frequent anomalies.

[0334] β is the adjustment factor for historical validation results;

[0335] H t This is a result verified historically;

[0336] In the system, the verification burden B t To verify the strength V t Measured, higher V t This indicates a high verification burden, and the system dynamically adjusts its approach by calculating the verification burden:

[0337] B t =γ×V t

[0338] Where γ is a constant coefficient.

[0339] When the frequency of anomalies increases or data quality issues occur frequently, the system automatically increases the verification intensity to ensure that these issues are fully checked. The system not only relies on current anomaly detection results but also adjusts the verification strategy based on historical verification results. The solution comprehensively assesses the quality of data by considering both anomaly detection results and the frequency of data quality issues. The system automatically adjusts the data verification strategy based on feedback from anomaly frequency and historical verification results, reducing the need for manual intervention. By combining historical results, anomaly frequency, and data quality issue frequency, the system can accurately identify potential quality issues in the data and reduce the risk of missed detections.

[0340] The system incorporates an automated monitoring and report generation module to monitor the data verification process in real time and periodically generate data quality reports for administrator review. These reports include statistical information on data quality issues, tracking records of abnormal events, and recommendations for targeted improvement measures.

[0341] Data quality issues include missing data, duplicate data, and format errors. The quantity and ratio of these issues are defined by the following formula.

[0342] Q t : Represents the total number of quality problems present in the data collected at time t;

[0343] Q total : Indicates the total amount of data in the dataset;

[0344] Data quality issue frequency P t Represented as:

[0345]

[0346] For each abnormal event, the system records the occurrence time, event type, and number of records affected;

[0347] E t : Indicates the number of abnormal events detected at time t;

[0348] E types : Indicates the type of abnormal event;

[0349] For each abnormal event e i Statistically analyze its occurrence frequency F i :

[0350]

[0351] The system should generate targeted improvement measures to address data quality issues and anomalies.

[0352] S t : Suggestions for improvement measures generated at time t

[0353] Each question type e i The corresponding suggested weight W i ;

[0354] The proposed improvement measures are generated using the following formula:

[0355]

[0356] n is the number of types of abnormal events, W i For each exception event type e i The weight of relevant improvement measures;

[0357] Monitor the real-time status of the data verification process, define real-time feedback indicators, and set up the system to generate a real-time monitoring report (R). t It includes statistical information on data quality issues and detailed tracking records of abnormal events;

[0358] Monitoring Report R t It consists of the following parts:

[0359] R t ={P t F1, F2, ..., F n S t}

[0360] Among them, P t F1, F2, ..., F are the ratios of data quality problems. n The frequency of occurrence of various abnormal events, S t Improvement suggestions based on abnormal events;

[0361] Regular reports should include a summary of data quality issues, trend analysis of abnormal events, and recommendations for targeted improvement measures. Regular reports should be generated periodically using the following formula:

[0362] The summary of data quality issues in periodic reports is based on the average data quality issue rate over a past period:

[0363]

[0364] Where T is the period length;

[0365] For each type of abnormal event, the system should track its changing trend and use a moving average method to smooth the frequency changes of abnormal events:

[0366]

[0367] Regular reports should also include a summary of improvement measures for each type of anomaly, calculated based on suggested weights and frequencies:

[0368]

[0369] The system can monitor the data verification process in real time, detect data quality issues and anomalies, and record relevant information. By statistically analyzing the frequency of different types of data quality issues and anomalies, the system helps managers clearly understand the source and impact of data problems. The system can regularly generate data quality reports, summarizing the average data quality issue rate over a period of time, trend analysis of anomalies, and corresponding improvement suggestions. Through automated monitoring and report generation, the need for manual intervention is reduced, lowering the risk of human error. The system automatically generates relevant improvement suggestions based on the type and frequency of each anomaly, with each anomaly having a corresponding suggestion weight. This allows for flexible adjustment of improvement measures based on actual conditions, ensuring that the suggestions are practically actionable.

[0370] When processing data, the system employs data anonymization and encryption technologies to ensure that data in the sandbox environment is not leaked during the verification process. Furthermore, the system includes inspection tools that automatically verify whether data usage in the sandbox environment complies with relevant data protection regulations.

[0371] Data anonymization involves modifying or hiding sensitive information in data. An original dataset D = {d1, d2, ..., d...} is established. n}, where each data point d i Contains sensitive information, which is de-identified through the operation Δ(d) i ), to obtain a de-identified dataset D′={d′1,d′2,...,d′ n}, where d′ i =Δ(d) i () indicates a desensitized version;

[0372] The de-identification operation Δ employs various methods for data field f. iDesensitization, the desensitization rules are defined as follows:

[0373] Δ(f i ) = mask(f i )

[0374] Encryption transforms sensitive data into unreadable ciphertext using encryption algorithms, ensuring that even if the data is leaked, it cannot be maliciously used. Let the original data be D = {d1, d2, ..., d...}. n The encrypted data is D′={e1, e2, ..., e}. n}, where e i =Encrypt(d i );

[0375] Encryption algorithms include symmetric encryption and asymmetric encryption, and the encryption formula is:

[0376] e i =Encrypt k (d i )

[0377] Where k is the encryption key, Encrypt k For encryption operations, output e i The data is encrypted;

[0378] Ensure that data usage in the sandbox environment complies with relevant data protection regulations, and verify compliance through automated inspection tools. Establish an inspection function to verify whether dataset D′ complies with relevant regulations.

[0379] The specific inspection items include:

[0380] Data access control: Confirms whether data access is restricted to authorized users;

[0381] Data processing purpose: To ensure that data is used only for legitimate purposes;

[0382] Data minimization: Ensure that the amount of data processed complies with regulations and avoid over-collection;

[0383] Define a compliance verification formula as follows:

[0384]

[0385] To further ensure data security, the system monitors and records all data access behaviors, assuming A = {a1, a2, ..., a...} m} represents all access records, where a i =(u i , t i d i ) represents user u i At time ti For data d i Access behavior;

[0386] The system should generate an access audit report. t And conduct compliance checks on data access:

[0387] L t ={a1, a2, ..., a m Verify compliance with access control policies by checking access logs:

[0388]

[0389] The system introduces a function to detect potential data breaches. By monitoring anonymized and encrypted data, it identifies data breach risks. The data breach warning trigger formula is as follows:

[0390]

[0391] Before entering the sandbox environment, the data undergoes desensitization and encryption processing, using the following formula:

[0392] D′={Δ(d1), Δ(d2),..., Δ(d n )}

[0393] Or encrypt:

[0394] D′={Encrypt k (d1), Encrypt k (d2), ..., Encrypt k (d n )}

[0395] Ensure data usage complies with relevant regulations through automated inspection tools:

[0396] CheckCompliant(D′)→True / False

[0397] All data access is logged and subject to compliance checks:

[0398] CheckAccess(L t → True / False

[0399] And monitor potential risks through a leak warning mechanism:

[0400] LeakPrevention(D′)→True / False.

[0401] Through encryption and anonymization, the system ensures that even if data is leaked, sensitive information will not be exposed. Automated inspection tools ensure that data usage in the sandbox environment complies with relevant data protection regulations. By verifying compliance in areas such as data access control, data processing purposes, and data minimization, the system ensures that all data operations are legal and compliant. The system ensures that data access is restricted to authorized users, preventing unauthorized access to sensitive data. By monitoring encrypted and anonymized data, the system can identify potential leakage risks and trigger leakage alerts. The system ensures that the amount of data processed complies with relevant regulations, avoiding excessive collection. The system monitors all encrypted and anonymized data and access behavior in real time, enabling timely detection of any potential risks. The system regularly generates access audit reports to check the compliance of each data access, helping administrators understand whether data operations are compliant.

[0402] This embodiment also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the quality governance and real-time data verification method based on the data sandbox as described above.

[0403] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the data sandbox-based quality governance and real-time data verification method as described above.

[0404] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0405] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0406] The above embodiments of the present invention are not intended to limit the scope of protection of the present invention. The implementation of the present invention is not limited thereto. All other modifications, substitutions or alterations made to the above structure of the present invention based on the above content of the present invention, in accordance with ordinary technical knowledge and common practice in the field, without departing from the basic technical idea of ​​the present invention, shall fall within the scope of protection of the present invention.

Claims

1. A method for quality governance and real-time data verification based on a data sandbox, characterized in that: Includes the following steps: A data sandbox environment is built, employing distributed storage technology and a real-time data synchronization mechanism to synchronize data from the production environment to the sandbox environment in real time. This synchronization process is achieved through streaming data transmission technology, ensuring high efficiency and low latency in the data synchronization process. In the sandbox environment, a multi-level quality governance rule engine is adopted. This engine automatically assigns different verification rules according to the characteristics of the data, and the rule engine dynamically adjusts the verification rules according to real-time data changes. The system monitors each piece of data synchronized to the sandbox in real time, automatically identifies abnormal patterns in the data based on the anomaly detection model, detects potential data quality problems, and automatically sends an alarm to the administrator after an anomaly is detected, prompting relevant personnel to handle the situation. Based on changes in real-time data streams, anomaly detection results, and historical verification results, the system automatically adjusts its data verification strategy. When the system detects that certain data quality issues occur frequently, the verification strategy will increase the verification intensity; conversely, the system will automatically reduce the verification burden. The system introduces an automated monitoring and report generation module to monitor the data verification process in real time and generate data quality reports for managers to review regularly. The report content includes statistical information on data quality issues, tracking records of abnormal events, and suggestions for targeted improvement measures. When processing data, the system uses data anonymization and encryption technologies to ensure that data in the sandbox environment is not leaked during the verification process. In addition, the system also includes inspection tools to automatically verify whether the use of data in the sandbox environment complies with relevant data protection regulations. In the sandbox environment, a multi-level quality governance rule engine is adopted. This engine automatically assigns different verification rules based on the characteristics of the data. The method by which the rule engine dynamically adjusts the verification rules according to real-time data changes is as follows: The characteristics of the data are used to assign corresponding validation rules based on a predefined set of rules. The following variables are set: The data characteristic vector represents the various characteristics of the data, and is represented by a multi-dimensional vector: ; in, The first in the data characteristics One dimension; The rule set for validation contains multiple rules: ; in, For the first Verification rules; For data characteristics With rules Fit, representing the degree of matching between a feature and a rule, is measured using a scoring function: ; in, To be set according to the relationship between characteristics and rules; The verification rule assignment process is completed using the following formula: ; in, A threshold is defined as the degree of fit between data characteristics and rules. When it exceeds the threshold, the rule Assigned to data characteristics ; As data changes in real time, the rule engine dynamically adjusts validation rules, involving real-time monitoring of data changes and adjustments based on new data characteristics or quality standards, establishing: This indicates the amount of change in data within a specific time period; This indicates the amount of change in the validation rules caused by changes in the data; The formula for dynamically adjusting the verification rules is: ; in, This is a function used to dynamically adjust the rule set based on data changes and system settings. This is a sensitivity parameter, representing the sensitivity of rule adjustments to changes in data. An adjustment threshold is used to determine when to change the rules; During the rule adjustment process, the following was established: ; in, For the new set of verification rules; The rule engine adjusts future rule assignments based on the validation results, establishing: For rules The verification results; For rules The weight represents the importance of the rule in the overall verification process; The adjustment formula for the rule feedback mechanism is: ; in, Sensitivity factor adjusted for feedback; In a multi-layered rule engine, rules are assigned and adjusted through a hierarchical decision tree, with the following variables set: Indicates the hierarchy of rules; hierarchical Priority; The decision formula is: ; in, Indicates the priority of the rule hierarchy. This is a priority threshold to ensure that only rules with high priority are selected into the final rule set. .

2. The method for quality governance and real-time data verification based on a data sandbox according to claim 1, characterized in that, A data sandbox environment is constructed, employing distributed storage technology and a real-time data synchronization mechanism to synchronize data from the production environment to the sandbox environment in real time. This synchronization process is achieved through streaming data transmission technology. The methods to ensure the high efficiency and low latency of the data synchronization process are as follows: Data transmission latency is expressed as the time required for data to travel from the source system to the target system, and is defined as follows: This represents the total delay time of the data synchronization process. This refers to network transmission time. For data processing time; The message queue wait time; The formula is: ; in, Calculated using network bandwidth and data volume: ; in, For data size, For network bandwidth; For the time required for data processing; Related to the queue length and throughput of streaming data, establish The system processes messages in a queue at a rate of [number]. messages / second, then: ; Throughput is the amount of data successfully transmitted per unit of time, measured in data per second, and is defined as follows: For throughput; Size of each data block; This refers to the number of data blocks processed per second. The formula is: ; in, For the size of the data block, This represents the processing speed of data blocks per second. Streaming data synchronization efficiency is measured as the ratio of actual data transmission efficiency to processing efficiency, defined as: To improve data synchronization efficiency; This represents the total time for data synchronization. This represents the ideal data synchronization time. The formula is: ; When synchronizing data in real time, ensure the consistency and integrity of the data in the sandbox environment. Establish a data verification process using a checksum method, with the following formula: For checksums of data in the production environment; For the checksum of synchronized data in the sandbox environment; The formulas for consistency and integrity verification are: ; when and When they are equal, the data is synchronized and kept consistent in the sandbox environment.

3. The method for quality governance and real-time data verification based on a data sandbox according to claim 1, characterized in that, The system monitors each piece of data synchronized to the sandbox in real time, automatically identifies abnormal patterns in the data based on an anomaly detection model, and detects potential data quality issues. Upon detecting an anomaly, the system automatically sends an alarm to the administrator, prompting relevant personnel to take appropriate action. The system monitoring data stream is set as a time series or multidimensional dataset, where each data point... It will undergo anomaly detection; , indicating at time The collected data; This represents an anomaly detection model that outputs whether an anomaly exists. It can be built using a statistical model, a machine learning model, or a deep learning model. The anomaly detection method is selected from one or more of the following methods: distance-based anomaly detection methods; isolated forest methods; probability-based models, including Gaussian distribution assumption models; The formula for anomaly detection is expressed as: ; in, For at any time Data If the abnormal detection results are... This indicates that there is an anomaly in the data. This indicates that the data is normal; The anomaly detection formula based on the Gaussian distribution assumption model is expressed as follows: ; in, This is the current data point; The mean vector of the data; The covariance matrix of the data; For data The probability density; For multidimensional or sequential data, anomaly patterns are sudden fluctuations in a certain feature or changes in the relationship between multiple features. Anomaly patterns can be identified by calculating the data's deviation or fluctuation. ; in, This refers to the observation data at the current moment; These are predicted values ​​based on historical data. For data Standard deviation; Data quality issues include missing values, duplicate values, and incorrect formatting. The detection of data quality issues is modeled using the following methods: Missing value detection: ; in, Indicates time Does the value exist? 1 is an indicator function. If the value is missing, the value is 1; otherwise, the value is 0. Duplicate value detection: ; in, Indicates time Check for duplicate data; if the data items are equal, return 1; otherwise, return 0. Once an abnormal pattern or data quality issue is detected, the system will automatically issue an alarm. The alarm conditions are set as follows: ; like or or If so, an alarm will be triggered; Alarm priority is calculated based on the severity of the anomaly, with a weight defined for each anomaly type. : ; in, Weights for abnormal types, The severity score is used to indicate the abnormality. The formula for generating alarm notifications is: ; This formula generates a notification sent to the relevant administrator, indicating the abnormal data and its priority level.

4. The method for quality governance and real-time data verification based on a data sandbox according to claim 3, characterized in that, Based on changes in real-time data streams, anomaly detection results, and historical verification results, the system automatically adjusts its data verification strategy. When the system detects frequent occurrences of certain data quality issues, the verification strategy will increase its intensity; conversely, the system will automatically reduce the verification burden. Define model variables: : Indicates time The collected data stream; : Indicates time Abnormal detection results; : Indicates time Frequency of data quality issues; : Indicates time Data validation strength; : Indicates time Historical verification results; The system should calculate the frequency of anomalies based on anomaly detection results and the frequency of data quality issues, and define the anomaly frequency at each time point. for: ; For a moment Abnormal detection results For a moment The frequency of data quality issues, among which, The values ​​have been normalized to the range of [0,1]. Using this formula, the system makes a comprehensive judgment based on the frequency of anomalies and data quality issues. Based on historical verification results and current anomaly frequency The system automatically adjusts its data verification strategy, employing a non-linear adjustment strategy that adjusts the verification intensity based on frequency and historical results. ; The formula for adjusting the verification strength is expressed as: ; in, The verification strength at the previous time step; This is the abnormal frequency adjustment coefficient, which controls the impact of abnormal frequencies on the verification strength. This is the threshold for the frequency of anomalies, representing the minimum standard for frequent anomalies. Adjust the coefficients based on historical verification results; This is a historical verification result; In the system, the verification burden To determine the strength of the verification Measurement, higher This indicates a high verification burden, and the system dynamically adjusts its approach by calculating the verification burden: in, It is a constant coefficient.

5. The method for quality governance and real-time data verification based on a data sandbox according to claim 4, characterized in that, The system incorporates an automated monitoring and report generation module to monitor the data verification process in real time and periodically generate data quality reports for administrator review. These reports include statistical information on data quality issues, tracking records of abnormal events, and recommendations for targeted improvement measures. Data quality issues include missing data, duplicate data, and format errors. The quantity and ratio of these issues are defined by the following formula. : Indicates time The total number of quality issues present in the collected data; : Indicates the total amount of data in the dataset; Data quality issue frequency Represented as: ; For each abnormal event, the system records the occurrence time, event type, and number of records affected; : Indicates time The number of abnormal events detected; : Indicates the type of abnormal event; For each type of abnormal event Statistical analysis of its frequency of occurrence : ; The system should generate targeted improvement measures to address data quality issues and anomalies. : Indicates time The generated improvement suggestions; Each question type Corresponding suggested weights ; The proposed improvement measures are generated using the following formula: ; This represents the number of types of abnormal events. For each exception event type The weight of relevant improvement measures An abnormal event The frequency of occurrence; For improvements at time t; Monitor the real-time status of the data verification process, define real-time feedback indicators, and set up the system to generate a monitoring report in real time. It includes statistical information on data quality issues and detailed tracking records of abnormal events; Monitoring Report It consists of the following parts: in, The ratio of data quality issues. , ,..., The frequency of occurrence of various abnormal events, Improvement suggestions based on abnormal events; Regular reports should include a summary of data quality issues, trend analysis of abnormal events, and recommendations for targeted improvement measures. Regular reports should be generated periodically using the following formula: The summary of data quality issues in periodic reports is based on the average data quality issue rate over a past period: in, The period length; For each type of abnormal event, the system should track its changing trend and use a moving average method to smooth the frequency changes of abnormal events: Regular reports should also include a summary of improvement measures for each type of anomaly, calculated based on suggested weights and frequencies: 。 6. The method for quality governance and real-time data verification based on a data sandbox according to claim 5, characterized in that, When processing data, the system employs data anonymization and encryption technologies to ensure that data in the sandbox environment is not leaked during the verification process. Furthermore, the system includes inspection tools that automatically verify whether data usage in the sandbox environment complies with relevant data protection regulations. Data anonymization involves modifying or hiding sensitive information in the data to create an original dataset. Each data point Contains sensitive information, which is then de-identified. Obtain a de-identified dataset ,in This indicates a desensitized version; Desensitization procedure Various methods are used for data fields Desensitization, the desensitization rules are defined as follows: Encryption transforms sensitive data into unreadable ciphertext using encryption algorithms, ensuring that even if the data is leaked, it cannot be used maliciously. Let the original data be... The encrypted data is ,in = ; Encryption algorithms include symmetric encryption and asymmetric encryption, and the encryption formula is: = in For encryption key, For encryption operations, output The data is encrypted; Ensure that data usage in the sandbox environment complies with relevant data protection regulations, perform compliance verification through automated inspection tools, and establish an inspection function to verify the dataset. Does it comply with relevant regulations? The specific inspection items include: Data access control: Confirms whether data access is restricted to authorized users; Data processing purpose: To ensure that data is used only for legitimate purposes; Data minimization: Ensure that the amount of data processed complies with regulations and avoid over-collection; Define a compliance verification formula as follows: To further ensure data security, the system monitors and records all data access activities and establishes... For all access records, among which Indicates user In time Data Access behavior; The system should generate an access audit report. And conduct compliance checks on data access: Verify compliance with access control policies by checking access logs: The system introduces a function to detect potential data breaches. By monitoring anonymized and encrypted data, it identifies data breach risks. The data breach warning trigger formula is as follows: Before entering the sandbox environment, the data undergoes desensitization and encryption processing, using the following formula: Or encrypt: Ensure data usage complies with relevant regulations through automated inspection tools: All data access is logged and subject to compliance checks: And monitor potential risks through a leak warning mechanism: 。 7. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the data sandbox-based quality governance and real-time data verification method as described in any one of claims 1-6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the data sandbox-based quality governance and real-time data verification method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Communication information security risk early warning management and control method and system based on big data

    CN117955712A

  • Optimization system of data security sandbox

    CN119004446A

  • Intelligent market supervision data management system and method based on multi-stage data sharing and exchange

    CN120013333A