Problem data processing method and device, computer equipment and storage medium
By extracting data from multiple data sources, integrating it, locating anomalies, and verifying rules, it solves the shortcomings of problem data identification and repair in traditional data governance methods, achieves accurate identification and automatic repair, and improves data quality and the accuracy of business decisions.
Patent Information
- Application Number
- CN202510594745.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-10-03
AI Technical Summary
Traditional data governance methods are unable to accurately identify and repair problematic data when faced with diverse data scenarios with varying formats and quality, leading to data silos and business decision deviations, and increasing financial risks.
By extracting business data from multiple data sources and performing data integration processing, clustering and classification algorithms are used to locate anomalies, combined with the rule engine to filter problem data, and rectification scripts are constructed for automatic repair.
It achieves accurate identification and automatic repair of problematic data, improves data quality, reduces financial risks, and enhances the accuracy of business decisions.
Smart Images

Figure CN120743891A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology and can be applied to fields such as financial technology and digital medicine, and in particular to methods, devices, computer equipment and storage media for processing problem data. Background Art
[0002] Against the backdrop of the rapid development of digital business, the amount of data reported by various business systems is increasing day by day, and data governance issues have gradually become a core bottleneck restricting the efficient operation of enterprises. Traditional data governance systems rely on manual rules and simple verification logic, which makes it difficult to cope with complex scenarios with diversified data sources, uneven formats and quality, resulting in significantly insufficient accuracy and efficiency in identifying and repairing problem data. Specifically, traditional methods are usually based on fixed rules (for data quality testing, and lack in-depth analysis of data semantics, business logic and cross-system relationships. This shallow verification model cannot dynamically adapt to the dynamic changes of data sources, resulting in a large amount of potential problem data being missed, further exacerbating data silos and business decision-making deviations.
[0003] For example, in credit approval scenarios in the financial sector, traditional data governance methods may only perform basic field verification on the application forms submitted by customers (such as the format of the ID number and the range of income amounts), without in-depth analysis of the business logic relationships between the data. If a customer has multiple loan records at the same time, traditional methods may miss risk signals of excessive total debt-to-income ratios due to a lack of cross-system data correlation capabilities; or due to the failure to identify unstructured data, high-risk customers may be mistakenly judged as low-risk. This governance shortcoming may not only lead to credit default risks, but may also cause model distortion due to the influx of problematic data into risk control models, further exacerbating asset losses for financial institutions.
[0004] Therefore, there is an urgent need to build a dynamic data governance system based on intelligent algorithms and multi-source data integration to achieve accurate identification and automatic repair of problem data, thereby improving data quality, ensuring the accuracy of business decisions, and reducing the risk losses caused by data problems. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to propose a method, apparatus, computer equipment and storage medium for processing problem data to solve the technical problem of insufficient accuracy and efficiency in existing problem data identification and repair.
[0006] In a first aspect, a method for processing problem data is provided, comprising:
[0007] Extract business data corresponding to business needs from multiple preset data sources;
[0008] Performing data integration processing on the business data to obtain corresponding target business data;
[0009] Based on the preset clustering algorithm and classification algorithm, the target business data is anomaly located to obtain corresponding abnormal data;
[0010] Perform rule verification on the abnormal data based on a preset rule engine, and filter out problematic data that does not comply with the rules from the abnormal data;
[0011] Obtaining a rule number corresponding to the problem data, and constructing a rectification script corresponding to the problem data based on the rule number;
[0012] A rectification process corresponding to the problem data is executed based on the rectification script.
[0013] In a second aspect, a device for processing problem data is provided, comprising:
[0014] The extraction module is used to extract business data corresponding to business needs from multiple preset data sources;
[0015] An integration module, configured to perform data integration processing on the business data to obtain corresponding target business data;
[0016] A positioning module is used to locate abnormalities in the target business data based on a preset clustering algorithm and classification algorithm to obtain corresponding abnormal data;
[0017] A verification module is used to perform rule verification on the abnormal data based on a preset rule engine, and filter out problematic data that does not comply with the rules from the abnormal data;
[0018] A construction module, configured to obtain a rule number corresponding to the problem data, and construct a rectification script corresponding to the problem data based on the rule number;
[0019] A processing module is used to perform rectification processing corresponding to the problem data based on the rectification script.
[0020] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method for processing the above-mentioned problem data when executing the computer program.
[0021] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for processing the above-mentioned problem data are implemented.
[0022] The solution implemented by the above-mentioned problem data processing method, device, computer equipment and storage medium first extracts business data corresponding to business needs from multiple preset data sources; then performs data integration processing on the business data to obtain corresponding target business data; then, based on the preset clustering algorithm and classification algorithm, anomalies of the target business data are located to obtain corresponding abnormal data; subsequently, based on the preset rule engine, rule verification is performed on the abnormal data to filter out problem data that does not comply with the rules from the abnormal data; further, the rule number corresponding to the problem data is obtained, and a rectification script corresponding to the problem data is constructed based on the rule number; finally, the rectification processing corresponding to the problem data is executed based on the rectification script. This application extracts business data corresponding to business needs from multiple preset data sources, performs data integration processing on the business data to obtain target business data, and then locates anomalies in the target business data based on the combined use of clustering algorithms and classification algorithms to obtain corresponding anomaly data, and then performs rule verification on the anomaly data based on the use of a rule engine, filters out problem data that does not comply with the rules from the anomaly data, and then constructs a rectification script corresponding to the problem data based on the rule number corresponding to the obtained problem data, and finally executes the rectification processing corresponding to the problem data based on the rectification script. Through the above-mentioned data processing process, the dynamic data governance method based on intelligent algorithms and multi-source data fusion provided by this application can quickly and intelligently realize the accurate identification and automatic repair of problem data, improve the recognition efficiency and repair accuracy of existing problem data, and thus help improve data quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;
[0025] Figure 2 is a flow chart of an embodiment of a method for processing question data according to the present application;
[0026] Figure 3 is a structural diagram of an embodiment of a device for processing question data according to the present application;
[0027] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0029] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0030] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0031] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0032] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0033] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.
[0034] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .
[0035] It should be noted that the method for processing problem data provided in the embodiment of the present application is generally executed by a server / terminal device, and accordingly, the device for processing problem data is generally set in the server / terminal device.
[0036] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0037] Continue to refer Figure 2 , shows a flowchart of an embodiment of the method for processing problem data according to the present application. According to different needs, the order of the steps in the flowchart can be changed, and some steps can be omitted. The method for processing problem data provided in the embodiment of the present application can be applied to any scenario that requires data governance, and the method for processing problem data can be applied to products in these scenarios, for example, data governance in the financial field and the medical field. The method for processing problem data includes the following steps:
[0038] Step S201 : extracting business data corresponding to business requirements from a plurality of preset data sources.
[0039] In this embodiment, the method for processing problem data is executed on the electronic device (eg Figure 1The server / terminal device shown in the figure) can obtain business data corresponding to business needs through a wired connection or a wireless connection. It should be pointed out that the above-mentioned wireless connection methods may include but are not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultrasound) connection, and other wireless connection methods currently known or to be developed in the future. The executive subject of this application is specifically a problem handling system, or a data governance system, which can be referred to as a system for short. Problem handling scenarios in data governance business in the financial and medical fields. For example, in the financial field, the above-mentioned business data may include the customer account table in the insurance core system (fields: customer ID, account balance, account opening date), the bank's credit approval report (PDF format, including customer credit score, loan purpose description), and the securities company's real-time transaction data (Kafka stream, fields: transaction time, amount, counterparty ID). In the medical field, the above business data may include: electronic medical records in the hospital system (fields: patient ID, diagnosis results, operation time), CT imaging reports (DICOM format, including lesion location and size), wearable device data (MQTT protocol, fields: heart rate, blood pressure, timestamp), etc.
[0040] The specific implementation process of extracting business data corresponding to business needs from the preset multiple data sources will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0041] Step S202: performing data integration processing on the business data to obtain corresponding target business data.
[0042] In this embodiment, the specific implementation process of performing data integration processing on the business data to obtain the corresponding target business data will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0043] Step S203 : locating abnormalities in the target business data based on a preset clustering algorithm and classification algorithm to obtain corresponding abnormal data.
[0044] In this embodiment, a modular intelligent data analysis engine is pre-designed in the middle core processing layer of the system. The intelligent data analysis engine is scalable and can integrate a variety of machine learning algorithm libraries. Specifically, algorithm library integration: by introducing mature machine learning algorithm libraries, such as clustering algorithms (K-means, hierarchical clustering, etc.) and classification algorithms (decision trees, random forests, etc.). Ensure that the algorithm library can efficiently process large-scale data and provide a friendly interface for the engine to call. Engine function implementation: realize data loading, preprocessing, algorithm calling, result analysis and other functions. The engine should be able to automatically identify the data type, select the appropriate algorithm for analysis, and generate an analysis report.
[0045] Among them, the specific implementation process of locating anomalies in the target business data based on the preset clustering algorithm and classification algorithm to obtain corresponding abnormal data will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0046] Step S204: performing rule verification on the abnormal data based on a preset rule engine, and filtering out problematic data that does not comply with the rules from the abnormal data.
[0047] In this embodiment, the rule engine is pre-built into the core processing layer of the system. The specific implementation process of the preset rule engine performing rule verification on the abnormal data and filtering out problematic data that does not conform to the rules from the abnormal data will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0048] Step S205: Obtain a rule number corresponding to the problem data, and construct a rectification script corresponding to the problem data based on the rule number.
[0049] In this embodiment, when filtering out problematic data from the abnormal data, the problematic data is marked and the specific rule number and violation details are recorded. The specific implementation process of constructing a remediation script corresponding to the problematic data based on the rule number will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0050] Step S206: executing rectification processing corresponding to the problem data based on the rectification script.
[0051] In this embodiment, the above-mentioned rectification script can be executed to automatically perform rectification processing on the problem data identified from the abnormal data, thereby completing accurate repair of the problem data.
[0052] The present application first extracts business data corresponding to business needs from multiple preset data sources; then performs data integration processing on the business data to obtain corresponding target business data; then performs anomaly location on the target business data based on a preset clustering algorithm and a classification algorithm to obtain corresponding anomaly data; subsequently, performs rule verification on the anomaly data based on a preset rule engine to filter out problem data that does not conform to the rules from the anomaly data; further obtains a rule number corresponding to the problem data, and constructs a rectification script corresponding to the problem data based on the rule number; finally, executes rectification processing corresponding to the problem data based on the rectification script. The present application extracts business data corresponding to business needs from multiple preset data sources, and performs data integration processing on the business data to obtain target business data, then performs anomaly location on the target business data based on the combination of a clustering algorithm and a classification algorithm to obtain corresponding anomaly data, then performs rule verification on the anomaly data based on the use of a rule engine to filter out problem data that does not conform to the rules from the anomaly data, subsequently constructs a rectification script corresponding to the problem data based on the rule number corresponding to the obtained problem data, and finally executes rectification processing corresponding to the problem data based on the rectification script. Through the above-mentioned data processing process, the dynamic data governance method based on intelligent algorithms and multi-source data fusion provided by this application can quickly and intelligently realize the accurate identification and automatic repair of problem data, thereby improving the identification efficiency and repair accuracy of existing problem data, and thus helping to improve data quality.
[0053] In some optional implementations, step S201 includes the following steps:
[0054] Acquire update frequency information of each of the data sources.
[0055] In this embodiment, the data source at least includes a structured database, an unstructured file library, a streaming data interface, etc. The update frequency information refers to the frequency of data updates of the data source.
[0056] Based on the update frequency information, all the data sources are divided into a first data source that is updated frequently and a second data source that is updated infrequently.
[0057] In this embodiment, for any data source, the update frequency information of the data source is compared with a preset frequency threshold. If the update frequency is greater than or equal to the frequency threshold, the data source is determined to be frequently updated, and the data source is classified as a first data source with frequent updates. If the update frequency is less than the frequency threshold, the data source is determined to be infrequently updated, and the data source is classified as a second data source with infrequent updates. The value of the frequency threshold is not specifically limited and can be set based on actual business needs.
[0058] Based on a preset incremental extraction strategy, first business data corresponding to the business requirement is extracted from the first data source.
[0059] In this embodiment, full or incremental extraction can be selected based on business needs. For data sources with large volumes of data and frequent updates, incremental extraction is used to extract only the data that has been added or modified since the last extraction. For example, for an insurance company's core business system, incremental extraction is performed daily to obtain only the transaction data for that day.
[0060] Based on a preset full-data extraction strategy, second business data corresponding to the business requirement is extracted from the second data source.
[0061] In this embodiment, for data sources with infrequent data updates, full extraction is used to obtain all data at one time; for example, for some configuration information tables of an insurance company, a full extraction is set to be performed once a month.
[0062] The first service data and the second service data are combined to obtain corresponding combined data.
[0063] In this embodiment, the first business data and the second business data extracted from different sources and formats are unified in data format (such as date format, numerical format, etc.), and then integrated into the data storage area of the system to obtain combined data, thereby obtaining the final business data.
[0064] The combined data is used as the service data.
[0065] The present application obtains the update frequency information of each of the data sources; then based on the update frequency information, all the data sources are divided into a first data source with frequent updates and a second data source with infrequent updates; then, based on a preset incremental extraction strategy, the first business data corresponding to the business needs is extracted from the first data source; and based on a preset full extraction strategy, the second business data corresponding to the business needs is extracted from the second data source; subsequently, the first business data and the second business data are combined to obtain corresponding combined data, and the combined data is used as the business data. The present application obtains the update frequency information of each of the data sources, then, based on the update frequency information of each of the data sources, all the data sources are divided into a first data source with frequent updates and a second data source with infrequent updates; then, based on the use of the incremental extraction strategy, the first business data corresponding to the business needs is extracted from the first data source; and based on the use of the full extraction strategy, the second business data corresponding to the business needs is extracted from the second data source, and then the first business data and the second business data are combined, so that the business data that matches the business needs can be accurately and completely extracted from multiple data sources, ensuring the accuracy and completeness of the obtained business data.
[0066] In some optional implementations of this embodiment, step S202 includes the following steps:
[0067] The business data is preprocessed to obtain corresponding first processed data.
[0068] In this embodiment, the preprocessing may include removing noise data and missing values to ensure data quality, so that the generated first processed data can be correctly identified and processed by the intelligent mapping algorithm.
[0069] Call the preset intelligent mapping algorithm.
[0070] In this embodiment, the master data system is designed in advance according to industry norms and enterprise business needs. The scope of master data is clarified, such as customer information, product information, supplier information, etc. Detailed standards are defined for each master data, including data name, data definition, data format, value range, data source, etc. For example, for customer information master data, the customer number is defined as a string type with a length of 10 digits and a value range of numbers that comply with specific coding rules; the customer name is a string type with a length of no more than 50 characters, etc. Then, a reference data system is designed, including reference data such as interest rates, exchange rates, and region codes. Data standards are established for each reference data, such as the expression method of interest rates (annual interest rate, monthly interest rate, etc.), the source and update frequency of exchange rates, and the coding rules of region codes.
[0071] Subsequently, intelligent mapping algorithms are constructed based on the established data standards. Specifically, an intelligent rule engine is built to store and manage mapping rules. This specified rule engine can automatically select appropriate mapping rules based on the characteristics of the data source and the requirements of the data standard. Suitable intelligent mapping algorithms include rule-based mapping algorithms and machine learning-based mapping algorithms. Rule-based mapping algorithms map data according to predefined rules, while machine learning-based mapping algorithms can automatically discover mapping relationships between data by learning from a large number of data samples.
[0072] The first processed data is mapped and converted based on the intelligent mapping algorithm to obtain corresponding second processed data.
[0073] In this embodiment, the intelligent mapping algorithm is used to convert data from different sources into standardized data. For example, the formats of customer names in different systems are unified, and the expression of interest rates in different regions is converted.
[0074] Data verification is performed on the second processed data.
[0075] In this embodiment, the data verification refers to verifying the mapped and transformed data (i.e., the second processed data) to check whether it meets the requirements of the data standard. If the second processed data is detected to meet the requirements of the data standard, the second processed data is determined to have passed the data verification. If the second processed data is detected to not meet the requirements of the data standard, the second processed data is determined to have failed the data verification.
[0076] If the second processed data passes the data verification, the second processed data is used as the target business data.
[0077] In this embodiment, if it is detected that the second processed data fails the data verification, that is, it is detected that data that does not meet the requirements exists in the second processed data, adjustments and corrections are made in a timely manner.
[0078] This application obtains the corresponding first processed data by pre-processing the business data; then calls a preset intelligent mapping algorithm; then performs mapping and conversion processing on the first processed data based on the intelligent mapping algorithm to obtain the corresponding second processed data; subsequently performs data verification on the second processed data; if the second processed data passes the data verification, the second processed data is used as the target business data. This application obtains the corresponding first processed data by pre-processing the business data; then performs mapping and conversion processing on the first processed data based on the use of an intelligent mapping algorithm to obtain the corresponding second processed data, and uses the second processed data as the target business data when it is detected that the second processed data passes the data verification, thereby automatically and accurately completing the data integration processing of the business data, and effectively ensuring the compliance of the obtained target business data.
[0079] In some optional implementations, step S203 includes the following steps:
[0080] Feature extraction is performed on the target business data to obtain corresponding feature data.
[0081] In this embodiment, the feature extraction includes numerical feature extraction and categorical feature extraction. Specifically, numerical feature extraction includes calculating statistical features such as the mean, variance, maximum, and minimum values of the data. Categorical feature extraction includes analyzing categorical features such as the type and distribution of the data.
[0082] The characteristic data is clustered based on the clustering algorithm to generate corresponding data clusters.
[0083] In this embodiment, the characteristic data can be clustered according to the selected clustering algorithm to cluster similar data points together to form different data clusters. The distribution and potential patterns of the data can be understood by analyzing the characteristics of each data cluster.
[0084] The feature data is classified based on the classification algorithm to generate corresponding data categories.
[0085] In this embodiment, the characteristic data may be classified according to a selected classification algorithm and divided into different data categories, for example, customer transaction data may be divided into categories such as high-value customers and low-value customers.
[0086] Get the preset feature model of normal data.
[0087] In this embodiment, a feature model of normal data is pre-established based on historical data and normal business patterns.
[0088] The data clusters and the data categories are respectively compared with the feature model to determine the designated data that deviates from a preset normal range.
[0089] In this embodiment, by analyzing the characteristics of the aforementioned data clusters and data categories and comparing them with the characteristic model of normal data, data points or data clusters that deviate from the normal range can be identified and used as the final abnormal data. The detected abnormal data can be marked, and information such as the type and severity of the abnormality can be recorded. For example, in the case of customer transaction data, if the transaction amount of a customer group significantly deviates from the normal range, the transaction data of this customer group will be marked as abnormal data.
[0090] The designated data is regarded as the abnormal data.
[0091] The present application obtains corresponding feature data by extracting features from the target business data; then clustering the feature data based on the clustering algorithm to generate corresponding data clusters; and classifying the feature data based on the classification algorithm to generate corresponding data categories; then obtaining a preset feature model of normal data; subsequently comparing the data clusters and the data categories with the feature model to determine the specified data that deviates from the preset normal range; and finally using the specified data as the abnormal data. The present application obtains corresponding feature data by extracting features from the target business data; then clustering the feature data based on the clustering algorithm to generate corresponding data clusters; and classifying the feature data based on the classification algorithm to generate corresponding data categories, and then comparing the data clusters and the data categories with the feature model to accurately determine the abnormal data that deviates from the preset normal range, thereby improving the accuracy of abnormal location of the target business data, and providing important data support for subsequent data verification, and helping to determine the data range that needs to be focused on verification.
[0092] In some optional implementations, step S204 includes the following steps:
[0093] Call the preset rule library.
[0094] In this embodiment, the rule base is a pre-built database for classifying, storing and managing pre-built verification rules including data quality rules and compliance rules, and provides the functions of adding, deleting, modifying and checking rules to facilitate users to maintain and update the rules.
[0095] Among them, in the rule engine of the core processing layer of the system, data quality rules and compliance rules are predefined according to actual business needs. Data quality rules: including data integrity rules (such as fields cannot be empty), data accuracy rules (such as numerical range restrictions), data consistency rules (such as logical relationships between data), etc. Compliance rules: including regulatory requirements rules (such as anti-money laundering rules in the financial industry), industry norms rules (such as data security standards), etc. In addition, by using machine learning algorithms to analyze historical data and governance tasks, the potential laws and patterns in the data are extracted, and new rules are generated through self-learning. For example, by analyzing historical transaction data, it is found that certain transaction patterns have risks, and corresponding risk control rules are automatically generated. Furthermore, a rule library is recommended to classify, store and manage all the constructed rules.
[0096] Verification rules are obtained from the rule base; wherein the verification rules include data quality rules and compliance rules.
[0097] In this embodiment, the stored verification rules can be retrieved from the rule library, and the verification rules at least include data quality rules and compliance rules.
[0098] The abnormal data is matched with the verification rules based on the rule engine to detect whether the abnormal data meets the rule requirements of the verification rules and obtain a corresponding detection result.
[0099] In this embodiment, the rule engine matches the collected abnormal data with the verification rules in the rule base to check whether the abnormal data meets the rule requirements. A test result is then generated based on the matching results. The test result includes information such as the verified data range, the number of rules verified, the amount of data that meets the rules, and the amount of data that does not meet the rules.
[0100] The detection results are analyzed to filter out designated abnormal data that does not meet the verification rules from the abnormal data.
[0101] In this embodiment, the above-mentioned detection results can be analyzed, and the data that does not meet the above-mentioned verification rules can be marked to obtain the above-mentioned specified abnormal data, and the specific rule number and violation details of the violation can be recorded at the same time.
[0102] The designated abnormal data is used as the problem data.
[0103] In this embodiment, after obtaining the problem data, it can be further classified and graded based on the severity of the violation and the scope of impact. For example, problems can be categorized as serious, general, and minor, with each assigned a priority. Furthermore, the marked problems are stored in a problem database to facilitate subsequent query and analysis.
[0104] The present application obtains verification rules based on the use of a preset rule base; then obtains verification rules from the rule base; wherein the verification rules include data quality rules and compliance rules; then matches the abnormal data with the verification rules based on the rule engine to detect whether the abnormal data meets the rule requirements of the verification rules and obtains corresponding detection results; subsequently analyzes the detection results to filter out the specified abnormal data that does not meet the verification rules from the abnormal data; finally uses the specified abnormal data as the problem data. The present application obtains verification rules based on the use of a rule base; then matches the abnormal data with the verification rules based on the use of a rule engine to detect whether the abnormal data meets the rule requirements of the verification rules and obtains corresponding detection results, and then analyzes the detection results to filter out the specified abnormal data that does not meet the verification rules from the abnormal data and uses them as the required problem data, thereby realizing efficient and accurate screening of problem data from abnormal data, improving the intelligent screening of problem data, and ensuring the accuracy of the obtained problem data, so as to provide accurate data basis for subsequent problem rectification.
[0105] In some optional implementations of this embodiment, step S205 includes the following steps:
[0106] Call the preset script template library.
[0107] In this embodiment, the script template library is a pre-built set of rectification script templates (referred to as script templates) containing various common data problems. Each template corresponds to a specific problem type and rule number.
[0108] A designated script template matching the rule number is retrieved from the script template library.
[0109] In this embodiment, the key features of the problem data are obtained, and then the most appropriate correction script template is matched from the script template library based on the rule number and the key features of the problem data. For example, if a data accuracy rule is violated and the problem data is concentrated in a specific field, the corresponding field correction script template is selected.
[0110] Get detailed information about the problem data.
[0111] In this embodiment, the above-mentioned detailed information refers to the specific information of the problematic data, including the data table name, field name, violation value, etc.
[0112] Fill the detailed information into the corresponding position in the specified script template to obtain the filled specified script template.
[0113] In this embodiment, a customized remediation script can be generated by filling in the detailed information of the problem data into a matching designated script template. After the remediation script is generated, a script preview function is provided, allowing users to view the specific content of the remediation script. If the remediation script is found to not meet actual needs, it can be manually adjusted.
[0114] The filled designated script template is used as the rectification script corresponding to the problem data.
[0115] In this embodiment, the generated rectification script can be performance analyzed to check its execution efficiency. If the rectification script is found to be slow, it can be optimized, such as by reducing unnecessary queries and optimizing the order of data operations. Furthermore, the rectification script can be optimized for readability to make it easier for developers to understand and maintain. For example, comments can be added and meaningful variable names can be used.
[0116] In addition, the generated rectification script can also be audited. Specifically, professional auditors are organized to conduct a logical audit of the rectification script to check whether the logic of the rectification script is correct and whether there are loopholes or errors. Auditors need to carefully analyze each line of code in the script to ensure that the script can accurately achieve the rectification goals. And evaluate the impact of the rectification script on other data. When executing the rectification script, multiple data tables and fields may be involved, and it is necessary to ensure that the script does not damage other normal data. For example, when updating the value of a field, it is necessary to check whether the field has an association with other fields to avoid data inconsistencies. Subsequently, the auditor will feedback the audit results to the script generator, and the script generator will modify and improve the rectification script based on the feedback until the rectification script passes the review.
[0117] This application calls a preset script template library; then queries the specified script template that matches the rule number from the script template library; then obtains the detailed information of the problem data; subsequently fills the detailed information into the corresponding position in the specified script template to obtain the filled-in specified script template; finally, uses the filled-in specified script template as the rectification script corresponding to the problem data. This application queries the specified script template that matches the rule number of the problem data based on the use of the script template library, and then fills the obtained detailed information of the problem data into the corresponding position in the specified script template, thereby achieving efficient and intelligent construction of the required rectification script, effectively improving the construction efficiency and accuracy of the rectification script. In addition, since the rectification script records the entire process of data rectification, including information such as problem data characteristics, rectification methods, and executors, in the subsequent process tracing, the rectification script can be used to quickly locate the root cause of the problem, evaluate the rectification effect, and provide support for the continuous improvement of data governance.
[0118] In some optional implementations of this embodiment, after step S206, the electronic device may further perform the following steps:
[0119] Obtaining a processing feedback result corresponding to the problem data.
[0120] In this embodiment, a rectification process corresponding to the problem data is executed based on the rectification script, and a corresponding processing feedback result is generated after the rectification process is completed.
[0121] A corresponding processing evaluation report is generated based on the processing feedback result.
[0122] In this embodiment, based on the aforementioned processing feedback results, a processing evaluation report can be automatically generated, including the execution status and completion time of the rectification task, the quality status of the problem data (such as data accuracy, completeness, consistency, and other indicators). The generated processing evaluation report is a summary and feedback of the entire data governance process for problem data rectification. The evaluation results can be used to understand the effectiveness of problem data rectification and provide improvement directions for subsequent steps such as data collection and data analysis.
[0123] Call the preset visual interface.
[0124] In this embodiment, a visualization interface can be designed in the top-level management and display layer of the system for user interaction. The visualization interface can display data asset maps, processing assessment reports, governance process progress and other information in the form of charts, reports, maps, etc.
[0125] The processing evaluation report is displayed based on the visual interface.
[0126] In this embodiment, the processing evaluation report can be displayed on the above-mentioned visual interface in the form of charts, reports, maps, etc.
[0127] This application obtains the processing feedback results corresponding to the problem data; then generates a corresponding processing evaluation report based on the processing feedback results; then calls a preset visual interface; and subsequently displays and processes the processing evaluation report based on the visual interface. After executing the rectification processing corresponding to the problem data based on the rectification script, this application will automatically obtain the processing feedback results corresponding to the problem data, and generate a corresponding processing evaluation report based on the processing feedback results, thereby realizing the automatic generation of the processing evaluation report and improving the generation efficiency of the processing evaluation report. In addition, the processing evaluation report will be intelligently displayed based on the use of the visual interface, so that relevant personnel can understand the rectification effect of the problem data by consulting the processing evaluation report, and can provide improvement directions for the subsequent data collection and data analysis steps, which is conducive to improving the work efficiency and work experience of relevant personnel.
[0128] In some optional implementations of this embodiment, the present application implements the functions of problem feedback and governance model selection, including:
[0129] Standardize feedback pages: Design business feedback and technology feedback pages, and standardize input formats, such as required fields, data types, and format requirements, to improve the accuracy and rationality of feedback data.
[0130] Simplified governance model: Based on past governance experience, the system has simplified the governance model into two modes: self-check and designated problem troubleshooting. The self-check mode is suitable for problems discovered and solved by business personnel, such as customer information errors. The designated problem troubleshooting mode is suitable for more complex problems or those requiring technical assistance, such as data errors caused by system failures.
[0131] Development of feedback templates: Develop fixed feedback templates for different governance models and roles, clarify the content and format of feedback, and facilitate feedback and communication between business and technical personnel.
[0132] In addition, this application also implements the function of modular versioning operation, including:
[0133] Table Versioning Reminder: This reminds business personnel to confirm data and entry criteria by finalizing the table. Before finalizing the table, the system checks the data in the table to ensure its completeness and accuracy.
[0134] Flexible versioning: Compared with the overall versioning in the incremental mode, modular versioning is more flexible. You can version specific parts of a table based on business needs, such as data for a certain time period or data for a certain business module.
[0135] Version record and review: record the operation information of each version, including the version time, version setter, version content, and review to ensure the compliance and effectiveness of the version operation.
[0136] In addition, this application also implements the functions of task management and process tracing, including:
[0137] Task association and push: Associating task execution with personnel roles, automatically pushes tasks to the corresponding business personnel or technical personnel based on the nature of the task and the personnel's responsibilities. For example, data collection tasks can be pushed to data collection personnel, and problem feedback tasks can be pushed to business personnel.
[0138] To-do reminder: Before the task ends, the system will send email reminders to the unprocessed to-do items according to the time set by the initiator to ensure that the task is completed on time.
[0139] Operation log recording: The entire operation log is recorded, including the operation information of each link such as task creation, assignment, execution, feedback, etc., to achieve process traceability.
[0140] Accountability: Based on the operation log, problems encountered during task execution are investigated and the responsible persons for each link are clearly identified.
[0141] In some optional implementations, the user information obtained is obtained with the user's consent and complies with relevant laws and policies.
[0142] In addition, any software tools or components not provided by our company that appear in the embodiments of this application are merely examples and do not represent actual use.
[0143] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0144] It should be emphasized that in order to further ensure the privacy and security of the above-mentioned problem data, the above-mentioned problem data can also be stored in a node of a blockchain.
[0145] The blockchain referred to in this application refers to a new application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (to prevent counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.
[0146] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0147] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0148] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0149] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0150] Further references Figure 3 , as a response to the above Figure 2 In order to realize the method shown in the figure, the present application provides an embodiment of a device for processing problem data. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0151] like Figure 3As shown, the problem data processing device 300 of this embodiment includes: an extraction module 301, an integration module 302, a positioning module 303, a verification module 304, a construction module 305 and a processing module 306.
[0152] The extraction module 301 is used to extract business data corresponding to business needs from multiple preset data sources;
[0153] Integration module 302, used to perform data integration processing on the business data to obtain corresponding target business data;
[0154] A positioning module 303 is configured to locate anomalies in the target business data based on a preset clustering algorithm and classification algorithm to obtain corresponding anomaly data;
[0155] A verification module 304 is configured to perform rule verification on the abnormal data based on a preset rule engine, and filter out problematic data that does not comply with the rules from the abnormal data;
[0156] A construction module 305 is configured to obtain a rule number corresponding to the problem data and construct a rectification script corresponding to the problem data based on the rule number;
[0157] The processing module 306 is configured to execute a rectification process corresponding to the problem data based on the rectification script.
[0158] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the method for processing problem data in the aforementioned embodiment, and are not described in detail here.
[0159] In some optional implementations of this embodiment, the extraction module 301 includes:
[0160] A first acquisition submodule is used to obtain update frequency information of each of the data sources;
[0161] A division submodule, configured to divide all the data sources into a first data source with frequent updates and a second data source with infrequent updates based on the update frequency information;
[0162] a first extraction submodule, configured to extract first business data corresponding to the business requirement from the first data source based on a preset incremental extraction strategy;
[0163] a second extraction submodule, configured to extract second business data corresponding to the business requirement from the second data source based on a preset full-data extraction strategy;
[0164] a combining submodule, configured to combine the first service data and the second service data to obtain corresponding combined data;
[0165] The first determining submodule is configured to use the combined data as the service data.
[0166] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the method for processing problem data in the aforementioned embodiment, and are not described in detail here.
[0167] In some optional implementations of this embodiment, the integration module 302 includes:
[0168] A preprocessing submodule, configured to preprocess the service data to obtain corresponding first processed data;
[0169] The first calling submodule is used to call a preset intelligent mapping algorithm;
[0170] A mapping submodule, configured to perform mapping conversion processing on the first processed data based on the intelligent mapping algorithm to obtain corresponding second processed data;
[0171] a verification submodule, configured to perform data verification on the second processed data;
[0172] The second determining submodule is configured to use the second processed data as the target business data if the second processed data passes the data verification.
[0173] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the method for processing problem data in the aforementioned embodiment, and are not described in detail here.
[0174] In some optional implementations of this embodiment, the positioning module 303 includes:
[0175] An extraction submodule, configured to extract features from the target business data to obtain corresponding feature data;
[0176] A clustering submodule, configured to perform clustering processing on the feature data based on the clustering algorithm to generate corresponding data clusters;
[0177] A classification submodule, configured to classify the feature data based on the classification algorithm to generate corresponding data categories;
[0178] The second acquisition submodule is used to obtain a preset feature model of normal data;
[0179] a comparison submodule, configured to compare the data cluster and the data category with the feature model respectively, to determine the specified data that deviates from a preset normal range;
[0180] The third determining submodule is configured to use the designated data as the abnormal data.
[0181] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the method for processing problem data in the aforementioned embodiment, and are not described in detail here.
[0182] In some optional implementations of this embodiment, the verification module 304 includes:
[0183] The second calling submodule is used to call the preset rule base;
[0184] A second acquisition submodule is configured to acquire verification rules from the rule base; wherein the verification rules include data quality rules and compliance rules;
[0185] A matching submodule, configured to match the abnormal data with the verification rules based on the rule engine to detect whether the abnormal data meets the rule requirements of the verification rules and obtain a corresponding detection result;
[0186] A screening submodule, configured to analyze the detection results to screen out designated abnormal data that does not conform to the verification rule from the abnormal data;
[0187] The fourth determining submodule is configured to use the designated abnormal data as the problem data.
[0188] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the method for processing problem data in the aforementioned embodiment, and are not described in detail here.
[0189] In some optional implementations of this embodiment, the construction module 305 includes:
[0190] The third calling submodule is used to call the preset script template library;
[0191] A query submodule, configured to query the script template library for a specified script template that matches the rule number;
[0192] The third acquisition submodule is used to obtain detailed information of the problem data;
[0193] A filling submodule, configured to fill the detailed information into a corresponding position in the specified script template to obtain a filled specified script template;
[0194] The fifth determining submodule is configured to use the filled designated script template as the rectification script corresponding to the problem data.
[0195] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the method for processing problem data in the aforementioned embodiment, and are not described in detail here.
[0196] In some optional implementations of this embodiment, the apparatus for processing question data further includes:
[0197] An acquisition module, configured to acquire a processing feedback result corresponding to the problem data;
[0198] A generating module, configured to generate a corresponding processing evaluation report based on the processing feedback result;
[0199] Calling module, used to call the preset visual interface;
[0200] A display module is used to display the processing evaluation report based on the visual interface.
[0201] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the method for processing problem data in the aforementioned embodiment, and are not described in detail here.
[0202] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0203] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 with components 41-43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0204] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.
[0205] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk equipped on the computer device 4, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, etc. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for problem data processing methods. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.
[0206] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or process data, such as computer-readable instructions for executing the method for processing the problem data.
[0207] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.
[0208] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0209] In an embodiment of the present application, business data corresponding to business needs are extracted from a plurality of preset data sources, and data integration processing is performed on the business data to obtain target business data, and then the target business data is anomaly located based on the combined use of clustering algorithm and classification algorithm to obtain corresponding abnormal data, and then the abnormal data is rule-verified based on the use of a rule engine, and problem data that does not conform to the rules is screened out from the abnormal data, and then a rectification script corresponding to the problem data is constructed based on the rule number corresponding to the obtained problem data, and finally the rectification processing corresponding to the problem data is executed based on the rectification script. Through the above-mentioned data processing process, the dynamic data governance method based on the fusion of intelligent algorithms and multi-source data provided by the present application can quickly and intelligently realize the accurate identification and automatic repair of problem data, improve the recognition efficiency and repair accuracy of existing problem data, and thus help improve data quality.
[0210] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the problem data processing method as described above.
[0211] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0212] In an embodiment of the present application, business data corresponding to business needs are extracted from a plurality of preset data sources, and data integration processing is performed on the business data to obtain target business data, and then the target business data is anomaly located based on the combined use of clustering algorithm and classification algorithm to obtain corresponding abnormal data, and then the abnormal data is rule-verified based on the use of a rule engine, and problem data that does not conform to the rules is screened out from the abnormal data, and then a rectification script corresponding to the problem data is constructed based on the rule number corresponding to the obtained problem data, and finally the rectification processing corresponding to the problem data is executed based on the rectification script. Through the above-mentioned data processing process, the dynamic data governance method based on the fusion of intelligent algorithms and multi-source data provided by the present application can quickly and intelligently realize the accurate identification and automatic repair of problem data, improve the recognition efficiency and repair accuracy of existing problem data, and thus help improve data quality.
[0213] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0214] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. A method for processing problem data, characterized in that: The steps include: Extract business data corresponding to business needs from multiple preset data sources; Performing data integration processing on the business data to obtain corresponding target business data; Based on the preset clustering algorithm and classification algorithm, the target business data is anomaly located to obtain corresponding abnormal data; Perform rule verification on the abnormal data based on a preset rule engine, and filter out problematic data that does not comply with the rules from the abnormal data; Obtaining a rule number corresponding to the problem data, and constructing a rectification script corresponding to the problem data based on the rule number; A rectification process corresponding to the problem data is executed based on the rectification script.
2. The method for processing problem data according to claim 1, characterized in that: The step of extracting business data corresponding to business requirements from a plurality of preset data sources specifically includes: Obtaining update frequency information of each of the data sources; Based on the update frequency information, all the data sources are divided into a first data source that is frequently updated and a second data source that is not frequently updated; Extracting first business data corresponding to the business requirement from the first data source based on a preset incremental extraction strategy; Extracting second business data corresponding to the business requirement from the second data source based on a preset full-data extraction strategy; Combining the first service data and the second service data to obtain corresponding combined data; The combined data is used as the service data.
3. The method for processing problem data according to claim 1, characterized in that: The step of performing data integration processing on the business data to obtain corresponding target business data specifically includes: Preprocessing the business data to obtain corresponding first processed data; Call the preset intelligent mapping algorithm; Performing mapping conversion processing on the first processed data based on the intelligent mapping algorithm to obtain corresponding second processed data; performing data verification on the second processed data; If the second processed data passes the data verification, the second processed data is used as the target business data.
4. The method for processing problem data according to claim 1, characterized in that: The step of locating anomalies in the target business data based on a preset clustering algorithm and classification algorithm to obtain corresponding abnormal data specifically includes: Extracting features from the target service data to obtain corresponding feature data; Performing clustering processing on the feature data based on the clustering algorithm to generate corresponding data clusters; Classify the feature data based on the classification algorithm to generate corresponding data categories; Obtaining a preset feature model of normal data; Comparing the data clusters and the data categories with the feature model respectively to determine the designated data that deviates from a preset normal range; The designated data is regarded as the abnormal data.
5. The method for processing problem data according to claim 1, characterized in that: The step of performing rule verification on the abnormal data based on a preset rule engine and filtering out problematic data that does not conform to the rules from the abnormal data specifically includes: Call the preset rule library; Obtaining verification rules from the rule base; wherein the verification rules include data quality rules and compliance rules; Matching the abnormal data with the verification rules based on the rule engine to detect whether the abnormal data meets the rule requirements of the verification rules and obtain corresponding detection results; Analyzing the detection results to filter out designated abnormal data that does not comply with the verification rules from the abnormal data; The designated abnormal data is used as the problem data.
6. The method for processing problem data according to claim 1, characterized in that: The step of constructing a rectification script corresponding to the problem data based on the rule number specifically includes: Call the preset script template library; Querying the specified script template that matches the rule number from the script template library; Obtain detailed information about the problem data; Fill the detailed information into the corresponding position in the specified script template to obtain the filled specified script template; The filled designated script template is used as the rectification script corresponding to the problem data.
7. The method for processing problem data according to claim 1, characterized in that: After the step of performing the rectification process corresponding to the problem data based on the rectification script, the method further includes: Obtaining a processing feedback result corresponding to the problem data; generating a corresponding processing evaluation report based on the processing feedback result; Call the preset visual interface; The processing evaluation report is displayed based on the visual interface.
8. A device for processing problem data, characterized in that: include: The extraction module is used to extract business data corresponding to business needs from multiple preset data sources; An integration module, configured to perform data integration processing on the business data to obtain corresponding target business data; A positioning module is used to locate abnormalities in the target business data based on a preset clustering algorithm and classification algorithm to obtain corresponding abnormal data; A verification module is used to perform rule verification on the abnormal data based on a preset rule engine, and filter out problematic data that does not comply with the rules from the abnormal data; A construction module, configured to obtain a rule number corresponding to the problem data, and construct a rectification script corresponding to the problem data based on the rule number; A processing module is used to perform rectification processing corresponding to the problem data based on the rectification script.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the method for processing problem data according to any one of claims 1 to 7 when executing the computer-readable instructions.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the method for processing problem data according to any one of claims 1 to 7.