A data fusion management method, a data management platform, a medium and a program product
By adopting the latest data standards and standard word root mapping library to screen financial data, and combining multi-dimensional scoring and refined permission management, the challenges of financial data management in banking data governance have been solved, the standardization and reliability of data have been improved, and the accuracy and security of data interaction have been ensured.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2026-03-27
AI Technical Summary
The lack of unified data governance and control products in the banking industry has increased the difficulty of data management for financial institutions and reduced the accuracy and timeliness of interactive information.
The latest data standards are used to screen Rongxin data. A standard word root mapping library is built through field naming and type validation. Data quality management requirements are identified and rectification processes are implemented. Multiple preset progress points are set for sampling and testing, and data access vouchers are generated for refined permission management.
It improves the standardization and reliability of data, reduces the risk of missing data quality defects, achieves data availability and security, and enhances the flexibility and efficiency of data services.
Smart Images

Figure CN119919221B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of data management, and particularly relates to a data management method, a data management platform, a medium and a program product. BACKGROUND
[0002] With the promotion of data as an important asset value of the banking industry, the degree of attention of companies to data is increasing. In order to effectively improve the management ability of data, and mine internal and external information resources to provide support for bank business decision-making, major banks have carried out data governance work.
[0003] At present, banks lacking unified data governance control products manage their own metadata information in the internal management of each data platform, and generally interact through company information technology center in the form of Excel documents in response to metadata demands of each data platform and IT user. In terms of data quality and data standardization, ETL check jobs and storage processes developed by custom are mainly added to the ODS system to complete the verification and check of data dictionary. In this way, multiple applications and multiple communications are required, which increases the difficulty of data management and reduces the accuracy and timeliness of interactive information. SUMMARY
[0004] The application provides a data management method, a data management platform, a medium and a program product, which are used to reduce the difficulty of data management and improve the accuracy and timeliness of interactive information.
[0005] In a first aspect, the application provides a data management method, which extracts qualified data from data based on the latest data standard, and the latest data standard is obtained by adding field naming verification and field type verification to the original data standard through standard roots;
[0006] Obtain data quality management requirements, and determine a data quality range according to the data quality management requirements;
[0007] In a case where the qualified data does not meet the data quality management requirements based on the data quality range, a rectification process is performed on the qualified data;
[0008] Data sampling detection is performed at preset progress points of the rectification process, and a plurality of sampling detection results are obtained, the detection standard of the data sampling detection is the data quality range, the sampling detection result is qualified or unqualified, and the number of preset progress points is a preset first number;
[0009] If the sampling detection result corresponding to any preset progress point is unqualified, the step of performing the rectification process on the qualified data is re-executed;
[0010] If the sampling detection results corresponding to the preset first number of preset progress points are all qualified, the rectified qualified data is outputted after the rectification process.
[0011] The rectified qualified data is stored in the data management library of the corresponding user.
[0012] By using the above technical solutions, the latest data standard is used to filter the data, improving the standardization of the data in field naming and type. After determining the data quality management requirements and scope, the data that does not meet the requirements is rectified, and multiple preset progress points are set for sampling detection during the rectification process, improving the controllability and effectiveness of the rectification process. Through multiple sampling verification of the rectification results, only when the detection results of all preset progress points are qualified, the rectified data is outputted, improving the overall quality level of the data. When any detection point is found to be unqualified, the rectification process is triggered to be re-executed, forming a closed-loop quality management system, reducing the risk of missing data quality defects. Finally, the rectified qualified data is stored in the user data management library, making the data more usable and reliable.
[0013] In some embodiments of the first aspect, based on the latest data standard, the qualified data in the data is extracted, specifically including:
[0014] A standard root mapping library is constructed, which contains standard roots of commonly used fields in business fields and corresponding field type sets;
[0015] Based on the standard root mapping library, the fields in the original data standard are processed for word segmentation, obtaining a field root sequence;
[0016] The field root sequence is matched with the standard root;
[0017] If the matching is successful, it is verified whether the verification field type of the field root sequence belongs to the field type set;
[0018] Fields that fail to match or whose field type verification fails are marked, and field correction suggestions are generated;
[0019] According to the field correction suggestion, the original data standard is updated to obtain the latest data standard.
[0020] By adopting the technical scheme, the standard word root mapping library is established, the standard word roots of the fields commonly used in the business field are associated with the field type set, and a basis support is provided for field standardization. The fields in the original data standard are subjected to word segmentation processing and matched with the standard word roots, and the standardization verification of the field naming is realized. Whether the field type belongs to the predetermined type set is verified, and the standardization of the field type is improved. For the fields that fail to match or do not pass the type verification, correction suggestions are generated and the data standard is updated according to the suggestions, so that the data standard is continuously optimized and improved. The standardization method based on the word root mapping reduces the subjectivity of human judgment, improves the accuracy and efficiency of the field standardization, and makes the data standard more scientific and standardized.
[0021] In combination with some embodiments of the first aspect, in some embodiments, data sampling detection is performed at a preset progress point of the rectification process, and a plurality of sampling detection results are obtained, specifically including:
[0022] A multi-dimensional scoring matrix is constructed based on the data quality range, and the multi-dimensional scoring matrix includes integrity, accuracy, consistency and timeliness indexes;
[0023] A stratified random sampling method is used to extract samples from qualified data;
[0024] The samples are scored in multiple dimensions based on the multi-dimensional scoring matrix, and dimension score results are obtained;
[0025] The dimension score results are weighted calculated based on preset dimension weights, and a comprehensive score is obtained;
[0026] The comprehensive score is compared with a preset qualified threshold, and a sampling detection result is obtained.
[0027] By adopting the technical scheme, the representativeness of the samples is improved by using the stratified random sampling method, and it is ensured that the evaluation result can accurately reflect the overall data quality condition. The comprehensive score obtained by scoring the samples in multiple dimensions and combining the preset dimension weights for weighted calculation more objectively reflects the overall level of data quality. The way of comparing the comprehensive score with the preset qualified threshold to obtain the detection result establishes a clear quality judgment standard, so that the quality evaluation result is measurable and comparable, and the scientificity and operability of data quality management are improved.
[0028] In combination with some embodiments of the first aspect, in some embodiments, after storing the rectified qualified data into the data management library corresponding to the user, the method further includes:
[0029] Receiving a data calling request sent by the user, the data calling request containing target data feature description information and use scenario description information;
[0030] The rectified Fusion data is labeled to obtain data label information, and the data label information includes a feature identifier matched with the target data feature description information.
[0031] A data use permission matrix is generated according to the use scenario description information, and the data use permission matrix includes data access level information.
[0032] A data call credential is generated based on the data access level information, and the data call credential includes an access time parameter, a use range parameter, and a desensitization requirement parameter.
[0033] The data access interface parameters are configured according to the data call credential, and the data access interface parameters are used to control the access mode of the rectified Fusion data.
[0034] By using the above technical solutions, the rectified Fusion data is labeled, which facilitates quick positioning of appropriate data according to the data feature requirements of users. Based on the use scenario description information, a data use permission matrix is generated, and a data call credential including an access time, a use range, and a desensitization requirement is generated, thereby establishing a fine-grained data access control mechanism. By configuring the data access interface parameters to control the data access mode, the manageability and security of data use are realized. This scene-based fine-grained permission management method improves the flexibility of data services while ensuring data security, so that data resources can be fully utilized within a controlled range.
[0035] In combination with some embodiments of the first aspect, in some embodiments, the rectified Fusion data is labeled to obtain data label information, specifically including:
[0036] The target data feature description information is parsed to obtain a feature keyword sequence.
[0037] A preset data label system is constructed, and the preset data label system includes a business attribute label subset, a time attribute label subset, and a scene attribute label subset.
[0038] The rectified Fusion data is analyzed based on the feature keyword sequence to generate an initial data label set.
[0039] The initial data label set is filtered using a preset label optimization rule to obtain a target data label set.
[0040] The correlation strength coefficients between labels in the target data label set are calculated to generate a label relationship network structure, and the data label information is obtained.
[0041] By adopting the technical solution, the feature keyword sequence is obtained by analyzing the target data feature description information, the initial data label set is generated by performing feature analysis on the rectified qualified fusion data based on the preset data label system, the target data label set is obtained by screening using the preset label optimization rule, and finally the label relationship network structure is generated by calculating the correlation strength coefficient between labels. This label processing method can accurately label and correlation analyze data from multiple dimensions. The label system covers three levels of business attributes, time attributes and scene attributes, making the data feature description more comprehensive and systematic. The screening of the label optimization rule can remove redundant and invalid labels, improving the accuracy and applicability of the labels. The establishment of the label relationship network structure reveals the correlation and strength between different labels, which helps better understand and utilize the internal relationship between data.
[0042] In combination with some embodiments of the first aspect, in some embodiments, after configuring the data access interface parameter according to the data calling credential, the method further comprises:
[0043] Collecting data calling behavior feature information, the data calling behavior feature information including calling timestamp, calling frequency statistical value and data usage mode identifier;
[0044] Setting a calling early warning threshold parameter according to a preset calling early warning rule;
[0045] Calculating a data reuse rate index based on the data calling behavior feature information;
[0046] Constructing a data calling scoring rule set, and quantitatively evaluating data usage compliance according to the data calling scoring rule set to obtain a compliance evaluation result;
[0047] Updating the data usage permission matrix according to the compliance evaluation result.
[0048] By adopting the technical solution, the data usage is dynamically monitored and the permission is adjusted by collecting the data calling behavior feature information and setting the calling early warning threshold parameter, combining the data reuse rate index calculation and the quantitative evaluation of data usage compliance. The system can record the calling timestamp, calling frequency and other behavior characteristics to timely discover abnormal data access patterns. The setting of the early warning threshold parameter provides a quantitative basis for early identification of potential risks. The calculation of the data reuse rate index reflects the efficiency of data resource use, and the data calling scoring rule set evaluates the compliance of data usage from multiple dimensions. Based on the compliance evaluation result, the data usage permission matrix is dynamically updated, which not only ensures the security of data access, but also provides a flexible permission adjustment mechanism, enhances the data security management capability, reduces the risk of data abuse, and optimizes the allocation efficiency of data resources.
[0049] In some embodiments of the first aspect, in some embodiments, the data call behavior feature information is collected, and specifically includes:
[0050] The data call log record is obtained, and the caller identity information and operation behavior sequence are extracted from the data call log record;
[0051] The operation behavior sequence is subjected to time sequence analysis to obtain data call time parameters and call frequency parameters, and a data call link graph is constructed based on the operation behavior sequence;
[0052] The data usage mode feature is extracted from the data call link graph to obtain data usage mode identification;
[0053] The data call time parameters, call frequency parameters and data usage mode identification are combined to obtain the data call behavior feature information.
[0054] By adopting the above technical solution, the construction of the data call link graph visualizes the data flow process, facilitating the tracking of the data usage path. By analyzing the time feature and frequency feature of the operation behavior sequence, the regularity and abnormal mode of data usage can be identified. This all-round data call behavior analysis method improves the traceability and controllability of data usage behavior, and provides strong technical support for data security management.
[0055] In the second aspect, the embodiments of the present application provide a data management platform, which includes one or more processors and a memory; the memory is coupled with the one or more processors, and is used to store computer program codes, the computer program codes including computer instructions, and the one or more processors invoke the computer instructions to enable the system to execute the method described in the first aspect and any possible implementation manner of the first aspect.
[0056] In the third aspect, the embodiments of the present application provide a computer readable storage medium, which includes instructions, and when the instructions run on the data management platform, enable the data management platform to execute the method described in the first aspect and any possible implementation manner of the first aspect.
[0057] In the fourth aspect, the embodiments of the present application provide a computer program product, which is characterized in that when the computer program product runs on the data management platform, enables the data management platform to execute the method described in any possible implementation manner of the first aspect.
[0058] The one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0059] 1. The application provides a kind of data management method of fusing, by using the latest data standard to filter fusing data by using the above technical solution, improve the standardization of data in field naming and type.After determining the data quality management demand and range, the data that does not meet the requirements is reformed, and multiple preset progress points are set during the rectification process to carry out sampling detection, improve the controllability and effectiveness of the rectification process. Through multiple sampling verification of rectification result, only when the detection results of all preset progress points are qualified, the data after rectification is output, improve the overall quality level of fusing data. When any detection point is found to be unqualified, rectification process is triggered to re-execute, forming a closed-loop quality management system, reducing the risk of missing data quality defects. Finally, the data that is rectified qualified is stored in user data management library, so that data has higher usability and reliability.
[0060] 2. The application provides a kind of data management method of fusing, by carrying out label processing to the fusing data that is rectified qualified, it is convenient to quickly locate suitable data according to the data characteristic demand of user. Based on the use scene description information generation data use permission matrix, and generate data calling voucher containing access time limit, use range and desensitization requirement etc. parameter, establish fine data access control mechanism. Through the configuration data access interface parameter to control data access mode, realize the manageability and security of data use. This kind of fine right management mode based on scene, while guaranteeing data security, improve the flexibility of data service, so that data resources can be fully utilized in controlled range.
[0061] 3. The application provides a kind of data management method of fusing, by collecting data calling behavior characteristic information and setting calling early warning threshold parameter, combined with data reuse rate index calculation and data use compliance quantitative evaluation, realize the dynamic monitoring and right adjustment of data use. System can discover abnormal data access mode in time by recording calling time stamp, calling frequency and other behavior characteristics. The setting of early warning threshold parameter provides quantitative basis for early identification of potential risks. The calculation of data reuse rate index reflects the use efficiency of data resources, and data calling score rule set evaluates the compliance of data use from multiple dimensions. Based on the compliance evaluation result, dynamically update data use permission matrix, which not only guarantees the security of data access, but also provides flexible right adjustment mechanism, enhances data security control ability, reduces data abuse risk, and also optimizes the allocation efficiency of data resources. BRIEF DESCRIPTION OF DRAWINGS
[0062] Figure 1 It is the first page schematic view of a kind of data management platform provided by the application.
[0063] Figure 2is a functional module structure schematic diagram of a data management platform provided by the application.
[0064] Figure 3 is a flow schematic diagram of a data management method provided by the application.
[0065] Figure 4 is another flow schematic diagram of a data management method provided by the application.
[0066] Figure 5 is an entity device structure schematic diagram of a data management system provided by the application. DETAILED DESCRIPTION
[0067] The terms used in the following embodiments of the application are only for the purpose of describing the specific embodiments and are not intended to be limiting on the application. As used in the specification and the appended claims of the application, the singular forms "a," "an," and "the" are intended to include both singular and plural forms, unless the context clearly indicates otherwise. It will be further understood that the terms "and / or" used in the application means any or all possible combinations of one or more of the listed items.
[0068] Hereinafter, the terms "first" and "second" are only for the purpose of description, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features, and in the description of the embodiments of the application, the meaning of "a plurality of" is two or more, unless otherwise specified.
[0069] It should be noted that the data management method provided by the application is applied to a data management platform. The data management platform is an effective technical tool for data standard management, metadata management, and data quality management work promoted by financial institutions for enterprise information collection and management work, and is also a unified work platform for financial institution business personnel and IT staff to exchange, feedback, track, and solve information collection and management problems. The home page interface of the data management platform is as shown in Figure 1 .
[0070] As shown in Figure 2 , the data management platform includes an application layer, a data service layer, a data platform layer, a data layer, and an infrastructure layer.
[0071] The application layer includes data asset view, customer view, risk compliance, and business analysis modules.
[0072] The data service layer comprises several data function components, and the data function components are specifically report, autonomous analysis, online query, batch data, quality report, analysis mining, data cache, and service log.
[0073] The data platform layer comprises a data visualization module, a data collection module, a data integration module, a data calculation module, an intelligent scheduling module, and a workflow engine module. The data visualization module comprises data report, data analysis, visual large screen, and data dashboard configurable graphics. The data collection module comprises streaming collection, real-time collection, batch collection, and crawler. The data integration module comprises data cleaning, data checking, and data conversion. The data calculation module comprises offline calculation, streaming calculation, and real-time calculation. The intelligent scheduling module comprises resource scheduling, task scheduling, and data scheduling.
[0074] The data layer comprises a metadata area, a standard market, an archive area, an existing warehouse / market, and a source data area. The metadata area comprises a meta model, metadata, and a meta process. The standard market comprises account information, product information, asset information, and customer information. The existing warehouse / market comprises a risk market, a marketing market, and an evaluation market.
[0075] The infrastructure layer comprises Hadoop, a graph database, an MPP database, a relational database, a full-text search engine database, computing resources, network resources, and storage resources.
[0076] In addition, the data management platform further comprises a data management module and a data security module. The data management module comprises metadata management, data directory management, data standard management, data source management, blood relationship analysis, index system management, data distribution, and data quality. The data security module comprises data permission, data security, data download, data encryption, and data desensitization.
[0077] The following describes a data management method of the application by means of an embodiment and in combination with Figure 3 The data management method of the application is described as follows:
[0078] Please refer to Figure 3 The data management method of the application is described as follows:
[0079] S101, extracting qualified data from the data of the data management platform based on the latest data standard.
[0080] The data management platform extracts qualified data from the data of the data management platform based on the latest data standard. The latest data standard is obtained by adding field naming verification and field type verification to the original data standard through a standard root. Specifically, a standard root mapping library is constructed, and the standard root mapping library comprises standard roots of fields commonly used in business fields and corresponding field type sets.
[0081] The field in the original data standard is processed by word segmentation based on a standard root mapping library to obtain a field root sequence;
[0082] The field root sequence is matched with the standard root;
[0083] If the matching is successful, it is verified whether the verification field type of the field root sequence belongs to the field type set;
[0084] The fields that fail to match or whose field type verification fails are marked to generate field correction suggestions;
[0085] The original data standard is updated according to the field correction suggestions to obtain the latest data standard.
[0086] In this step, the data management platform first needs to obtain the latest data standard. The data standard refers to the specification and requirement for the format, content, quality, etc. of data, which ensures the consistency, accuracy and integrity of data. The latest data standard is obtained by adding field naming verification and field type verification based on the standard root on the basis of the original data standard. This process can optimize the data standard and improve the standardization and accuracy of the data standard. After obtaining the latest data standard, the data management platform will extract the data of Rongxin based on the standard to screen out qualified data of Rongxin that meets the standard requirements.
[0087] In order to efficiently obtain the latest data standard, the data management platform can construct a standard root mapping library. The mapping library contains the standard roots of commonly used fields in business fields and the corresponding field type set. By processing the fields in the original data standard by word segmentation, a field root sequence is obtained, and then the field root sequence is matched with the standard root. If the matching is successful, it is further verified whether the verification field type of the field root sequence belongs to the field type set. For the fields that fail to match or whose field type verification fails, the data management platform can be marked and the corresponding field correction suggestions are generated. Finally, the original data standard is updated according to the field correction suggestions, so as to obtain the latest data standard.
[0088] S102, acquire data quality management requirements, and determine data quality range according to data quality management requirements;
[0089] In this step, the data management platform needs to obtain the data quality management requirements proposed by users or business parties. Data quality management requirements reflect the expectations and requirements of users or business parties for data quality, such as accuracy, completeness, consistency, timeliness, etc. The data management platform can understand and collect their data quality management requirements through communication and research with users or business parties. After clarifying the data quality management requirements, the data management platform will determine the scope of data quality according to these requirements, i.e. the data quality indicators and dimensions that need to be focused on and managed.
[0090] In order to better obtain and understand the data quality management requirements, the data management platform can provide a set of standardized requirement templates or questionnaires. By guiding users or business parties to fill out these templates or questionnaires, the data management platform can systematically and comprehensively collect their requirements for data quality. At the same time, the data management platform can also establish a data quality requirement library to record and manage historical data quality requirements, and perform classification, organization and analysis. Through the mining and reuse of these requirements, the data management platform can more efficiently and accurately identify and determine new data quality management requirements.
[0091] S103、In the case that the qualified fusion data determined based on the data quality range does not meet the data quality management requirements, the qualified fusion data is subjected to a rectification process;
[0092] In this step, the data management platform needs to determine whether the qualified fusion data determined based on the data quality range meets the data quality management requirements. If the quality level of the qualified fusion data does not meet the standards specified by the data quality management requirements, it indicates that the current fusion data still has certain defects and deficiencies in quality and needs to be rectified and improved. At this time, the data management platform will start the rectification process of the qualified fusion data, and through a series of data cleaning, conversion, correction, etc. operations, improve the quality of the qualified fusion data to meet the data quality management requirements.
[0093] In the process of rectifying the qualified fusion data, the data management platform can use various data quality improvement techniques and means. For example, for incomplete or missing data, data completion, data estimation, etc. techniques can be used for filling and correction; for inaccurate or erroneous data, data checking, data correction, etc. techniques can be used for identification and correction; for inconsistent data, data standardization, data coordination, etc. techniques can be used for unification and standardization. At the same time, the data management platform can also use the data quality rule engine to automatically detect and identify various quality problems in the qualified fusion data, and according to the pre-defined quality rules and strategies, automatically propose rectification suggestions and solutions.
[0094] During the process of rectifying qualified Rongxin data, some new quality issues may arise. For example, when completing missing data, new inaccurate data may be introduced; when unifying inconsistent data, changes or loss of data semantics may occur. To avoid and reduce these problems, the data management platform can establish a monitoring and verification mechanism for data rectification. This mechanism monitors the quality of newly added or changed data in real time during the rectification process, and promptly verifies and corrects any new quality issues discovered. Simultaneously, the data management platform can record and track various operations and changes during the data rectification process through technologies such as metadata management and data lineage analysis, ensuring the traceability and explainability of data rectification and providing support for locating and tracing data quality issues.
[0095] S104. Data sampling and testing are conducted at the preset progress points of the rectification process to obtain several sampling and testing results.
[0096] The data management platform conducts data sampling inspections at preset progress points in the rectification process, obtaining several sampling inspection results. The inspection standard for data sampling inspection is the data quality range, and the sampling inspection results are either qualified or unqualified. The number of preset progress points is a preset first number. Specifically: a multi-dimensional scoring matrix is constructed based on the data quality range. The multi-dimensional scoring matrix includes indicators of completeness, accuracy, consistency, and timeliness.
[0097] A stratified random sampling method was used to extract samples from qualified Rongxin data;
[0098] The sample is scored in multiple dimensions based on a multi-dimensional scoring matrix to obtain the scoring results for each dimension.
[0099] The scores for each dimension are weighted and calculated based on preset dimension weights to obtain a comprehensive score.
[0100] The comprehensive score is compared with the preset pass threshold to obtain the sampling test results.
[0101] In this step, the data management platform needs to conduct sampling inspections on the rectified and qualified Rongxin data at preset progress points in the rectification process. Preset progress points refer to key nodes or milestones predetermined in the rectification process. At these nodes, data quality is inspected and evaluated to understand the intermediate effects and progress of the rectification. The data management platform will randomly select a certain number of data samples at each preset progress point, conduct quality inspections on these samples according to the inspection standards and rules defined in the data quality range, and obtain the corresponding sampling inspection results to evaluate and judge the rectification effect at the current stage.
[0102] To ensure more comprehensive and reliable sampling results, the data management platform can employ stratified random sampling to select sample data. Stratified random sampling involves dividing the data population into different strata based on certain characteristics or attributes, and then randomly sampling within each stratum. This method ensures that the sample data has a certain degree of representativeness and balance across different strata, thereby improving the accuracy and reliability of the sampling results. Simultaneously, the data management platform can dynamically adjust the sampling ratio and frequency, flexibly setting sampling plans for different progress points based on the actual situation of rectification progress and effects, thus optimizing testing efficiency and cost.
[0103] S105. If any sampling inspection result corresponding to any preset progress point is unqualified, the steps of rectifying qualified Rongxin data shall be repeated.
[0104] In this step, the data management platform needs to comprehensively judge and evaluate the sampling inspection results at each preset progress point. If any of the preset progress points fails to meet the inspection results, it indicates that the currently rectified qualified Rongxin data still does not meet the requirements of data quality management. The rectification process cannot end there, and a new round of data rectification process needs to be restarted to address the problems and defects found during the inspection. The data management platform will retrospectively analyze the rectification process based on the unqualified inspection results, identify the causes of the quality problems, and formulate targeted rectification plans and measures to process and improve the qualified Rongxin data again until it meets the data quality management requirements.
[0105] To improve the efficiency and quality of the data rectification process, the data management platform can establish a comprehensive rectification problem analysis and decision-making mechanism. Through in-depth analysis and mining of non-conforming test results, the data management platform can identify common problems and key influencing factors in the rectification process, and design standardized rectification solutions and specifications for these problems and factors, forming a reusable rectification knowledge base and best practices. Simultaneously, the data management platform can utilize data quality monitoring and early warning mechanisms to monitor the changing trends of key quality indicators in real time during the rectification process. Once abnormal fluctuations or exceeding thresholds are detected, timely warnings are issued and countermeasures are taken to prevent the problem from escalating and spreading.
[0106] During the repeated execution of data rectification processes, situations may arise where rectification efforts stagnate or similar problems recur. This indicates that current rectification plans and methods may have limitations or blind spots, failing to effectively address certain fundamental or hidden quality issues. Therefore, the data management platform needs to optimize and improve the rectification process itself. On one hand, this can be achieved by introducing new data processing technologies and tools, such as artificial intelligence and machine learning, to enhance the intelligence and adaptability of the rectification process and improve its effectiveness. On the other hand, it can be achieved by strengthening communication and collaboration with business domain experts, fully absorbing and integrating their understanding and insights into business data, and optimizing key parameters and rules in the rectification process to better align with actual business needs. Simultaneously, the data management platform can regularly review and summarize the rectification process, learning from both successes and failures to continuously improve and innovate the rectification mechanism.
[0107] S106. If the sampling test results corresponding to the first number of preset progress points are all qualified, then output the qualified Rongxin data after the rectification process.
[0108] In this step, the data management platform needs to make a final judgment on the sampling inspection results of all preset progress points. If all the sampling inspection results corresponding to the first preset number of progress points are qualified, it indicates that the qualified Rongxin data after the rectification process has met the standards required for data quality management, and the entire data rectification process can be successfully completed and ended. At this time, the data management platform will officially output and release the rectified qualified Rongxin data as the basis for subsequent business applications and data analysis. Through this high-quality rectified data, the normal operation and stable functioning of the business can be effectively supported and promoted.
[0109] To facilitate the management and use of rectified and qualified Rongxin data, the data management platform can package, catalog, and publish the rectified data according to certain organizational methods and standards. For example, based on attributes such as business type, source system, and update time, rectified and qualified data can be divided into different datasets or data modules, and corresponding data directories and indexes can be created to facilitate users' quick retrieval and location of the required data. Furthermore, based on the data's security level and usage permissions, rectified and qualified data can be identified into different access categories, and corresponding access control and encryption measures can be implemented to ensure data security and confidentiality. Simultaneously, the data management platform can also provide flexible and diverse data publishing and subscription mechanisms, allowing users to obtain rectified and qualified data on demand, improving the convenience and real-time nature of data use.
[0110] During the process of outputting rectified and qualified data from Rongxin, issues such as data format or interface mismatches and poor data transmission performance may be encountered, affecting the effective use of data in business systems. To solve these problems, the data management platform needs to strengthen collaboration and communication with data users. On the one hand, this can be achieved by establishing unified data format and interface standards, clarifying the output specifications and requirements for rectified and qualified data, and ensuring data interoperability and consistency. On the other hand, it can be achieved by optimizing data transmission channels and protocols, increasing data transmission bandwidth and concurrency, reducing data arrival latency, and improving the real-time performance and stability of data supply. Simultaneously, the data management platform can establish a comprehensive data usage feedback and optimization mechanism, regularly collecting opinions and suggestions from data users, and continuously improving the content and form of data output to better meet the actual needs of the business.
[0111] S107. Store the rectified and qualified Rongxin data into the corresponding user's data management database.
[0112] In this step, the data management platform needs to persistently store and manage the rectified Rongxin data. By storing this high-quality data in the corresponding user's data management repository, a long-term, stable, and reliable data resource library can be provided, allowing users to easily access and use the business data they need at any time. A data management repository is a data warehouse or data platform specifically designed for storing, organizing, and managing enterprise data assets. It can systematically classify, catalog, retrieve, and access rectified data, making the data a truly usable, controllable, and revitalized valuable resource for the enterprise.
[0113] In building a user data management repository, the data management platform can employ various advanced data organization and storage technologies. For example, a data lake architecture can be used to store rectified qualified data in its original format within a distributed file system, enabling rapid retrieval and access to massive amounts of heterogeneous data through metadata management and data retrieval tools. Alternatively, a data warehouse architecture can be adopted, modeling and organizing rectified qualified data according to subject areas or business processes to form a highly integrated, analysis-oriented structured data system, facilitating multi-dimensional and multi-granular data analysis and mining for users. Simultaneously, the data management platform can leverage big data platforms and cloud computing platforms to provide distributed storage and elastic computing capabilities for massive amounts of data, ensuring the high scalability and availability of the data management repository.
[0114] In the above embodiments, by adopting the technical solution described above and using the latest data standards to screen Rongxin data, the standardization of data field naming and types is improved. After determining the data quality management requirements and scope, non-compliant data is rectified, and multiple preset progress points are set for sampling and testing during the rectification process, improving the controllability and effectiveness of the rectification process. By performing multiple sampling verifications on the rectification results, the rectified data is only output when the test results of all preset progress points are qualified, thus improving the overall quality level of Rongxin data. When any test point is found to be unqualified, the rectification process is triggered to be re-executed, forming a closed-loop quality management system and reducing the risk of missing data quality defects. Finally, the rectified data is stored in the user data management database, giving the data higher availability and reliability.
[0115] The first embodiment described above primarily outlines Rongxin Data's quality management and rectification process, focusing on the quality dimensions of data governance. Through standardized extraction, quality testing, and rectification verification, it ensures that Rongxin Data meets the expected quality standards. The following section combines... Figure 4 Another data management method for Rongxin in the embodiments of this application is described below:
[0116] Please see Figure 4 This is another flowchart illustrating a data management method for financial services in this application.
[0117] S201. Receive data access requests sent by users, and perform tagging processing on the rectified and qualified Rongxin data to obtain data tag information;
[0118] The data management platform receives data retrieval requests from users, tags the rectified and qualified Rongxin data, and obtains data tag information. The data retrieval request includes target data feature description information and usage scenario description information. The data tag information includes feature identifiers that match the target data feature description information. Specifically: The target data feature description information is parsed to obtain a sequence of feature keywords;
[0119] Construct a pre-defined data tagging system, which includes a subset of business attribute tags, a subset of time attribute tags, and a subset of scenario attribute tags;
[0120] Based on the feature keyword sequence, feature analysis is performed on the qualified Rongxin data after rectification to generate an initial data tag set;
[0121] The initial data label set is filtered using preset label optimization rules to obtain the target data label set;
[0122] Calculate the association strength coefficient between each label in the target data label set, generate the label relationship network structure, and obtain the data label information.
[0123] In this step, the data management platform first receives a data request from the user. This request includes a description of the target data's characteristics and a description of its usage scenario, reflecting the user's data needs. After clarifying the user's needs, the data management platform will tag the rectified Rongxin data. Tagging involves assigning various attributes and characteristics to the data, forming structured and standardized data tag information to facilitate subsequent data retrieval, matching, and application. The data tag information includes various feature identifiers that match the target data's characteristic description, reflecting the data's content, type, quality, timeliness, and other attribute characteristics.
[0124] To achieve efficient and accurate data tagging, the data management platform can employ various automated semantic analysis and feature extraction techniques. First, it parses the user-provided target data feature descriptions, extracting keywords to form a sequence of feature keywords. Then, it uses a pre-built data tagging system to match and map this sequence of feature keywords. The data tagging system is a multi-dimensional, multi-level tag classification framework, including subsets such as business attribute tags, time attribute tags, and scenario attribute tags, covering key features of data in different application scenarios. By comparing feature keywords with the tagging system, the data management platform can automatically generate an initial set of data tags. Next, using pre-defined tag optimization rules, such as tag importance, coverage, and mutual exclusivity, the initial tag set is filtered and optimized, ultimately resulting in a high-quality target data tag set. Furthermore, the data management platform can calculate the correlation strength between tags in the target tag set, generating a tag relationship network structure that comprehensively presents the inherent connections and combination logic between data tags.
[0125] During the data tagging process, issues such as incomplete tag coverage, tag redundancy, and inconsistent tag semantics may arise, affecting tag quality and subsequent data retrieval performance. To address these issues, data management platforms can implement several optimization and improvement measures. For example, by introducing external knowledge sources such as industry-specific dictionaries and knowledge graphs, the tag system can be expanded and improved, enhancing its coverage and professionalism. Furthermore, by setting statistical thresholds for tag usage frequency and co-occurrence probability, redundant or low-quality tags can be periodically eliminated or merged, maintaining the simplicity and effectiveness of the tag system. Additionally, manual review and correction mechanisms can be used to conduct quality checks and calibrations on automatically generated tags, continuously optimizing the consistency and accuracy of tag semantics.
[0126] S202. Generate a data usage permission matrix based on the usage scenario description information;
[0127] The data management platform generates a data access permission matrix based on the usage scenario description information. The data access permission matrix contains data access level information.
[0128] In this step, the data management platform will generate a data access permission matrix based on the usage scenario descriptions provided by the user. The data access permission matrix is a permission configuration scheme used to regulate and control user access to and use of data. It comprehensively considers multiple factors such as user roles, responsibilities, and business scenarios, subdividing and classifying user data usage permissions to form a hierarchical authorization management mechanism. The core content of the permission matrix is data access level information, clearly defining the specific permission settings for different levels of users regarding data visibility, operability, and shareability. Through the data access permission matrix, the data management platform can more granularly manage data resources, ensuring the security and compliance of data during its flow and use.
[0129] To dynamically generate a permission matrix based on scenario descriptions, the data management platform can design a rule-based scenario analysis and permission mapping mechanism. First, the user-provided scenario descriptions are semantically parsed and refined to identify key scenario elements, such as business processes, usage purposes, and involved roles. Then, these scenario elements are matched against predefined scenario models and automatically categorized into a standard scenario type. Different scenario types correspond to different permission definition templates, which pre-define common user roles and corresponding data permission configuration items for that scenario. Based on the scenario matching results, the data management platform can automatically populate and modify the role types and permission details in the templates until a complete, customized data permission matrix is generated. During this process, certain policy rules and constraints can be used to automatically check and correct potential errors and conflicts in the permission matrix, such as permission leaks and excessive permissions, to further enhance the rationality of the permission matrix.
[0130] S203. Generate data access credentials based on data access level information;
[0131] The data management platform generates data access credentials based on data access level information. The data access credentials include access validity parameters, scope of use parameters, and de-identification requirements parameters.
[0132] In this step, the data management platform will further generate data access credentials based on the data access level information generated in the previous step. A data access credential is an authorization and authentication token automatically generated by the system when a user requests data access permissions. It encrypts the user's identity information and access permission details, requiring the user to present the credential to verify their identity and control access permissions when subsequently accessing data. Data access credentials typically include multiple control parameters such as access validity period, scope of use, and anonymization requirements, providing comprehensive constraints and management of the user's data access behavior, further refining and strengthening data access security.
[0133] To dynamically generate access credentials based on data access levels, the data management platform can employ a series of encryption algorithms and security protocols. First, the data access level information is encoded and serialized to form standardized authorization text. Then, symmetric or asymmetric encryption algorithms, such as AES and RSA, are used to encrypt the authorization text, generating a digital signature or encryption token. The digital signature verifies the integrity and credibility of the authorization text, while the encryption token protects the confidentiality of the authorization information. Next, the encrypted authorization information is bound to the user's identity ID, timestamp, etc., to generate the final data access credential. Multiple access control parameter values in the credential can also be automatically converted and extracted from the data access level according to a pre-defined mapping relationship. Throughout this process, the data management platform can flexibly select and customize various security measures, such as key management, access log auditing, and multi-factor authentication, to enhance the anti-counterfeiting and security of the credential.
[0134] S204. Configure data access interface parameters according to the data retrieval credential;
[0135] The data management platform configures data access interface parameters based on data access credentials. These parameters are used to control the access method for rectified and qualified Rongxin data.
[0136] In this step, the data management platform will configure the relevant parameters of the data access interface based on the data retrieval credentials generated in the previous step. The data access interface is a set of standardized data retrieval APIs provided by the data management platform for users. Users submit data retrieval requests and obtain data return results through these interfaces. Interface parameters are a series of configuration options that control data access behavior, such as table name, field list, query conditions, and result format. By dynamically setting and managing these interface parameters, security restrictions on data access can be further strengthened, ensuring that data is exposed strictly according to the permission boundaries defined by the credentials, thus preventing unauthorized access and data leakage.
[0137] To enable automatic adaptation of interface parameters based on the calling credentials, the data management platform can maintain a set of parameter mapping and conversion rules. First, the data calling credentials are parsed to extract key parameter information, including data range, access validity period, and anonymization requirements. Then, this parameter information is matched against the backend data model and interface definition specifications to identify the corresponding interface control parameters. During the matching process, pre-configured adaptation rules come into play, achieving automatic mapping and conversion from credential parameters to interface parameters, ensuring consistency between credential requirements and interface behavior. Finally, the converted parameter values are filled back into the interface definition. When a data request is submitted, these parameters will automatically take effect, controlling the actual data access behavior. In this process, the data management platform can also promptly detect and block unauthorized parameter modification behavior through parameter validity verification and security check mechanisms, further strengthening the interface's protection capabilities.
[0138] S205. Collect data retrieval behavior characteristic information;
[0139] The data management platform collects data call behavior characteristic information, including call timestamps, call frequency statistics, and data usage method identifiers. Specifically, it obtains data call log records and extracts caller identity information and operation behavior sequences from these log records.
[0140] Perform time-series analysis on the operation behavior sequence to obtain data call time parameters and call frequency parameters, and construct a data call link diagram based on the operation behavior sequence;
[0141] Extract data usage characteristics from the data call chain diagram to obtain data usage identifiers;
[0142] The data call time parameter, call frequency parameter, and data usage method identifier are combined to obtain data call behavior characteristic information.
[0143] In this step, the data management platform will collect and record users' data access behavior throughout the entire process, extracting various data access behavior characteristics. These characteristics comprehensively reflect the user's data usage across various dimensions, including the time, frequency, and method of access. By collecting these behavioral characteristics, the data management platform can monitor data usage dynamics and patterns in real time, assess the rationality and legitimacy of data access, and provide foundational data support for subsequent behavior auditing, risk analysis, and access optimization measures.
[0144] During the data call behavior characteristics collection process, the data management platform can obtain raw data call information through multiple channels using various methods such as log analysis and traffic monitoring. First, the data management platform needs to monitor the entire data call process, deploying collection probes at different locations such as gateways, interfaces, and databases to capture data request and response messages in real time and record them as call logs. Then, the raw call logs are parsed and extracted to extract structured behavior records, including elements such as the caller's identity ID, timestamp, operation type, and data object. Next, multi-dimensional statistical analysis is performed on the behavior records. On the one hand, time-series aggregation of call records generates frequency statistics reflecting call intensity, including real-time concurrency and total call volume per hour / day / week / month. On the other hand, the call process is traced based on the call chain to determine the final data usage method, such as caching, downloading, or re-sharing. Finally, the statistical results are correlated with the raw behavior records to comprehensively compile a complete profile of data call behavior characteristics.
[0145] S206. Set the call warning threshold parameter according to the preset call warning rules;
[0146] In this step, the data management platform needs to set corresponding call alert threshold parameters based on the system's preset call alert rules. Call alert rules are a set of built-in criteria for judging abnormal data access behavior. These rules can promptly detect and identify various unauthorized call behaviors such as misoperations, malicious crawling, and unauthorized access. The alert threshold parameters are a series of quantified critical values based on the alert rules, used to define whether a call behavior triggers an alert, thus automating the alerting process. Through dynamic configuration of the alert threshold parameters, the data management platform can flexibly adjust the sensitivity and granularity of alerts, finding the optimal balance between prevention and convenience.
[0147] To transform early warning rules into early warning threshold parameters, the data management platform first needs to comprehensively review and structure the early warning rules, forming a triplet format, such as <calling subject, data object, early warning condition>. The early warning condition part needs to be as detailed and quantified as possible, clearly defining the criteria and numerical boundaries for triggering an early warning. For example, the original rule "prohibit high-frequency calls to sensitive data within a short period" can be quantified as "the frequency of calls to the same data table by the same user cannot exceed 100 times per minute." During the rule quantification process, the data call behavior characteristics collected in previous steps can be fully referenced to statistically analyze the distribution of various indicators of normal call behavior, such as mean, variance, and median. Based on this, an appropriate safety margin can be set to ultimately determine the early warning threshold parameters. The setting of threshold parameters should also fully consider factors such as data security level and business importance, setting differentiated early warning trigger thresholds for different data and different scenarios. After the thresholds are set, they can be used to judge in real time whether user call behavior is compliant.
[0148] S207. Calculate the data reuse rate index based on data call behavior characteristic information;
[0149] In this step, the data management platform will analyze and calculate the data reuse rate metric based on the collected data access behavior characteristics. Data reuse rate is an important data asset management metric used to assess the repeated use of the same data across different scenarios and users, reflecting the efficiency of data asset utilization and value creation capabilities. By calculating the data reuse rate, the data management platform can gain insights into the current state of data sharing applications, identify data usage hotspots, and pinpoint access bottlenecks, thereby optimizing data governance strategies and maximizing data value.
[0150] When calculating data reuse rate, data management platforms can employ various quantitative analysis models and statistical algorithms. The basic idea is to statistically analyze the usage of specific data over a certain period, measuring the reuse level by calculating indicators such as the number of users accessing the data, the number of accesses, and the number of usage scenarios. Specifically: First, it can calculate the concentration indicators of data access users, such as the percentage of users contributing the most accesses, visually representing the "80 / 20" distribution effect. Second, it can perform cluster analysis on data usage scenarios, measuring the frequency and proportion of cross-scenario calls, characterizing the degree of data sharing among vertical businesses. Third, it can statistically analyze the temporal distribution characteristics of data access, such as monthly active users and average daily access volume, examining the data's support for business continuity. Fourth, it can track and statistically manage the secondary processing and utilization of data based on data lineage, evaluating the effectiveness of data value-added applications. Finally, the various sub-indicators can be weighted and summed to form a comprehensive data reuse rate scoring system.
[0151] S208. Construct a data access scoring rule set, and quantitatively evaluate the compliance of data use based on the data access scoring rule set to obtain the compliance evaluation result;
[0152] In this step, the data management platform needs to construct a comprehensive set of data access scoring rules and conduct a systematic compliance assessment of data access behavior on the platform based on these rules, ultimately generating quantified compliance assessment results. The data access scoring rules are a set of clearly defined and standardized criteria for judging data usage behavior, covering all key aspects of data access and comprehensively characterizing the compliance of data access from multiple dimensions such as permission management, access control, and data usage. Through this set of objective and impartial scoring rules, the data management platform can efficiently identify and quantify various risks and problems in data access, promoting a shift in data management from passive prevention to proactive auditing, thereby better regulating user behavior and ensuring data security.
[0153] To establish a scientific and comprehensive set of scoring rules, the data management platform needs to systematically review existing data usage management systems, industry standards and specifications, and compliance audit requirements. It must extract and summarize the key evaluation points, and select representative and operable scoring dimensions and indicators, such as data access frequency, access time period, data usage scope, and usage purpose, to construct the overall framework of the rule set. Based on this, for each scoring indicator, a quantitative scoring standard is defined, and a unified scoring scale and weight are established. For example, for data access frequency, multiple thresholds can be set; access exceeding the threshold will deduct corresponding points, and differentiated deduction levels can be set according to the importance and sensitivity of different data. During actual evaluation, the system will automatically obtain the associated information for each data call, such as data content, user role, and usage stage, substitute it into the scoring rules to calculate the score, and finally summarize the results to obtain a comprehensive compliance score for the data call operation, used to measure its compliance risk level.
[0154] S209. Update the data usage permission matrix based on the compliance assessment results.
[0155] After the compliance assessment in the previous step is completed, the data management platform needs to dynamically update and adjust the data access permission matrix based on the assessment results. The data access permission matrix records the access permissions for different types of users to different data and serves as the foundation for controlling data access behavior. Dynamically optimizing the permission matrix based on compliance scores allows for targeted reduction or expansion of data access permissions for specific users, preventing data abuse or omissions at the source and effectively improving data security. This closed-loop matrix optimization mechanism makes the platform's permission management more intelligent, precise, and flexible, truly achieving dynamic authorization and timely control based on user behavior.
[0156] Specifically, after obtaining a compliance score for each data access, the data management platform correlates the score with the user's past scoring history to continuously track the compliance trend of the user's data usage behavior. On the one hand, if a user's compliance score continues to rise and is better than the average level of similar users, their permissions can be appropriately relaxed, such as granting access to more data or increasing the scope of data visibility, to incentivize compliant use. On the other hand, if a user's compliance score remains low or fluctuates abnormally, the platform can reduce their permissions accordingly based on preset punitive measures, such as limiting access frequency, narrowing the scope of accessible data, or even completely prohibiting access, to effectively curb violations. Simultaneously, when a user's data access behavior undergoes significant adjustments, such as a change in access objects or a sudden increase in usage frequency, even if the score has not yet reached the red line, the system can pre-adjust their permission matrix to avoid potential risks from over-authorization. Furthermore, the data management platform can also use visualization and other technologies to intuitively present the current scoring status and permission distribution of different users, allowing administrators to gain real-time insight into security dynamics and providing a reference for further optimizing permission strategies.
[0157] In the above embodiments, the rectified Rongxin data is tagged to facilitate the rapid location of suitable data based on users' data characteristics and needs. A data access permission matrix is generated based on usage scenario descriptions, and data retrieval credentials containing parameters such as access timeliness, usage scope, and anonymization requirements are generated, establishing a refined data access control mechanism. By configuring data access interface parameters to control data access methods, manageability and security of data use are achieved. This scenario-based refined permission management approach enhances the flexibility of data services while ensuring data security, enabling data resources to be fully utilized within a controlled scope.
[0158] The data management platform in the embodiments of this invention is described below from the perspective of hardware processing. Please refer to [link / reference]. Figure 5 This is a schematic diagram of the physical device structure of a Rongxin data management platform provided in an embodiment of this application.
[0159] It should be noted that, Figure 5 The structure of the data management platform shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0160] like Figure 5As shown, the data management platform includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 302 or programs loaded from storage section 308 into Random Access Memory (RAM) 303, such as executing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.
[0161] The following components are connected to I / O interface 305: input section 306 including a camera, infrared sensor, etc.; output section 307 including a liquid crystal display (LCD) and speakers, etc.; storage section 308 including a hard disk, etc.; and communication section 309 including a network interface card such as a LAN (Local Area Network) card and a modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.
[0162] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the various functions defined in the present invention.
[0163] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, wherein a computer-readable computer program is carried. The transmitted data signal can take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof.
[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0165] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the system described in the above embodiments; or it may exist independently and not assembled into the system. The storage medium carries one or more computer programs that, when executed by a processor of a system, cause the system to implement the methods provided in the above embodiments.
[0166] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0167] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0168] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0169] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A data management method for financial services, applied to a data management platform, characterized in that: include: Qualified Rongxin data is extracted from Rongxin data based on the latest data standards. These latest data standards are derived by adding field naming and field type validation to the original data standards using standard terminology. Specifically, they include: Construct a standard word root mapping library, which contains standard word roots for commonly used fields in the business domain and their corresponding field type sets; Based on the standard word root mapping library, the fields in the original data standard are segmented to obtain the field word root sequence; Match the field word root sequence with the standard word root; If the match is successful, then verify whether the verification field type of the root sequence of the field belongs to the set of field types; Mark fields that fail to match or whose field type validation fails, and generate field correction suggestions; The original data standard is updated based on the field correction suggestions to obtain the latest data standard; Obtain data quality management requirements, and determine the data quality range based on the data quality management requirements; If, based on the data quality range, it is determined that the qualified credit data does not meet the data quality management requirements, a rectification process will be implemented for the qualified credit data. Data sampling and testing are performed at preset progress points in the rectification process to obtain several sampling and testing results. The testing standard for the data sampling and testing is the data quality range. The sampling and testing results are qualified or unqualified. The number of preset progress points is a preset first number. If any of the preset progress points results in a non-compliant sampling inspection, then the steps of the rectification process for the compliant Rongxin data are repeated. If the sampling test results corresponding to the preset first number of preset progress points are all qualified, then the qualified Rongxin data after the rectification process is output. The rectified and qualified Rongxin data will be stored in the corresponding user's data management database.
2. The method according to claim 1, characterized in that, The process involves data sampling and testing at preset progress points in the rectification process to obtain several sampling and testing results, specifically including: A multi-dimensional scoring matrix is constructed based on the data quality range, and the multi-dimensional scoring matrix includes indicators of completeness, accuracy, consistency and timeliness. A stratified random sampling method was used to extract samples from the qualified credit data; The sample is scored in multiple dimensions based on the multi-dimensional scoring matrix to obtain the scoring results for each dimension. The scores of each dimension are weighted and calculated based on preset dimension weights to obtain a comprehensive score. The comprehensive score is compared with the preset qualified threshold to obtain the sampling test results.
3. The method according to claim 1, characterized in that, After storing the rectified and qualified Rongxin data in the corresponding user's data management database, the method further includes: Receive a data request sent by a user, the data request containing target data feature description information and usage scenario description information; The rectified and qualified Rongxin data is tagged to obtain data tag information, which includes feature identifiers that match the target data feature description information. A data usage permission matrix is generated based on the usage scenario description information, and the data usage permission matrix contains data access level information. A data access credential is generated based on the data access level information. The data access credential includes access validity parameters, scope of use parameters, and de-identification requirements parameters. Configure data access interface parameters according to the data retrieval credentials. The data access interface parameters are used to control the access method of the rectified and qualified Rongxin data.
4. The method according to claim 3, characterized in that, The step of tagging the rectified and qualified Rongxin data to obtain data tag information specifically includes: The target data feature description information is parsed to obtain a feature keyword sequence; Construct a preset data tagging system, which includes a subset of business attribute tags, a subset of time attribute tags, and a subset of scene attribute tags; Based on the characteristic keyword sequence, feature analysis is performed on the rectified and qualified credit data to generate an initial data tag set; The initial data tag set is filtered using preset tag optimization rules to obtain the target data tag set; Calculate the association strength coefficient between each tag in the target data tag set, generate a tag relationship network structure, and obtain data tag information.
5. The method according to claim 3 or 4, characterized in that, After configuring the data access interface parameters based on the data credential, the method further includes: Collect data call behavior characteristic information, which includes call timestamp, call frequency statistics and data usage method identifier; Set the call warning threshold parameters according to the preset call warning rules; Calculate the data reuse rate index based on the data call behavior characteristic information; Construct a data access scoring rule set, and quantitatively evaluate the compliance of data use based on the data access scoring rule set to obtain the compliance evaluation result; The data usage permission matrix is updated based on the compliance assessment results.
6. The method according to claim 5, characterized in that, The collected data retrieval behavior characteristic information specifically includes: Obtain data call log records, and extract the caller's identity information and operation behavior sequence from the data call log records; A time-series analysis is performed on the operation behavior sequence to obtain data call time parameters and call frequency parameters, and a data call link diagram is constructed based on the operation behavior sequence; Data usage characteristics are extracted from the data call chain diagram to obtain data usage identifiers; The data call time parameter, the call frequency parameter, and the data usage method identifier are combined to obtain data call behavior feature information.
7. A data management system for financial services, characterized in that, The system includes: One or more processors and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the system to perform the method as described in any one of claims 1-6.
8. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on the system, the system performs the method as described in any one of claims 1-6.
9. A computer program product, characterized in that, When the computer program product is run on the system, the system performs the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Data sampling detection method and device, equipment and storage medium
CN112860741A
Automatic data element management method
CN117313714A