Credit investigation data quality management method and system based on privacy computing technology
By adopting privacy request technology and SCQL engine in credit data quality management, the problem of insufficient security and timeliness of data verification work at both ends in the existing technology is solved, and efficient and secure data quality management is achieved.
Patent Information
- Application Number
- CN202510119352.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
In the existing credit data quality management, data verification work at both ends relies on manual operations, and there are problems of insufficient data interaction security and calculation timeliness, especially in high data security and large-scale data processing scenarios.
The credit data quality management method based on privacy computing technology is adopted, and the Privacy Receiverable Query Language (SCQL) system engine is used to realize the automation and efficiency of data verification tasks at both ends, ensuring the privacy protection of data during processing.
Through privacy computing technology, the security and computing timeliness of data interaction are significantly improved, manual operations are reduced, the efficiency and accuracy of data quality management are improved, and the risk of data leakage is reduced.
Smart Images

Figure BDA0005258515220000101 
Figure BDA0005258515220000111 
Figure BDA0005258515220000112
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, belonging to the category of information technology security and credit investigation system optimization. Specifically, it relates to a credit investigation data quality management method and system based on privacy computing technology. Background Art
[0002] The data quality management work of the credit investigation system needs to regularly conduct quantitative evaluations on the data quality of each credit investigation access institution, guide the credit investigation access institutions to attach importance to data quality work, and reduce misreported and unreported data. The two-end data verification work is an important means for the credit investigation system to carry out data quality management work, that is, by comparing the data at both ends of the access institution and the credit investigation system, the accuracy, timeliness, and integrity of the data reported by the access institution are always maintained at a high level. With the introduction of laws and regulations such as the Personal Information Protection Law and the Data Security Law, higher requirements have been put forward for the security, convenience, and computing efficiency of the data comparison work of credit investigation access institutions.
[0003] At present, the two-end data verification work of credit investigation data quality management mainly relies on manual operations, has many data transfer links, and involves problems such as the processing of a large amount of unmasked identification data. The data interaction security and computing timeliness need to be improved urgently.
[0004] Currently, there are mainly the following four technologies for two-end verification of data quality management:
[0005] (1) Direct data exchange
[0006] In this method, the two parties directly share the complete data set and then perform matching or comparative analysis. The direct data exchange method is intuitive and easy to implement conceptually. However, this method is not applicable to scenarios with high data security. Directly exchanging the data set means that both parties must fully expose their data content, which is extremely likely to lead to the leakage of sensitive information, especially in the absence of effective encryption measures.
[0007] (2) Hash-based comparison
[0008] In this method, by converting the data into hash values, the two parties can compare the hash values to identify common elements. As a method aimed at promoting the identification of common elements between the two parties by converting the data into hash values, its original intention is to improve the efficiency and security of data verification. However, although this method protects data privacy to a certain extent, it is vulnerable to rainbow table attacks and collision attacks. Especially when the data set is small, the original data is more likely to be collided out.
[0009] (3) Database query matching
[0010] This method relies on a third-party server to store data and conducts data verification through this server. This technology aims to achieve efficient data management and matching through a centralized platform. However, this technical path usually requires relying on a third-party database or cloud service, increasing the risk of centralized data storage and making it vulnerable to hacker attacks. Moreover, long-term use of third-party database services or cloud storage also incurs high costs.
[0011] (4) Traditional encryption technology
[0012] This method uses traditional encryption for raw data to ensure the security of data during transmission and storage. However, encryption and decryption often require a large amount of computing resources, especially when dealing with large-scale data sets. At the same time, the generation, distribution, and management of keys are complex and prone to becoming weak links in system security.
[0013] Therefore, the need for two-end verification work in data quality management of credit reporting access institutions is to comprehensively improve work efficiency by innovating work models and relying on innovative technologies to enhance data interaction security and computing timeliness. Summary of the Invention
[0014] The purpose of the present invention is to provide a credit reporting data quality management method based on privacy computing technology, with privacy computing technology as the underlying technical support, to innovate the data quality assessment work model, thereby further enhancing the automation of work tools, data localization, global management, and process convenience in data assessment work, achieving one-stop work implementation and dynamic supervision and management, thereby enhancing the security and timeliness of two-end verification work in credit reporting system data quality management, and further comprehensively improving the effectiveness of data quality assessment work for non-bank financial institutions.
[0015] Another purpose of the present invention is to provide a credit reporting data quality management system that adopts the above-mentioned credit reporting data quality management method based on privacy computing technology.
[0016] A credit reporting data quality management method based on privacy computing technology proposed in the present invention application specifically focuses on using Private Set Intersection (PSI) in privacy computing technology to further enhance privacy protection and efficiency in two-end verification of data quality management in the credit reporting system.
[0017] According to the project requirements, evaluate the matching degree between various privacy computing technologies and the project requirements, considering whether the technology can meet the requirements in aspects such as data protection, computing power, communication efficiency, and scalability. In the two - end verification business scenario, the credit investigation system side and the institutional side verify the integrity and consistency of the data at both ends through detailed message data. This scenario requires supporting the regular evaluation of credit investigation institutions connected to the credit investigation system without the original data of both parties leaving their respective domains. Considering that the possible number of institutions to be covered is in the thousands, concurrent computing is required during peak hours, and the computing tasks need to be queued, scheduled, and flow - limited. The user scale of institutions ranges from tens of thousands to tens of millions, and for institutions with a relatively large scale, the data volume is about one hundred million. During the process of data verification at both ends, multiple different multi - party secure computing logics need to be run for indicators such as the completeness rate and consistency rate. At the same time, in terms of performance, it is necessary to ensure that a single verification Structured Query Language (SQL) can be executed within 4 hours, otherwise network environment jitter may have a greater impact on the success rate; at the same time, it is necessary to ensure that all institutions that have been connected for two - end verification can complete the verification within the time window required by the credit investigation system. Therefore, the ability to horizontally expand is required when the business concurrency is high.
[0018] Based on the analysis of the above - mentioned business scenario, an open - source Secure Collaborative Query Language (SCQL) system engine can be used to build a privacy computing platform to conduct joint analysis without the original data of all parties leaving their respective domains.
[0019] A method for credit investigation data quality management based on privacy computing technology. This method builds a credit investigation data quality management system based on SCQL. The credit investigation data quality management system includes a credit investigation system side and an institutional side. A privacy computing node is respectively deployed on the credit investigation system side and the institutional side. The staff on the credit investigation system side and the staff on the institutional side respectively log in to their respective privacy computing nodes. On the premise that the pre - data authorization is completed, the verification task at both ends is completed by the initiative of one end of the credit investigation system and the cooperation of the other end of the institution.
[0020] The process of initiating from one end of the credit investigation system and cooperating with the other end of the institution to complete the two - end verification task includes: ordinary users on the credit investigation system side upload data within the credit investigation system of a single or multiple institutions to the local node in the system through single - item or batch upload methods. After passing data verification, a task list is generated and sent to the task management menu; the credit investigation system side automatically initiates an evaluation task to the institution side. After receiving the evaluation task, institution - side users upload the business system data of their own institutions to the institution - side node. After the system data verification passes, the system automatically matches the data uploaded by the credit investigation system side and the institution side, and starts calculating evaluation indicators for the two - end verification task. After the two - end verification task is completed, the two - end verification data results and difference files are generated for users on the credit investigation system side and the institution side to view and / or download.
[0021] The credit investigation data quality management system at least includes two layers: the platform application layer and the privacy computing engine layer. The platform application layer is used to provide a user interaction page to realize product function expression, and the privacy computing engine encapsulates the entire SQL execution logic.
[0022] The credit investigation system side is mainly used for system user management, institution node management, two - end data verification, evaluation report management, and evaluation statistics monitoring. The users on the credit investigation system side include system administrator users and system ordinary users. The system administrator users are responsible for creating system ordinary users and assigning corresponding permissions, and the system ordinary users are responsible for relevant business operations.
[0023] The institution side is mainly used for institution users to conduct two - end data verification and evaluation report management, etc. Institution users are non - bank small and micro financial institutions, and institution users are created by system administrator users.
[0024] Among them, the users on the credit investigation system side include one system administrator user and several system ordinary users. System ordinary users are assigned corresponding function permissions by system administrator users. Among institution users, each institution is assigned one or more institution - side users, and institution users are default configured with full permissions on the institution side.
[0025] The present invention also provides a credit investigation data quality management system that adopts the above - mentioned credit investigation data quality management method based on privacy computing technology.
[0026] A credit investigation data quality management system based on privacy computing technology. The credit investigation data quality management system is a privacy computing system built based on SCQL. The credit investigation data quality management system includes a credit investigation system side and an institution side. A privacy computing node is respectively deployed on each of the credit investigation system side and the institution side. Staff members on the credit investigation system side and staff members on the institution side log in to their respective privacy computing nodes. On the premise that pre-data authorization is completed, the credit investigation system side initiates and the institution side cooperates to complete the two-end verification task. Among them: The credit investigation system side is used for system user management, institution node management, data two-end verification, evaluation report management, evaluation statistics monitoring. Credit investigation system side users include system administrator users and system ordinary users. The system administrator users are responsible for creating system ordinary users and assigning corresponding permissions. System ordinary users are responsible for relevant business operations. The institution side is used for institution users to conduct data two-end verification and evaluation report management, etc. The institution users are non-bank small and micro financial institutions, and the institution users are created by system administrator users.
[0027] Among them, the credit investigation data quality management system includes: a user interface layer, a business logic layer, an adaptation layer, a privacy computing engine layer, a data access layer, and a data layer.
[0028] The user interface layer is mainly responsible for interacting with users and providing a friendly operation interface. The user interface layer usually includes front-end technologies such as HTML, CSS, and JavaScript, which are used to build web page layouts, styles, and dynamic effects.
[0029] The business logic layer is responsible for processing core functions and business rules. The business logic layer includes function modules such as login authentication, evaluation statistics monitoring, institution node management, user management, verification data management, verification task management, evaluation report management, and log management to ensure the correct execution of the system's business processes.
[0030] The adaptation layer, as a bridge between the business logic layer and the underlying services, is responsible for converting upper-layer requests into forms that the underlying services can understand. The adaptation layer includes functions such as engine data registration, engine task submission, engine result query, and other related functions.
[0031] The privacy computing engine layer is responsible for providing privacy computing capabilities. The privacy computing engine layer includes functions such as SQL parsing, multi-party secure computing (MPC) protocol interaction, and task running to ensure the protection of data during processing.
[0032] The data access layer is responsible for interacting with data and providing read and write operations for data. The data access layer includes functions such as data caching, transaction management, and reading and writing databases to ensure data consistency and integrity.
[0033] The data layer is responsible for storing actual data. The data layer usually uses a relational database (such as MySQL) or other file storage methods. The data layer is the foundation of the entire system, ensuring the secure storage and efficient retrieval of data.
[0034] Through these hierarchical divisions, the system can better achieve functional modularization, improving development efficiency and maintainability.
[0035] Among them, the credit investigation system end-users include a system administrator user and several system ordinary users. The system ordinary users are assigned corresponding functional permissions by the system administrator user. The functional permissions assigned to these system ordinary users can be the same, different from each other, or overlapping. The specific assignment rules can be defined according to actual operation requirements.
[0036] Among them, in the institutional users, each institution is assigned one or more institutional end-users, and the institutional users are configured with full permissions for the institutional end.
[0037] The credit investigation data quality management system based on privacy computing technology as described above includes the following modules:
[0038] (1) User login module
[0039] Users log in by entering their accounts and passwords. After the password is verified correctly, they are redirected to the user homepage corresponding to their user permissions.
[0040] (2) Evaluation statistics and monitoring module
[0041] Provides aggregated display and drill-down analysis of evaluation scores for the whole country and each region: After the data evaluation report is generated, the verified results will be finally aggregated and displayed according to the evaluation batch dimension; the system integrates the evaluation reports of each institution into statistical data (such as statistical values like average score, highest score, lowest score, etc.), supporting drill-down according to dimensions such as region, institution type, business type, etc.; through this way of display, the overall situation can be intuitively viewed, and the regions and institutions that need to be focused on can be located.
[0042] (3) System user management module
[0043] Used to manage system user permissions and related information, and implement functions such as system user creation, query, information maintenance, password reset, and deletion.
[0044] (4) Institutional node management module
[0045] The institutional node management module includes: institutional information management and institutional user management. Among them: institutional information management realizes the maintenance of basic information of the evaluation institution, automatically creates an institutional administrator account, and automatically allocates privacy computing nodes on the institutional side; institutional user management is used to maintain the user information of the institution, and provides functions for deleting, editing, locking, and resetting passwords for institutional administrator accounts.
[0046] (5) Data two - end verification module
[0047] Data two - end verification has three functions on the credit investigation system side, including: uploading verification data, verification task management, and verification result query. Ordinary users on the credit investigation system side upload the data in the credit investigation systems of single or multiple institutions to the local node in the system through single - pen or batch upload methods. After passing data verification, a task list is generated and sent to the task management menu; the credit investigation system side automatically initiates an evaluation task to the institutional side. After the evaluated institution receives the evaluation task, users on the institutional side upload the data of their institution's business system to the institution's node. After the system data verification passes, the system automatically matches the data uploaded by the credit investigation system side and the institutional side, and starts the two - end verification task to calculate evaluation indicators; after the two - end verification task is completed, the two - end verification data results and difference files are generated for users on the credit investigation system side and the institutional side to view and / or download.
[0048] Among them, data verification includes: file name verification, file content verification, and file format verification. The evaluation indicators include two aspects: integrity and accuracy. The integrity indicators are the credit business integrity rate and the lender integrity rate, and the accuracy indicators are the account consistency rate under the borrower's name and the two - end consistency rate of account balances.
[0049] Correspondingly, the data two - end verification module includes: an upload verification data sub - module, a verification task management sub - module, and a verification result query sub - module.
[0050] Upload verification data sub - module: Ordinary users on the credit investigation system side upload the data in the credit investigation systems of single or multiple institutions to the local node in the system through single - pen or batch upload methods. After passing data verification, a task list is generated and sent to the task management menu.
[0051] Verification task management sub - module: The credit investigation system side automatically initiates an evaluation task to the institutional side. After the evaluated institution receives the evaluation task, users on the institutional side upload the data of their institution's business system to the institution's node. After the system data verification passes, the system automatically matches the data uploaded by the credit investigation system side and the institutional side, and starts the two - end verification task to calculate evaluation indicators.
[0052] The task status of data two - end verification is as follows:
[0053] Institution not responding: Waiting for users on the institutional side to upload the verification data of their institution.
[0054] During verification in progress: When the institutional user successfully uploads institutional data and passes the data verification, the system matches the two-end nodes to generate a verification task and waits for the system to perform the two-end verification. In this state, the user can click "View Details" to view the status and results of the indicator calculation.
[0055] Verification failed: For tasks in progress during verification, if the execution fails, it enters the state of verification failure. In this state, the error reason can be viewed and the evaluation can be initiated again. If a retry is required due to system reasons, you can choose not to clear the evaluation data and retry directly; if it is due to data error reasons and the credit information system side needs to re-upload the data, you can choose to clear the data and re-evaluate. In this state, the user can click "View Details" to view the specific situation of the operation log.
[0056] Verification completed: The verification task is completed. In this state, the user can click "View Details" to view the specific situation of the operation results.
[0057] Verification result query sub-module: After the two-end verification task is completed, the two-end verification data results and difference files are generated for the credit information system side and institutional users to view and / or download.
[0058] The core process of two-end data verification includes the following:
[0059] When the message on the credit information system side is uploaded successfully and passes the data verification, the system automatically forms a two-end verification task and distributes it to the corresponding institutional side; the institutional side uploads the institutional-side verification message in accordance with the corresponding verification format requirements. After the institutional-side message completes the verification and confirmation, the credit information system side automatically matches the task at a fixed time, and the system automatically starts the two-end verification task to calculate the evaluation indicators; after the two-end verification task is completed, it supports the credit information system side users to view and download the two-end verification data results and difference files of single and batch institutions, and the verification result situation table in EXCEL format can be downloaded according to dimensions such as time point and business type.
[0060] (6) Evaluation report management module
[0061] After the verification task is completed, the ordinary user on the credit information system side will upload and import the timeliness indicator. After passing the data verification, the system accepts the timeliness indicator data. After verification and matching with the institutional batch and task, the timeliness indicator data will be combined with the results of the two-end data verification to generate the original data of the evaluation report; when the user sends a viewing request at the system front end, the system front end renders and generates the evaluation report according to the original data returned by the system back end; after the credit information system side generates the evaluation report, the system will synchronize the report to the corresponding institutional side according to the evaluation batch.
[0062] The rendered and generated evaluation report supports Word and / or PDF formats.
[0063] (7) System Log Management Module
[0064] Record all operations of all system users during the period from login to logout; users on the credit investigation system side can view and download corresponding system operation logs according to selected evaluation batches, operation types, institutions, etc.
[0065] (8) System Information Management Module
[0066] Used for the credit investigation system side to complete the addition, editing, and release management of system information through the system information management module, and the institution side supports corresponding operations after receiving the information.
[0067] A credit investigation data quality management method and a credit investigation data quality management system based on privacy computing technology proposed by the present invention build a privacy computing system based on SCQL, construct a multi-party data collaboration network, and realize the task of two-end verification without the source files of data quality evaluation at the credit investigation system side and the institution side leaving the local node, strengthening the security and efficiency of data flow, processing, and calculation, being safe and efficient; outputting data quality evaluation results, and feeding back the data integrity and accuracy of each institution to the credit investigation system side and the institution side, not only greatly reducing the risk of privacy leakage during data flow, processing, and calculation, but also greatly improving the accuracy and efficiency of data comparison and verification.
[0068] A credit investigation data quality management system based on privacy computing technology proposed by the present invention meets the basic functional requirements of data quality management, including evaluation statistics monitoring, system user management, institution information management, two-end data verification, evaluation report management, system log management, system announcement management, etc., supports users on the credit investigation system side to conveniently create institution and user information, the system automatically sends evaluation task information according to conditions such as evaluation batches, institution regions, and institution types, and real-time views evaluation results and manages evaluation reports.
[0069] A credit investigation data quality management system based on privacy computing technology proposed by the present invention has an innovative design at the architecture level, uses the SCQL engine as the core component, and realizes the efficient execution of customized index calculation. In terms of functional expansion, this system not only supports the accurate verification of two-end data, but also outputs the two-end verification difference result file, as well as the ability to automatically generate evaluation reports, thus comprehensively improving the efficiency of data quality management.
[0070] In addition, a credit investigation data quality management system based on privacy computing technology proposed by the present invention is applicable to the automatic verification of two-end data with different business magnitudes, supports configuring corresponding institution-side nodes according to business scale, network resources, and the informatization level of institutions, and gives test conclusions on the calculation performance of 50,000 - 80 million data under different software and hardware configuration environments.
[0071] A credit investigation data quality management method and a credit investigation data quality management system proposed by the present invention are supported by privacy computing technology as the underlying technology, innovate the data quality evaluation work mode, further improve the tool automation, data localization, management globalization, and process convenience of the data evaluation work, realize one-stop work implementation and dynamic supervision and management, and comprehensively improve the effectiveness of the data quality evaluation work of credit investigation access institutions.
[0072] A credit investigation data quality management method and a credit investigation data quality management system proposed by the present invention, due to adopting the above technical solutions, have the following advantages compared with the prior art:
[0073] High security: The privacy intersection (PSI) technology in privacy computing is adopted, significantly reducing the risk of data leakage, ensuring the secrecy of data during transmission, and providing a solid guarantee for information security.
[0074] Efficient algorithm: Combining with the scenario requirements, the ECDH-PSI algorithm is used to achieve efficient verification of large-scale databases with low consumption of computing resources and network resources, greatly improving the data processing speed.
[0075] Strong scalability: SCQL is used as the data analysis engine. SCQL can convert SQL statements into a mixed plaintext and ciphertext execution graph and execute on a joint database system. This language inherits the popularity, learnability, and high maturity of SQL as a common data analysis language.
[0076] Intelligent operation: Paying attention to the user experience, an automated and intelligent operation process is built, significantly reducing the learning threshold and ensuring that various users can easily get started.
[0077] In summary, the credit investigation data quality management method based on privacy computing technology proposed by the present invention shows advantages over the prior art in terms of privacy protection, security mechanism, processing efficiency, and user-friendliness. This technical solution has great novelty, practicability, and expansibility.
[0078] A credit investigation data quality management method and a credit investigation data quality management system proposed by the present invention can effectively improve the security management of evaluation data relying on privacy computing technology, provide technical guarantees for all aspects such as data transmission, calculation, and storage, and thus comprehensively improve the data security of the data evaluation work; by innovating the data quality evaluation mode and building a privacy computing data quality evaluation system to enhance the tool automation, data localization, management globalization, and process convenience of the evaluation work, it brings certain economic and social benefits.
[0079] Social benefits
[0080] Innovate the work mode through data quality assessment, develop a data comparison tool for the application of privacy computing technology, and achieve secure transmission of ciphertext and timely and efficient calculation in data quality assessment work, which can further improve the data security transmission efficiency in data quality assessment work, and reduce the manual approval and transfer links in the assessment work, so as to enhance the security and timeliness of data assessment work from both aspects of data transmission security and manual operation security.
[0081] On the one hand, the credit reporting system can always keep track of the overall situation of data quality assessment work in which its own role participates in real time, comprehensively improving the efficiency of data quality assessment work. On the other hand, it guides institutions to attach importance to data quality work, reduces misreported and unreported data, effectively helps the data governance work of institutions, enhances the awareness of credit reporting data security of institutions, and improves the credit reporting data quality management level from the source end of institutional data.
[0082] In addition, introducing systematic and automated data quality assessment means helps to improve the data quality supervision level, improve the credit reporting service infrastructure, strengthen the awareness of data leakage risk prevention, and then enhance the role of the credit reporting system in financial risk prevention and control, promote the construction of the social credit system, help the development of inclusive finance, and achieve positive results in optimizing the credit environment and supporting the development of the real economy.
[0083] Economic benefits
[0084] At present, the two-end verification work of data quality assessment in the credit reporting system mainly relies on manual operation, with low automation and room for improvement in timeliness, and heavy work tasks, resulting in a large demand for manpower. After the completion of the privacy computing technology application project for data quality assessment of non-bank financial institutions, the credit reporting system will significantly reduce the manual interaction process in the assessment work, improve the timeliness of data quality assessment work, and then relieve the manual cost pressure at all ends participating in data management, achieving an increase in economic benefits. When the innovative assessment mode is comprehensively promoted and continuously applied in the market, with the increase in the number of covered institutions, it can effectively reduce the data quality management cost of accessing institutions and generate certain economic benefits.
[0085] Using the new assessment work mode to regularly carry out data quality assessment work for institutions helps to improve the credit reporting data quality, guide institutions to attach importance to data submission and subsequent data quality management work, thus effectively reducing the occurrence of credit reporting disputes and complaints, effectively relieving the work pressure of the dispute handling and related work teams, and correspondingly reducing the corresponding workload; at the same time, the credibility and satisfaction of the credit reporting system's product services at the social level will also be improved, and then a certain scale of indirect economic benefits will be generated. Description of the Drawings
[0086] Figure 1This is the system hierarchical packaging architecture diagram of a credit investigation data quality management system based on privacy computing technology according to the present invention.
[0087] Figure 2 This is the core flowchart of data two - end verification of a credit investigation data quality management method based on privacy computing technology according to the present invention.
[0088] Figure 3 This is the flowchart for generating an evaluation report of a credit investigation data quality management method based on privacy computing technology according to the present invention. Detailed implementation manners
[0089] Through the following embodiments of the present invention and the description in combination with the drawings, other advantages and features of the present invention are shown. The embodiments are given in the form of examples, but are not limited thereto.
[0090] The present invention provides a credit investigation data quality management system adopting the above - mentioned credit investigation data quality management method based on privacy computing technology. This credit investigation data quality management system is a privacy computing system built based on SCQL. The credit investigation data quality management system includes a credit investigation system end and an institution end. A privacy computing node is respectively deployed at the credit investigation system end and the institution end. Staff members at the credit investigation system end and the institution end respectively log in to their respective privacy computing nodes. On the premise that the pre - data authorization is completed, the two - end verification task is completed by initiating from one end of the credit investigation system and cooperating with the other end of the institution. Among them: The credit investigation system end is used for system user management, institution node management, data two - end verification, evaluation report management, and evaluation statistics monitoring. Credit investigation system end users include system administrator users and system ordinary users. The system administrator users are responsible for creating system ordinary users and granting corresponding permissions. System ordinary users are responsible for relevant business operations. The institution end is used for institution users to conduct data two - end verification and evaluation report management, etc. The institution users are non - bank small and micro financial institutions, and the institution users are created by system administrator users.
[0091] This credit investigation data quality management system includes: a user interface layer, a business logic layer, an adaptation layer, a privacy computing engine layer, a data access layer, and a data layer.
[0092] User interface layer: Mainly responsible for interacting with users and providing a friendly operation interface. The user interface layer usually includes front - end technologies such as HTML, CSS, and JavaScript, which are used to construct web page layouts, styles, and dynamic effects.
[0093] Business logic layer: Responsible for processing core functions and business rules. The business logic layer includes function modules such as login authentication, evaluation statistics monitoring, institution node management, user management, verification data management, verification task management, evaluation report management, and log management, ensuring the correct execution of the system's business processes.
[0094] The adaptation layer: As a bridge between the business logic layer and the underlying services, it is responsible for converting upper-layer requests into forms that the underlying services can understand. The adaptation layer includes functions such as engine data registration, engine task submission, engine result query, and other related functions.
[0095] The privacy computing engine layer: Responsible for providing privacy computing capabilities. The privacy computing engine layer includes functions such as SQL parsing, multi-party secure computing (MPC) protocol interaction, and task execution, etc., to ensure that data is protected during processing.
[0096] The data access layer: Responsible for interacting with data and providing read and write operations on data. The data access layer includes functions such as data caching, transaction management, and reading and writing databases, etc., to ensure data consistency and integrity.
[0097] The data layer: Responsible for storing actual data. The data layer usually uses a relational database (such as MySQL) or other file storage methods. The data layer is the foundation of the entire system, ensuring the secure storage and efficient retrieval of data.
[0098] Among them, the credit investigation system end-users include a system administrator user and several system ordinary users. The system ordinary users are assigned corresponding functional permissions by the system administrator user. The functional permissions assigned to these system ordinary users can be the same, or each can be different, or there can be overlaps. The specific assignment rules can be defined according to actual operation requirements.
[0099] Among them, among the institutional users, each institution is assigned one or more institutional end-users, and the institutional users are configured with full permissions for the institutional end.
[0100] This credit investigation data quality management system includes the following modules:
[0101] (1) User login module
[0102] Users log in by entering their account numbers and passwords. After the password is verified correctly, they are redirected to the user home page corresponding to their user permissions.
[0103] The specific functional permission distinctions for credit investigation system end-users and institutional end-users are shown in Table 1 and Table 2 respectively, and the user accounts and system data management scopes are shown in Table 3.
[0104] Table 1: Credit Investigation System End - User Functional Permissions
[0105]
[0106]
[0107] Table 2: Institutional End - User Functional Permissions
[0108]
[0109] Table 3: Scope of Account and System Data Management
[0110]
[0111] (2) Evaluation and Statistics Monitoring Module
[0112] Provide aggregated display and drill-down analysis of evaluation scores for the whole country and each region.
[0113] After the generation of the data evaluation report, the verified results will be finally aggregated and displayed according to the evaluation batch dimension; the system will integrate the evaluation reports of each institution into statistical data (statistical values such as average score, highest score, lowest score, etc.), and support drill-down according to dimensions such as region, institution type, business type, etc.; through this way of display, the overall situation can be intuitively viewed, and the regions and institutions that need to be focused on can be located.
[0114] (3) System User Management Module
[0115] Used to manage system user permissions and related information, and implement functions such as creating, querying, information maintenance, password reset, and deleting system users.
[0116] (4) Institution Node Management Module
[0117] Including: institution information management and institution user management. Among them: institution information management realizes the maintenance of basic information of evaluation institutions, automatically creates institution administrator accounts, and automatically allocates privacy computing nodes at the institution end; institution user management is used to maintain the user information of institutions, and provides functions such as deleting, editing, locking, and resetting passwords for institution administrator accounts.
[0118] (5) Data Verification Module at Both Ends
[0119] The data verification module at both ends includes: an upload verification data sub-module, a verification task management sub-module, and a verification result query sub-module.
[0120] Upload verification data sub-module: Ordinary users on the credit investigation system side upload data in the credit investigation systems of single or multiple institutions to the local nodes in the system through single or batch upload methods. After passing the data verification, a task list is generated and sent to the task management menu.
[0121] Verification Task Management Sub-module: The credit investigation system automatically initiates an evaluation task to the institutional side. After the evaluated institution receives the evaluation task, the user on the institutional side uploads the business system data of the institution to the institution node. After the system data verification passes, the system automatically matches the data uploaded by the credit investigation system side and the institutional side, and starts the two-end verification task to calculate the evaluation indicators.
[0122] The task status of the two-end data verification is as follows:
[0123] Institution Not Responding: Waiting for the user on the institutional side to upload the verification data of the institution.
[0124] Verification in Progress: The user on the institutional side has successfully uploaded the institutional data and passed the data verification. The system matches the two nodes to generate a verification task and waits for the system to perform the two-end verification. In this state, the user can click "View Details" to view the status and results of the indicator calculation.
[0125] Failed to Run: For a task in progress during verification, if the execution fails, it enters the state of failed to run. In this state, the error reason can be viewed and the evaluation can be initiated again. If it needs to be retried due to system reasons, you can choose not to clear the evaluation data and directly retry; if it is due to data error reasons and the credit investigation system side needs to re-upload the data, you can choose to clear the data and then re-evaluate. In this state, the user can click "View Details" to view the specific situation of the operation log.
[0126] Verification Completed: The verification task is completed. In this state, the user can click "View Details" to view the specific situation of the operation result.
[0127] Verification Result Query Sub-module: After the two-end verification task is completed, the two-end verification data result and the difference file are generated for the users on the credit investigation system side and the institutional side to view and / or download.
[0128] Among them, data verification includes: file name verification, file content verification, and file format verification.
[0129] File Name Verification Rule: QY / GR + 14-digit institution code + full institution name.txt (QY / GR represents the enterprise / individual business type); the example is "QYM21352900L0001 Wahaha Financial Leasing Co., Ltd..txt".
[0130] File Content Verification Rule: The content of the first line should contain "financial institution code, business number (account identification code), account opening date, detailed business type (major loan and borrowing business category), balance, name (enterprise name), document type (enterprise identity identification type), document number (enterprise identity identification number)" without abnormal values.
[0131] File format verification rules: TXT (UTF-8) or CSV (UTF-8). TXT text format is preferred, and UTF-8 encoding format is verified (or change the encoding format through "Save As").
[0132] Among them, the evaluation indicators include two aspects: integrity and accuracy. The integrity indicators include the credit business integrity rate and the lender integrity rate, and the accuracy indicators include the consistency rate of the accounts under the borrower's name and the consistency rate of the account balances at both ends.
[0133] Among them, the core process of data verification at both ends is as follows:
[0134] When the credit reporting system uploads the message successfully and passes the data verification, the system automatically forms a verification task at both ends and distributes it to the corresponding institutional end; the institutional end uploads the institutional end verification message in the corresponding verification format requirements. After the institutional end message completes the verification and confirmation, the credit reporting system end automatically matches the task regularly, and the system automatically starts to calculate the evaluation indicators for the verification task at both ends; after the verification task at both ends is completed, it supports the users of the credit reporting system end to view and download the verification data results and difference files of single and batch institutions, and the verification result situation table in EXCEL format can be downloaded according to dimensions such as time points and business types.
[0135] (6) Evaluation report management module
[0136] After the verification task is completed, ordinary users of the credit reporting system end will upload and import the timeliness index. After passing the data verification, the system accepts the timeliness index data. After verification and matching with the institutional batch and task, the timeliness index data will be combined with the results of data verification at both ends to generate the original data of the evaluation report.
[0137] When clicking to view, the front end renders and generates the evaluation report according to the original data returned by the back end (supporting Word and PDF formats).
[0138] After the evaluation report is generated at the credit reporting system end, the system will synchronize the report to the corresponding institutional end according to the evaluation batch.
[0139] (7) System log management module
[0140] Record all operations of all system users during the period from login to logout; users of the credit reporting system end can view and download the corresponding system operation logs according to the selected evaluation batch, operation type, institution, etc.
[0141] (8) System information management module
[0142] It is used for the credit reporting system end to complete the addition, editing and release management of system information through the system information management module, and the institutional end supports corresponding operations after receiving the information.
[0143] Architectural design
[0144] In the system architecture design, the system is divided into a user interface layer, a business logic layer, an adaptation layer, a privacy computing engine layer, a data access layer, and a data layer.
[0145] The user interface layer provides user interaction pages, realizes product function expression. The iterative changes mainly come from business requirement drive. In the iteration, it adopts a waterfall model to move forward quickly, focuses on horizontal expansion ability, and abstracts and precipitates the task status data model downward, and the centralized database storage is used to ensure consistency.
[0146] The privacy computing engine encapsulates the entire SQL execution logic, including task scheduling and underlying confidential communication protocols, etc., and has strict control over performance and stability. Through the division of these layers, the system can better realize function modularization and improve development efficiency and maintainability.
[0147] In the networking architecture, the credit investigation system end and the institutional end form a star topology structure through an interactive gateway.
[0148] In the database design, the database design specifications of this system are shown in Table 4.
[0149] Table 4: Database Design Specification Table
[0150]
[0151] Software and hardware configuration
[0152] The software and hardware configuration of the application system is shown in Table 5:
[0153] Table 5: Software and Hardware Environment Configuration Table
[0154] Category Specific content Programming language java, js, css, html, c++, go Integrated development tool Idea, antd, vsCode, etc. Server 8c16g500g specification, centos8.0 operating system Database MySQL 8.0 Management tool git Third-party main libraries springboot, sofa-ark, etc.
[0155] The application system runs in the docker container mode, and the specific running environment list is shown in Table 6.
[0156] Table 6: External Environment Dependence Table
[0157] Category Description CPU Support AVX512 instruction set Operating system CentOS Linux release 8+ or Red Hat Enterprise Linuex Server release 7.5+ docker 20.10+ openssl 1.1.1k jdk Open jdk1.8 mysql 8.0 Clock synchronization ntp HTTP protocol support http1.1 Operating user permissions Ordinary users, but need to be able to execute docker–privileged=true with super privileges Bandwidth requirement Expected minimum running bandwidth 2Mb, latency not exceeding 15ms Proxy gateway forwarding packet size The gateway supports sending content with a Body <= 2MB Firewall security policy Ensure that private computing network communication is not intercepted by the firewall Reverse proxy (optional) Multi-instance deployment, a reverse proxy (such as nginx, LVS) is required at the backend
[0158] The operation and maintenance items include the operation and maintenance of infrastructure such as servers, networks, and databases, and the operation and maintenance of the system platform. It is necessary to equip common operation and maintenance tools such as ssh and jmap to facilitate the collection of key indicators and logs of the system platform, which is convenient for problem troubleshooting and solution.
[0159] Performance evaluation
[0160] Currently, the frequency of two - end verification in the credit reporting system is once a month for each institution (the peak of verification operations is generally concentrated within 2 weeks). In terms of system performance, it is required to complete the verification of all institutions within 10 working days. If there are N access institutions, then on average, it should be able to complete the two - end data verification of N / 10 institutions per day. Among them, for institutions with a large customer scale, the number of data records to be verified is approximately tens of millions, and each multi - party security verification takes several hours; for institutions with a small customer scale, the number of verified data is approximately tens of thousands, and each multi - party security verification takes several minutes. From the perspective of system performance and business volume, multi - party security verification can meet the basic requirements of credit reporting data verification.
[0161] At the same time, when the evaluation system provides elastic scaling, task queuing mechanism, and network traffic limiting capabilities during peak and off - peak business hours, it can ensure the effective allocation of system resources and achieve business goals during the peak period of verification operations.
[0162] Taking the intersection calculation required for two - end verification as an example, under the recommended configurations corresponding to machines with different data scales, the measured computing performance is shown in Table 7.
[0163] Table 7: Measured Results under Different Data Volumes and System Specifications
[0164] Check the amount of data (number of rows) Server specification Network configuration Total time taken for eight rules Longest time taken for a single rule 5W 8C16G500G 1Mb 15ms latency 222s 81s 10W 8C16G500G 1Mb 15ms latency 464s 161s 100W 8C16G500G 1Mb 15ms latency Approximately 1.4h 31.2m 1000W 8C32G500G 10Mb 15ms latency Approximately 5h 1.4h 8000W 16C64G500G 20Mb 15ms latency Approximately 9h 1.5h 8000W 16C64G500G 100Mb 15ms latency 3.2h 0.8h
[0165] Remarks: The data in the table was obtained by deploying two servers in the intranet environment to simulate the credit reporting system side and the institution side respectively, and submitting requests for running verification metrics from the credit reporting system side. The same metric requests were run continuously 5 times, and then the average value of the duration was calculated. Under the minimum configuration of 8C16G (8 CPU cores, 16G of memory) and 1Mb bandwidth, the verifiable data volume that can run stably is 1 million (for data volumes of 50,000 - 100,000, performance - first settings can be adopted, and stability is guaranteed; for 1 million data, it is necessary to switch to a stability - first configuration to ensure the success rate, but a certain execution time needs to be sacrificed). When the data volume grows to the tens of millions level, the required network bandwidth and machine specifications need to be correspondingly improved. Among them, an 8C16G500G (8 CPU cores, 16G of memory, 500G of storage) server can support up to 10 million data verifications at most. Theoretically, it is estimated that for every additional 10 million data volume, 1C2G (1 CPU core, 2G of memory) of additional memory is required. After 12C, the memory increase is the main factor. Multi - task parallelism can increase the task parallelism by linearly expanding the CPU and memory.
[0166] Although the present invention has been described above according to the preferred embodiments, this does not mean that the scope of the present invention is only limited to the above structures. As long as those skilled in the art of this technology can easily develop equivalent alternative structures after reading the above description, all equivalent changes and modifications made without departing from the spirit and scope of the present invention should be covered within the scope of the present invention patent.
Claims
1. A credit data quality management method based on privacy computing technology, characterized in that: The method is to build a credit data quality management system based on SCQL, which includes a credit system end and an institutional end. A privacy computing node is deployed on each of the credit system end and the institutional end. The staff of the credit system end and the institutional end log in to their respective privacy computing nodes respectively. On the premise that the pre-data authorization is completed, the credit system end initiates and the institutional end cooperates to complete the verification task at both ends. The task of completing the verification at both ends is initiated by the credit reporting system at one end and coordinated by the institution at one end, including: ordinary users at the credit reporting system end upload the data in the credit reporting system of a single or multiple institutions to the local node in the system through single or batch uploading methods, and generate a task list to the task management menu after the data verification is passed; the credit reporting system end automatically initiates an evaluation task to the institution end, and after the evaluated institution receives the evaluation task, the user at the institution end uploads the business system data of the institution to the node of the institution, and after the system data verification is passed, the system automatically matches the data uploaded by the credit reporting system end and the institution end, and starts the verification task at both ends to calculate the evaluation indicators; after the verification task at both ends is completed, the verification data results and difference files at both ends are generated for viewing and / or downloading by users at the credit reporting system end and the institution end; The credit data quality management system includes at least two layers: a platform application layer and a privacy computing engine layer. The platform application layer is used to provide user interaction pages and realize product function expression. The privacy computing engine encapsulates the entire SQL execution logic. in: The credit reporting system end is used for system user management, institution node management, data verification at both ends, assessment report management, and assessment statistics monitoring. The credit reporting system end users include system administrator users and system ordinary users. The system administrator user is responsible for creating system ordinary users and granting corresponding permissions, and the system ordinary users are responsible for related business operations; The institutional end is used by institutional users to perform data verification at both ends and evaluation report management, etc. The institutional users are non-bank small and micro financial institutions, and the institutional users are created by system administrator users.
2. The credit information data quality management method according to claim 1, characterized in that: The credit reporting system end users include a system administrator user and several system ordinary users, and the system ordinary users are allocated corresponding functional permissions through the system administrator user.
3. The credit information data quality management method according to claim 1, characterized in that: Among the institutional users, each institution is assigned one or more institutional end users, and the institutional users configure full permissions of the institutional end.
4. A credit data quality management system based on privacy computing technology, characterized in that: The credit data quality management system is a privacy computing system built based on SCQL. The credit data quality management system includes a credit system end and an institutional end. A privacy computing node is deployed on each of the credit system end and the institutional end. The staff of the credit system end and the institutional end log in to their respective privacy computing nodes respectively. On the premise that the pre-data authorization is completed, the credit system end initiates and the institutional end cooperates to complete the verification task at both ends; wherein: The credit reporting system end is used for system user management, institution node management, data verification at both ends, assessment report management, and assessment statistics monitoring. The credit reporting system end users include system administrator users and system ordinary users. The system administrator user is responsible for creating system ordinary users and granting corresponding permissions, and the system ordinary users are responsible for related business operations; The institutional end is used for institutional users to check both ends of data and manage evaluation reports, etc. The institutional users are non-bank small and micro financial institutions, and the institutional users are created by system administrator users; The credit data quality management system includes: The user interface layer is mainly responsible for interacting with users and providing a friendly operation interface; Business logic layer, responsible for processing core functions and business rules; The adaptation layer, as a bridge between the business logic layer and the underlying services, is responsible for converting upper-layer requests into a form that the underlying services can understand; The privacy computing engine layer is responsible for providing privacy computing capabilities. The privacy computing engine layer includes SQL parsing, multi-party secure computing (MPC) protocol interaction, and task running to ensure that data is protected during processing; The data access layer is responsible for interacting with the data and providing data read and write operations. The data access layer includes data caching, transaction management, and reading and writing databases to ensure data consistency and integrity; and The data layer is responsible for storing the actual data.
5. The credit information data quality management system according to claim 4, characterized in that: The credit reporting system end users include a system administrator user and several system ordinary users, and the system ordinary users are allocated corresponding functional permissions through the system administrator user.
6. The credit information data quality management system according to claim 4, characterized in that: Among the institutional users, each institution is assigned one or more institutional end users, and the institutional users configure full permissions of the institutional end.
7. The credit information data quality management system according to claim 4, characterized in that: include: (1) User login module The user logs in by entering the account and password. After the password is verified, the user will be redirected to the user homepage corresponding to the user's permissions. (2) Assessment statistics monitoring module Provides aggregated display and drill-down analysis of national and regional assessment scores: After the data evaluation report is generated, the verification results will be aggregated and displayed according to the evaluation batch dimension; The system integrates the evaluation reports generated by each institution into statistical data, and supports drilling down by dimensions such as region, institution type, and business type; Through this kind of display, the overall situation can be viewed intuitively, and the areas and institutions that need special attention can be located; (3) System User Management Module Used to manage system user permissions and related information, and to implement system user creation, query, information maintenance, password reset and deletion functions; (4) Organization node management module Including: institutional information management and institutional user management; among which: Institutional information management realizes the maintenance of basic information of assessment institutions, automatically creates institutional administrator accounts, and automatically allocates privacy computing nodes on the institutional side; Institutional user management is used to maintain the user information of the institution, and provides functions for deleting, editing, locking, and resetting passwords for institution administrator accounts; (5) Data verification module at both ends Ordinary users of the credit reporting system upload data from the credit reporting systems of a single institution or multiple institutions to the local node in the system through single or batch upload. After the data is verified, a task list is generated and put into the task management menu; The credit reporting system automatically initiates an assessment task to the institution. After the assessed institution receives the assessment task, the institutional user uploads the institution's business system data to the institution's node. After the system data is verified, the system automatically matches the data uploaded by the credit reporting system and the institution, and starts the verification task at both ends to calculate the assessment indicators. After the verification task is completed, the verification data results and difference files are generated for users of the credit reporting system and institution to view and / or download; (6) Assessment report management module After the verification task is completed, ordinary users on the credit reporting system end will upload and import the timeliness index. After the data verification is passed, the system accepts the timeliness index data and verifies that it matches the institution's batch and task. Then, the timeliness index data is combined with the verification results at both ends of the data to generate the original data for the evaluation report. When the user issues a viewing request on the system front end, the system front end renders and generates an evaluation report based on the original data returned by the system back end; After the credit reporting system generates the assessment report, the system will synchronize the report to the corresponding institution according to the assessment batch; (7) System log management module Record all operations of all system users from login to logout; Users of the credit reporting system can view and download corresponding system operation logs according to the selected assessment batch, operation type, institution, etc. (8) System Information Management Module The credit reporting system is used to complete the addition, editing and publishing management of system information through the system information management module. The institution supports corresponding response operations after receiving the information.
8. The credit information data quality management system according to claim 7, characterized in that: When the user of the credit reporting system checks and / or downloads the verification data results and difference files at both ends, he can download the verification result table in EXCEL format according to dimensions such as time point and business type.
9. The credit information data quality management system according to claim 7, characterized in that: The evaluation indicators include integrity indicators and accuracy indicators. The integrity indicators include: credit business integrity rate, lender integrity rate, and the accuracy indicators include: consistency rate of accounts under the name of the borrower, and consistency rate of both ends of the account balance.
10. The credit information data quality management system according to claim 7, characterized in that: The evaluation report generated by the rendering supports Word and / or PDF formats.