Data circulation and security protection double-guarantee privacy computing conjoint analysis system
Through real-time data collection, dynamic storage and privacy computing technology, the problem that traditional data analysis systems cannot adapt to dynamic changes has been solved, the real-time nature and privacy protection of data have been achieved, and the decision-making ability and operational efficiency of enterprises have been improved.
Patent Information
- Application Number
- CN202510830961.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional data analysis systems are unable to adapt to the dynamically changing data environment in a timely manner, resulting in poor data real-time performance and delayed analysis results. They are unable to conduct effective evaluations based on the latest data and pose a risk of privacy leakage.
Real-time data collection, dynamic storage and privacy computing technologies are introduced, multi-channel data capture, distributed storage, encryption algorithms and enterprise labeling technologies are adopted, combined with machine learning models for in-depth analysis, to generate enterprise portraits and credit risk assessments, ensuring the real-time and privacy security of data.
It improves the real-time and accuracy of data analysis, protects user privacy, provides high-quality corporate portraits and credit risk assessments, optimizes resource allocation, and enhances the market competitiveness and operational efficiency of enterprises.
Smart Images

Figure CN120705894A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of privacy computing technology, and specifically to a privacy computing joint analysis system that provides dual guarantees for data circulation and security protection. Background Art
[0002] With the advent of the big data era, various enterprises and organizations generate a large amount of behavioral data and user information in their daily operations. This data provides enormous value for corporate decision-making, risk assessment, and market analysis. However, the widespread use of data also brings risks of privacy leakage and data security, especially when it involves user personal information and sensitive corporate data.
[0003] Traditional data analysis systems often rely on static processing methods and fixed parameter configurations, which makes them powerless in the face of an ever-changing data environment. This static approach cannot adapt to dynamic changes in data sources, such as shifts in user behavior patterns, fluctuations in market conditions, and the emergence of new data types. This limitation leads to poor real-time data and delayed analysis results, making it impossible for companies to make effective assessments based on the latest data when making key decisions. Summary of the Invention
[0004] (1) Technical problems solved
[0005] In response to the shortcomings of the existing technology, the present invention provides a privacy computing joint analysis system with dual guarantees for data circulation and security protection. By introducing real-time data collection, dynamic storage and privacy computing technologies, the efficiency and accuracy of data analysis are significantly improved. Compared with traditional static data analysis systems, the system can flexibly adapt to the ever-changing data environment and promptly reflect changes in user behavior patterns, market conditions and new data types. This dynamic response capability ensures the real-time nature of the data, enabling enterprises to make more accurate decisions based on the latest information. In addition, the privacy computing module uses encryption algorithms and enterprise labeling technology to effectively protect user privacy, and provide high-quality enterprise portraits and credit risk assessments during in-depth analysis. Therefore, enterprises can optimize resource allocation, reduce risks, enhance market competitiveness, and ultimately achieve higher operating efficiency and stronger decision-making capabilities while ensuring compliance.
[0006] (2) Technical solution
[0007] To achieve the above objectives, the present invention provides the following technical solutions: a privacy computing joint analysis system with dual guarantees for data circulation and security protection, comprising a data acquisition module, a data storage module, a privacy computing module, a data analysis module, and a result output module;
[0008] The data collection module is used to collect data in real time from different data sources, including basic enterprise information, enterprise behavior data, user access logs and system operation status, and transmit the real-time collected data to the data storage module;
[0009] The data storage module is used to store the real-time collected data transmitted by the data acquisition module, and calculate the data storage capacity utilization rate and data query response time to evaluate the usage of storage resources and the efficiency of data retrieval;
[0010] The privacy computing module is used to perform privacy calculation and analysis on the real-time collected data in the data storage module, using encryption algorithms to protect user privacy, executing enterprise tagging technology to generate enterprise profiles, and performing fraud detection based on the two-way elimination F test method. It also calculates the number of enterprise tags and fraud risk scores to assess the level of detail of enterprise profiles and identify potential fraudulent behavior;
[0011] The data analysis module is used to conduct in-depth analysis of the results generated by the privacy calculation module, use machine learning models to analyze corporate credit risks, calculate corporate credit scores, and generate corporate risk warning reports and corporate profile reports;
[0012] The result output module is used to send the enterprise risk warning report and the enterprise portrait report to the background application terminal.
[0013] Preferably, the data storage capacity utilization rate calculation formula is as follows:
[0014]
[0015] In the formula, Cs represents the data storage capacity utilization rate, C used Indicates the used storage capacity, C total Indicates the total storage capacity.
[0016] Preferably, the calculation formula for the data query response time is as follows:
[0017]
[0018] In the formula, T avg Indicates the data query response time, N indicates the total number of queries, T i Indicates the response time of the i-th query, where i is the index subscript.
[0019] Preferably, the data storage module is equipped with a dynamic storage optimization algorithm for automatically adjusting storage strategies according to real-time data changes, while automatically cleaning up expired and low-value data.
[0020] Preferably, the encryption algorithm used to protect user privacy is as follows:
[0021] C=Enc K (M)
[0022] In the formula, C represents the encrypted data, M represents the information to be encrypted, K represents the encryption key, and Enc represents the encryption function.
[0023] Preferably, the algorithm used to generate the enterprise portrait is as follows:
[0024] Q c =f(X)={l1,l2,...,l k}
[0025] Among them, Q c represents the enterprise portrait, X represents the enterprise feature vector, f represents the labeling algorithm model, l1, l2, ..., l k Indicates each acquired enterprise label.
[0026] Preferably, the calculation formula for the number of enterprise tags is as follows:
[0027]
[0028] In the formula, L represents the total number of labels generated by the enterprise, R represents the number of feature categories, and l j,k represents the kth label of the jth class feature, x j,k represents the kth specific eigenvalue in the jth class feature, w j,k Represents the weight coefficient of the kth feature of the jth category, which is obtained based on feature importance training, T j represents the threshold of the j-th class feature, and |*| represents the total number of labels.
[0029] Preferably, the calculation formula for the fraud risk score is as follows:
[0030]
[0031] In the formula, Fx represents the fraud risk score, x f represents the fraud-related feature vector, represents the model parameters corresponding to the fraud-related feature vector, b f Represents the bias term, σ represents the sigmoid function, and converts the output into a probability in the (0-1) range.
[0032] Preferably, the calculation formula for the enterprise credit score is as follows:
[0033]
[0034] In the formula, S represents the enterprise credit score, σ represents the sigmoid function, and converts the output into a probability in the range of (0-1).c represents the credit-related feature vector, represents the model parameters corresponding to the credit-related feature vector, b c Indicates the offset.
[0035] Preferably, the result output module is equipped with an intelligent report generator for automatically generating customized enterprise risk warning reports and enterprise portrait reports according to user needs.
[0036] Compared with the existing technology, this invention provides a privacy computing joint analysis system that ensures both data flow and security protection, and has the following beneficial effects:
[0037] This invention significantly improves the efficiency and accuracy of data analysis by introducing real-time data collection, dynamic storage and privacy computing technologies. Compared with traditional static data analysis systems, this system can flexibly adapt to the ever-changing data environment and promptly reflect changes in user behavior patterns, market conditions and new data types. This dynamic response capability ensures the real-time nature of the data, enabling enterprises to make more accurate decisions based on the latest information. In addition, the privacy computing module uses encryption algorithms and enterprise labeling technology to effectively protect user privacy, while providing high-quality enterprise portraits and credit risk assessments during in-depth analysis. Therefore, enterprises can optimize resource allocation, reduce risks, enhance market competitiveness, and ultimately achieve higher operating efficiency and stronger decision-making capabilities while ensuring compliance. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 Schematic diagram of the system flow of the present invention. DETAILED DESCRIPTION
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0040] Traditional data analysis systems often rely on static processing methods and fixed parameter configurations, failing to effectively cope with dynamically changing data environments. This leads to poor real-time data and delayed analysis results. Therefore, we propose a privacy-preserving joint analysis system that provides both data flow and security protection. Figure 1 ,The system includes a data acquisition module, a data storage module, a privacy ,computing module, a data analysis module and a result output module;
[0041] The data collection module uses multiple channels, multiple protocols, and multiple technologies to capture and transmit real-time data from different data sources. It utilizes a bridging tool based on ETL (Extract-Transform-Load) technology and employs multiple communication protocols, such as RESTful API, WebSocket, and MQTT, to ensure efficient and stable acquisition of multi-dimensional data, including basic enterprise information, enterprise behavior data, user access logs, and system operation status. During the data collection process, incremental synchronization and event-driven mechanisms are used to achieve real-time data updates. Pre-processing processes such as data filtering, cleaning, deduplication, and format normalization are used to ensure data integrity and consistency. To enhance the security and reliability of data transmission, the module uses SSL / TLS encryption protocols for communication, combined with message queue buffering technology and retry strategies to improve data transmission stability in high-concurrency and network fluctuation environments. After local pre-processing, the collected real-time data is pushed to the data storage module in real time via a high-performance API interface or message middleware, ensuring that the data is synchronously stored in the system in the shortest possible time, providing timely and accurate basic data support for subsequent analysis and processing.
[0042] The data storage module adopts a multi-level, multi-architecture storage system structure, combining relational databases (such as MySQL, PostgreSQL) and non-relational databases (such as MongoDB, Redis) and other storage technologies to ensure efficient, structured and diversified storage of enterprise profile information, enterprise credit data and monitoring logs. At the same time, it uses distributed storage technology and horizontal expansion capabilities to dynamically adjust storage capacity to adapt to the growing data scale, adopts a hot and cold separation strategy to achieve hierarchical management of data, and introduces a storage capacity monitoring mechanism within the system. By calculating the storage space utilization rate in real time, the formula is adopted:
[0043]
[0044] The ratio of occupied storage space to total capacity is monitored to provide timely warnings of resource bottlenecks or expansion needs and optimize resource allocation. In addition, the storage module is equipped with efficient indexing strategies and partition management technologies to reduce data retrieval response time. To ensure data retrieval efficiency, the system continuously monitors query response time according to the formula:
[0045]
[0046] Calculate the average response time, where T irepresents the response time of the i-th query, where N is the number of queries. By monitoring this metric, we can identify retrieval bottlenecks, optimize index structures, improve retrieval efficiency, and ensure that users and applications can quickly meet their real-time data access needs. Overall, this design not only ensures data storage capacity and efficiency but also provides a quantitative basis for the rational allocation of storage resources and system performance optimization through scientific monitoring indicators, thereby supporting the continuous and stable operation of enterprise data analysis and intelligent decision-making.
[0047] The privacy computing module utilizes a variety of advanced privacy protection technologies, including encryption algorithms such as homomorphic encryption, differential privacy, and multi-party secure computing, to ensure that user privacy is fully protected when analyzing and processing sensitive data stored in the data storage module, thereby avoiding the risk of information leakage. In the process of generating enterprise portraits, the system combines multi-source and multi-dimensional enterprise characteristics, extracts key indicators through enterprise labeling technologies such as deep learning and feature fusion, and automatically constructs a multi-level and multi-dimensional label set for the enterprise, thereby achieving a detailed description of enterprise characteristics. In order to reflect the richness and depth of detail of the enterprise portrait, the system introduces a calculation formula for the number of enterprise labels:
[0048]
[0049] Among them, R represents the number of feature categories (such as finance, market, behavior, etc.), and the label l under each category j,k According to the eigenvalue x j,k With weight w j,k , and pass the threshold T j Screening, filtering out significant labels, and finally obtaining the total number of labels of different dimensions by taking the union, reflecting the richness and diversity of corporate portraits. In order to enhance the ability to identify potential fraudulent behavior of enterprises, the system adopts a fraud detection method based on the two-way elimination F test to determine potential fraud risks. At the same time, the system calculates the fraud risk score of the enterprise:
[0050]
[0051] Among them, x f is the fraud-related feature vector, is the feature weight, b f is the bias term, and σ is the sigmoid function. Through the above technical means, not only can a detailed and accurate corporate profile be constructed while ensuring user privacy and security, but potential fraudulent behavior of enterprises can also be effectively identified, thus improving the intelligent level of risk management and supervision.
[0052] The data analysis module fully utilizes the enterprise portraits, risk detection indicators, and feature data generated by the privacy computing module, and uses advanced machine learning models (such as random forests, support vector machines, deep neural networks, etc.) to conduct in-depth analysis to assess the enterprise's credit risk. Through the trained model parameters, the system can integrate the enterprise's financial status, behavioral characteristics, and potential risk indicators to calculate the enterprise's credit score, the mathematical expression of which is:
[0053]
[0054] Among them, x c A multidimensional feature vector representing an enterprise, is the corresponding model weight vector, b c is the bias term, and σ is the sigmoid function, which is used to convert the model output into a credit risk probability between 0 and 1. In addition to credit scoring, the analysis module automatically generates enterprise risk warning reports based on multivariate analysis and feature correlation assessment. These reports detail key indicators such as the company's potential liquidity risk, financial health, potential fraud signs, and future growth potential. At the same time, the system also intelligently generates enterprise profile reports, which use visual charts and detailed analysis to showcase the company's financial performance, behavioral characteristics, industry status, and risk profile. These reports not only help enterprise management and risk control departments make more scientific decisions, but also provide regulatory authorities with detailed and reliable enterprise risk data support, effectively improving the overall intelligence and scientific level of enterprise credit risk management.
[0055] The result output module utilizes advanced message transmission and interface technologies, combined with multiple communication protocols such as RESTful API, WebSocket, and message queues, to securely and quickly transmit the enterprise risk warning reports and enterprise profile reports generated through analysis to the backend application terminal. During the implementation process, the system integrates an intelligent report generator, which uses natural language processing (NLP), template engine, and automatic typesetting technology to automatically generate standardized and aesthetically pleasing customized reports based on the user's personalized needs, role permissions, and specific concerns. The generated reports include key risk indicators, enterprise profile overviews, detailed financial analysis, and future warning forecasts, ensuring the integrity and readability of the information. Specifically, the report generator combines user demand parameters (such as focus on financial risks, industry comparisons, and development trends) with pre-defined content templates to recommend and optimize content using the following formula:
[0056] R=f(U,T,C)
[0057] Among them, R represents the final generated report content, U is the user's personalized demand description, T is the template structure parameter, and C is the key indicators and analysis results. The system also implements dynamic content adjustment and real-time update mechanisms to ensure that the report is both comprehensive and concise. Ultimately, these customized and structured enterprise risk warning reports and enterprise portrait reports are securely transmitted to the back-end application terminal for managers, risk control departments and decision-makers to quickly browse, analyze and apply, thereby improving the efficiency and effectiveness of enterprise risk management and realizing intelligent and automated report output and decision-making support.
[0058] Through the comprehensive application between modules, the system integrates data collection, storage, privacy protection, in-depth analysis and intelligent reporting, fully protecting user privacy and security while achieving efficient data circulation and intelligent analysis. The data collection module uses multi-protocol and multi-channel technologies to obtain data from multiple sources such as basic enterprise information, behavioral data, access logs, etc. in real time, and uses encryption, pre-processing and other technologies to transmit it to the storage module. Combined with distributed storage and capacity monitoring, it ensures efficient and secure storage of data. The privacy computing module uses high-end technologies such as homomorphic encryption, secure multi-party computing and differential privacy to generate enterprise labels, portraits and fraud detection indicators, while protecting sensitive information and ensuring that user privacy is not leaked during the data analysis process. The data analysis module combines machine learning models to deeply assess corporate credit risks, output credit scores, and automatically generate corporate risk warning reports and corporate portraits to support corporate management and risk control. The result output module relies on an intelligent report generator to automatically customize the generation of risk warning and corporate portrait reports based on user needs. It relies on templates and natural language processing technology to ensure that the content is comprehensive, intuitive, and in line with user concerns. Ultimately, these analysis results are provided to back-end application terminals through secure interfaces and protocols, providing enterprises with a full-process, intelligent risk assessment and decision-making support solution, achieving a perfect combination of efficient data flow, security and privacy protection, and providing solid technical support for corporate risk management and credit assessment.
[0059] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A privacy-preserving joint analysis system for data circulation and security protection, characterized by: It includes data acquisition module, data storage module, privacy calculation module, data analysis module and result output module; The data collection module is used to collect data in real time from different data sources, including basic enterprise information, enterprise behavior data, user access logs and system operation status, and transmit the real-time collected data to the data storage module; The data storage module is used to store the real-time collected data transmitted by the data acquisition module, and calculate the data storage capacity utilization rate and data query response time to evaluate the usage of storage resources and the efficiency of data retrieval; The privacy computing module is used to perform privacy calculation and analysis on the real-time collected data in the data storage module, using encryption algorithms to protect user privacy, executing enterprise tagging technology to generate enterprise profiles, and performing fraud detection based on the two-way elimination F test method. It also calculates the number of enterprise tags and fraud risk scores to assess the level of detail of enterprise profiles and identify potential fraudulent behavior; The data analysis module is used to conduct in-depth analysis of the results generated by the privacy calculation module, use machine learning models to analyze corporate credit risks, calculate corporate credit scores, and generate corporate risk warning reports and corporate profile reports; The result output module is used to send the enterprise risk warning report and the enterprise portrait report to the background application terminal.
2. The privacy computing joint analysis system with dual data circulation and security protection according to claim 1 is characterized by: The data storage capacity utilization rate calculation formula is as follows: In the formula, Cs represents the data storage capacity utilization rate, C used Indicates the used storage capacity, C total Indicates the total storage capacity.
3. The privacy computing joint analysis system with dual data circulation and security protection according to claim 2 is characterized by: The calculation formula for the data query response time is as follows: In the formula, T avg Indicates the data query response time, N indicates the total number of queries, T i Indicates the response time of the i-th query, where i is the index subscript.
4. The privacy computing joint analysis system with dual data circulation and security protection according to claim 3 is characterized by: The data storage module is equipped with a dynamic storage optimization algorithm for automatically adjusting storage strategies according to real-time data changes, while automatically cleaning out expired and low-value data.
5. The privacy computing joint analysis system with dual data circulation and security protection according to claim 4 is characterized by: The encryption algorithm used to protect user privacy is as follows: C=Enc K (M) In the formula, C represents the encrypted data, M represents the information to be encrypted, K represents the encryption key, and Enc represents the encryption function.
6. The privacy computing joint analysis system with dual data flow and security protection according to claim 5 is characterized by: The algorithm used to generate the enterprise portrait is as follows: Q c =f(X)={l1,l2,...,l k } Among them, Q c represents the enterprise portrait, X represents the enterprise feature vector, f represents the labeling algorithm model, l1, l2, ..., l k Indicates each acquired enterprise label.
7. The privacy computing joint analysis system with dual data circulation and security protection according to claim 6 is characterized by: The calculation formula for the number of enterprise tags is as follows: In the formula, L represents the total number of labels generated by the enterprise, R represents the number of feature categories, and l j,k represents the kth label of the jth class feature, x j,k represents the kth specific eigenvalue in the jth class feature, w j,k Represents the weight coefficient of the kth feature of the jth category, which is obtained based on feature importance training, T j represents the threshold of the j-th class feature, and |*| represents the total number of labels.
8. The privacy computing joint analysis system with dual data circulation and security protection according to claim 7 is characterized by: The calculation formula for the fraud risk score is as follows: In the formula, Fx represents the fraud risk score, x f represents the fraud-related feature vector, represents the model parameters corresponding to the fraud-related feature vector, b f Represents the bias term, σ represents the sigmoid function, and converts the output into a probability in the (0-1) range.
9. The privacy computing joint analysis system with dual data circulation and security protection according to claim 8 is characterized by: The calculation formula for the enterprise credit score is as follows: In the formula, S represents the enterprise credit score, σ represents the sigmoid function, and converts the output into a probability in the range of (0-1). c represents the credit-related feature vector, represents the model parameters corresponding to the credit-related feature vector, b c Indicates the offset.
10. The privacy computing joint analysis system with dual data circulation and security protection according to claim 9 is characterized by: The result output module is equipped with an intelligent report generator for automatically generating customized enterprise risk warning reports and enterprise portrait reports according to user needs.