Differential privacy algorithm optimization method and device for enterprise credit data privacy protection

By adaptively adjusting noise parameters, selectively adding noise, and merging queries, the differential privacy algorithm for enterprise credit data is optimized, solving the accuracy and comparability problems caused by noise introduction and achieving more efficient data privacy protection and analysis.

CN119416260BActive Publication Date: 2025-12-02CHINA CYBER SECURITY REVIEW CERTIFICATION AND MARKET SUPERVISION BIG DATA CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411550369.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-12-02
Estimated Expiration
2044-11-01

AI Technical Summary

Technical Problem

The introduction of noise in existing differential privacy algorithms for protecting corporate credit data leads to problems such as decreased accuracy, reduced signal-to-noise ratio, and poor comparability between different query results.

Method used

By analyzing the query types of enterprise credit data, the noise parameters are adaptively adjusted, noise is selectively added, multiple data queries are merged into one query, and the data after noise processing is preprocessed to optimize the noise mechanism and query process.

Benefits of technology

It significantly improves the accuracy of results while protecting privacy, reduces noise interference, enhances data query and analysis performance, strengthens the comparability of different query results, and supports better business decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416260B_ABST
    Figure CN119416260B_ABST
Patent Text Reader

Abstract

This application provides a differential privacy algorithm optimization method and apparatus for protecting enterprise credit data privacy. The method includes: analyzing the query type of enterprise credit data and adaptively adjusting noise parameters; selectively adding noise based on an improved noise mechanism; merging multiple data queries into one data query; preprocessing the noise-processed enterprise credit data to protect the privacy of the enterprise credit data through a noise-introduced differential privacy algorithm; flexibly adjusting the noise level according to the specific query situation, which can significantly improve the accuracy of the results while protecting privacy; reducing unnecessary noise interference by selectively adding noise, improving the performance of data query and analysis; and merging multiple queries into one query, reducing the frequency of noise introduction, which can improve the comparability between different query results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of differential privacy optimization technology, and more specifically, to a differential privacy algorithm optimization method and apparatus for protecting the privacy of enterprise credit data. Background Technology

[0002] Differential privacy algorithms for corporate credit data privacy protection are cryptographic techniques designed to safeguard corporate privacy by ensuring that individual information is not leaked when publishing statistical data or analyzing corporate credit results. The core idea is to introduce random noise into corporate credit data query results, making it impossible for external observers to accurately identify the participation of a specific individual, thereby protecting corporate credit data privacy.

[0003] The formal definition of differential privacy is based on the "ε-differential privacy" model. This model stipulates that for any two similar datasets (i.e., differing only in one record), the difference in the results for the same query will not exceed a small threshold (ε). Specifically, differential privacy requires that the distribution of query results should be almost identical regardless of whether a particular individual is included in the dataset. This means that even if an attacker knows the query results, they cannot determine whether a particular individual is in the dataset, thus protecting user privacy.

[0004] A key technique for achieving differential privacy is noise addition. Laplace noise or Gaussian noise is typically used, and the amount of noise is determined by calculating the sensitivity of the query results (i.e., the degree of influence of a single record in the enterprise credit dataset on the query results). The higher the sensitivity, the more noise needs to be added to ensure the security of individual data.

[0005] Noise introduction is a core mechanism in traditional differential privacy algorithms for protecting corporate credit data. Its purpose is to protect individual privacy by adding random noise to query results. However, this process inevitably reduces the usability of corporate credit data, specifically in the following ways: First, the addition of noise affects the accuracy of the results. When a large amount of noise is added to the query results, it may mask the true trend of corporate credit data, thus affecting the effectiveness of decision-making. Second, the introduction of noise reduces the signal-to-noise ratio of the data, leading to a decrease in the performance of the analysis model. In machine learning tasks, training models relies on accurate data. If the noise in the training data is too high, the model's learning process may be interfered with, ultimately leading to a decrease in the model's generalization ability in practical applications. Furthermore, the magnitude of noise often needs to be adjusted according to the sensitivity of specific queries, which may reduce the comparability between different query results. When different queries are applied to the same dataset, the inconsistency of noise complicates comparative analysis and reduces the usability of the data. Summary of the Invention

[0006] The purpose of this application is to provide a differential privacy algorithm optimization method and apparatus for protecting enterprise credit data privacy, so as to solve the problems of noise introduced by existing methods when performing differential privacy algorithms, causing unnecessary noise interference, and the comparability and accuracy of different query results.

[0007] In a first aspect, embodiments of this application provide a differential privacy algorithm optimization method for protecting enterprise credit data privacy, including:

[0008] Analyze the query types of enterprise credit data and adaptively adjust noise parameters;

[0009] Based on an improved noise mechanism, noise is selectively increased;

[0010] Merge multiple data queries into one data query;

[0011] The enterprise credit data, after noise processing, is preprocessed to protect its privacy through a differential privacy algorithm that incorporates noise.

[0012] In the above implementation process, the embodiments of this application analyze the query types of enterprise credit data and adaptively adjust the noise parameters; selectively add noise based on the improved noise mechanism; merge multiple data queries into one data query; preprocess the enterprise credit data after noise processing to protect the privacy of enterprise credit data through a differential privacy algorithm introduced by noise; flexibly adjust the noise level according to the specific situation of the query, which can significantly improve the accuracy of the results while protecting privacy; reduce unnecessary noise interference by selectively adding noise, improve the performance of data query and analysis; and merge multiple queries into one query to reduce the frequency of noise introduction, which can improve the comparability between different query results.

[0013] Furthermore, the query type for analyzing enterprise credit data adaptively adjusts noise parameters, including:

[0014] Analyze the query types of corporate credit data and assess the sensitivity of corporate credit data;

[0015] Identify the factors influencing the noise parameters, including the allocated privacy budget, the sensitivity measure of the query, and the distribution characteristics of the data;

[0016] Set noise parameters, dynamically adjust noise parameters, and adjust noise parameters according to the application scenario;

[0017] Verify the noise parameter settings and continuously optimize them.

[0018] In the above implementation process, by flexibly adjusting the noise level according to the specific circumstances of the query, the accuracy of the results can be significantly improved while protecting privacy.

[0019] Furthermore, the selective increase of noise based on the improved noise mechanism includes:

[0020] Classify and standardize enterprise credit data;

[0021] Select the desired noise type to add, and specify the noise parameters;

[0022] The magnitude of noise is determined based on data sensitivity and query type;

[0023] Generate random noise and add it to the data to update the data;

[0024] The data after adding noise was validated and evaluated.

[0025] In the above implementation process, selective noise addition reduces unnecessary noise interference and improves the performance of data query and analysis.

[0026] Furthermore, the merging of multiple data queries into one data query includes:

[0027] Analyze query requirements, determine query relevance, and merge multiple data queries into one data query;

[0028] Choose a merge strategy to optimize the query statement;

[0029] Modify the query interface, process query results, and monitor query performance.

[0030] In the above implementation process, merging multiple queries into one query reduces the frequency of noise introduction, improves the comparability between different query results, and allows for more accurate cross-query analysis, thereby helping enterprises make better decisions.

[0031] Furthermore, the preprocessing of the noise-processed enterprise credit data includes:

[0032] Identify the source of enterprise credit data and conduct a preliminary review of the enterprise credit data;

[0033] Detect outliers in enterprise credit data;

[0034] Remove noise from corporate credit data.

[0035] In the above implementation process, data cleaning and preprocessing remove outliers and noise, improve the quality of the original data, directly affect the accuracy and effectiveness of the final analysis results, and thus improve support for business decisions.

[0036] Furthermore, the analysis of query types for enterprise credit data and the assessment of the sensitivity of enterprise credit data include:

[0037] Identify the types of queries present in enterprise credit data, where different query types have different data sensitivity and privacy requirements;

[0038] The sensitivity of corporate credit data is assessed, including an evaluation of the company's financial condition, default history, and credit rating.

[0039] The factors influencing the determination of noise parameters include:

[0040] Allocate a privacy budget to control the extent of added noise;

[0041] Different privacy budgets are allocated based on the type of query and data sensitivity.

[0042] Establish a metric for query sensitivity and determine the sensitivity of a query by setting a method. The higher the sensitivity of a query, the larger the noise parameter should be added.

[0043] Establish the distribution characteristics of enterprise credit data and analyze the data distribution using statistical methods;

[0044] The step of setting noise parameters, dynamically adjusting noise parameters, and adjusting noise parameters according to the application scenario includes:

[0045] Different noise parameter setting strategies should be developed for different types of queries;

[0046] Specifically, for range queries, the noise parameter is determined based on the size of the query range and the distribution of the data; for count queries, the noise parameter is set based on the total number of companies in the dataset and the precision requirements of the query; and for summation queries, the noise parameter is set based on the numerical range of the data and the importance of the query.

[0047] Different noise parameters are set according to the sensitivity level of the data, where the sensitivity levels include high sensitivity level, medium sensitivity level and low sensitivity level;

[0048] Among them, the noise parameter set for the high sensitivity level is greater than that set for the medium sensitivity level, and the noise parameter set for the medium sensitivity level is greater than that set for the low sensitivity level.

[0049] Regularly re-evaluate the data, analyze changes in query types and data sensitivity, and adjust noise parameters based on the evaluation results;

[0050] Adjust noise parameters according to the application scenario;

[0051] The process of verifying and continuously optimizing the noise parameters includes:

[0052] Select a set number of enterprise credit data as samples, simulate different types of queries, and add noise according to the set noise parameters;

[0053] Compare the results of query with added noise with the actual results to evaluate the impact of noise parameters on the accuracy of query results;

[0054] If the difference exceeds the set threshold, adjust the noise parameters;

[0055] Continuously optimize noise parameter settings;

[0056] Establish a feedback mechanism to collect user feedback on the accuracy of query results and the degree of privacy protection;

[0057] Based on feedback and actual application, the noise parameter setting strategy is continuously adjusted.

[0058] Furthermore, the classification and standardization of enterprise credit data includes:

[0059] Classify enterprise credit data by categorizing it into different types based on data sensitivity and query type;

[0060] Standardize the categorized data;

[0061] The step of selecting and adding a set noise type and determining noise parameters includes:

[0062] Select the desired noise type to add;

[0063] The noise parameters are determined based on the sensitivity of the data and the type of query; these parameters include the mean, variance, and standard deviation of the noise.

[0064] The step of generating random noise and adding it to the data to update the data includes:

[0065] Use a random number generator to generate random noise, and add the generated random noise to the data;

[0066] Update the database with the data after adding random noise, replacing the original data;

[0067] The verification and evaluation of the data after adding noise includes:

[0068] Use data validation tools to validate the data after adding random noise;

[0069] Perform a privacy assessment on the data after adding random noise;

[0070] Performance evaluation was performed on the data after adding random noise.

[0071] Furthermore, the analysis of query requirements, determination of query relevance, and merging of multiple data queries into a single data query include:

[0072] Identify the use cases and business needs for obtaining enterprise credit data, and determine the relevance of the query.

[0073] Analyze the relevance between different queries;

[0074] If multiple queries involve the same data fields or the similarity of the query conditions exceeds a set threshold, they will be merged into one query.

[0075] The selection of a merging strategy to optimize the query statement includes:

[0076] Based on the relevance of the query and business requirements, select a merging strategy, which includes: condition merging, result merging, and function merging.

[0077] When designing merge queries, optimize the query statement. Methods for optimizing the query statement include avoiding duplicate queries, using indexes, and limiting the query result set.

[0078] The modification of the query interface, processing of query results, and monitoring of query performance include:

[0079] Based on the designed merge query, modify the data query interface so that the query interface can accept the merged query request;

[0080] After executing the merge query, the query results are processed.

[0081] If a result merging strategy is used, the results of multiple queries will be merged and organized.

[0082] If a function merging strategy is used, the corresponding function is called to process it;

[0083] After implementing the merge query, use database performance monitoring tools or log analysis tools to monitor query performance, which includes query execution time, CPU utilization, and memory utilization.

[0084] Furthermore, the process of clarifying the source of enterprise credit data and conducting a preliminary review of the enterprise credit data includes:

[0085] Determine the source channels of corporate credit data and understand the data collection methods and timeframes;

[0086] Conduct a preliminary review of the collected data to understand its format, field meanings, and data types;

[0087] The outlier detection of enterprise credit data includes:

[0088] Calculate the basic statistics of the data to understand the distribution of the data and identify outliers;

[0089] Use visualization tools to show the distribution of the data;

[0090] Outlier detection is achieved using model-based methods, which include cluster analysis and outlier detection algorithms.

[0091] The removal of noise from enterprise credit data includes:

[0092] Filtering methods are used to remove noise from enterprise credit data. These methods include mean filtering, median filtering, and Gaussian filtering.

[0093] If the data contains noise that is correlated with other variables, use regression analysis to remove the noise.

[0094] For time series data or data with continuous changing trends, data smoothing methods are used to remove noise. These methods include moving average and exponential smoothing.

[0095] Secondly, embodiments of this application provide a differential privacy algorithm optimization device for protecting enterprise credit data privacy, comprising:

[0096] The query analysis module is used to analyze the query types of enterprise credit data and adaptively adjust noise parameters.

[0097] A noise conditioning module for selectively increasing noise based on an improved noise mechanism;

[0098] The query merging module is used to merge multiple data queries into a single data query;

[0099] The data processing module is used to preprocess the noise-processed corporate credit data in order to protect the privacy of the corporate credit data through a differential privacy algorithm that introduces noise.

[0100] Thirdly, embodiments of this application provide an electronic device, including:

[0101] The system includes a processor, a memory, and a bus. The processor is connected to the memory via the bus. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, they are used to implement the differential privacy algorithm optimization method for protecting enterprise credit data privacy as described above.

[0102] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a server, implements the differential privacy algorithm optimization method for protecting enterprise credit data privacy as described above. Attached Figure Description

[0103] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0104] Figure 1 A flowchart illustrating a differential privacy algorithm optimization method for protecting enterprise credit data privacy provided in this application embodiment;

[0105] Figure 2 This is a schematic diagram of the structure of a differential privacy algorithm optimization device for protecting enterprise credit data privacy provided in an embodiment of this application;

[0106] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0107] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0108] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0109] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a differential privacy algorithm optimization method for protecting enterprise credit data privacy, provided in an embodiment of this application. The differential privacy algorithm optimization method for protecting enterprise credit data privacy includes:

[0110] 100. Analyze the query types of enterprise credit data and adaptively adjust noise parameters.

[0111] Specifically, the noise level is dynamically adjusted based on the specific data query. For example, noise parameters can be flexibly set for different types of queries or data sensitivity to balance privacy protection and data accuracy.

[0112] 110. Analyze the query types of enterprise credit data and assess the sensitivity of enterprise credit data.

[0113] Specifically, it involves identifying the types of queries present in enterprise credit data, as different query types have varying degrees of sensitivity and privacy requirements.

[0114] Optionally, identify the types of queries that may exist in the enterprise credit data, such as range queries (querying the number of enterprises within a specific credit score range), count queries (the total number of creditworthy enterprises in a specific industry), and summation queries (the total credit score of enterprises in a certain region). Different query types may have different levels of data sensitivity and privacy requirements. For example, range queries may more easily expose the distribution characteristics of the data, while count queries may have relatively lower data sensitivity.

[0115] Specifically, the sensitivity of corporate credit data is assessed, including an evaluation of the company's financial condition, default history, and credit rating.

[0116] Optionally, the sensitivity of corporate credit data can be assessed, considering specific information contained within the data, such as the company's financial condition, default history, and credit rating. Generally, data involving core financial data and sensitive information such as defaults are more sensitive, while some relatively broad credit rating information may be less sensitive.

[0117] 120. Identify the factors influencing the noise parameters, including the allocated privacy budget, the sensitivity measure of the query, and the distribution characteristics of the data.

[0118] 121. Allocate a privacy budget to control the extent of added noise.

[0119] Understandably, differential privacy protection typically allocates a privacy budget to control the degree of noise added. This privacy budget needs to be allocated appropriately when dealing with different types of queries and data sensitivities.

[0120] 122. Allocate different shares of the privacy budget for different types of queries and data sensitivity.

[0121] For example, for highly sensitive data and queries, a smaller privacy budget share can be allocated to increase the amount of noise and improve the level of privacy protection; for low-sensitivity data and queries, a larger privacy budget share can be allocated to reduce the impact of noise on the accuracy of query results.

[0122] 123. Establish a metric for query sensitivity. Determine the sensitivity of a query by setting a method. The higher the sensitivity of a query, the greater the noise parameter should be added.

[0123] Optionally, a metric for query sensitivity can be established. Query sensitivity is typically related to the degree of change in the query results. For example, for a counting query, if a change in a record in the dataset has a small impact on the query results, then the query sensitivity is low.

[0124] For example, the sensitivity of a query can be determined by analyzing its mathematical properties or by conducting experiments; the more sensitive the query, the larger the noise parameter needs to be added.

[0125] 124. Establish the distribution characteristics of enterprise credit data and use statistical methods to analyze the data distribution.

[0126] Specifically, consider the distribution characteristics of corporate credit data. Optionally, if the data is relatively concentrated, then a small amount of noise may be sufficient to provide privacy protection; if the data is relatively dispersed, a larger amount of noise may be needed to hide the specific characteristics of the data.

[0127] Statistical methods can be used to analyze the distribution of data, such as mean, variance, and standard deviation, to understand the degree of centralization and dispersion of the data.

[0128] 130. Set noise parameters, dynamically adjust noise parameters, and adjust noise parameters according to application scenarios.

[0129] Specifically, different noise parameter setting strategies should be developed for different types of queries.

[0130] For example, for range queries, the noise parameter is determined based on the size of the query range and the distribution of the data. If the query range is small and the data distribution is relatively concentrated, the noise parameter can be appropriately reduced; if the query range is large or the data distribution is relatively scattered, the noise parameter needs to be increased.

[0131] For example, for a counting query, the noise parameter is set according to the total number of companies in the dataset and the accuracy requirement of the query. If the total number of companies is large and the accuracy requirement of the query is not high, a smaller noise parameter can be used; if the total number of companies is small or the accuracy requirement of the query is high, the noise parameter needs to be increased.

[0132] For example, for a summation query, the noise parameter can be set according to the numerical range of the data and the importance of the query. If the numerical range of the data is large and the query result has a significant impact on the decision, the noise parameter can be increased to improve the level of privacy protection; if the numerical range of the data is small or the query result has a small impact on the decision, the noise parameter can be appropriately reduced.

[0133] Specifically, different noise parameters are set according to the sensitivity level of the data, which includes high sensitivity level, medium sensitivity level and low sensitivity level.

[0134] Among them, the noise parameter set for the high sensitivity level is greater than that set for the medium sensitivity level, and the noise parameter set for the medium sensitivity level is greater than that set for the low sensitivity level.

[0135] For example, for highly sensitive data, such as a company's core financial data and default records, a larger noise parameter can be set to ensure that the data privacy is fully protected.

[0136] For example, for moderately sensitive data, such as a company's credit rating and some financial indicators, a moderate noise parameter can be set to minimize the impact on the accuracy of the query results while protecting privacy.

[0137] For example, for low-sensitivity data, such as basic information about a company and some publicly available credit data, a smaller noise parameter can be set to improve the accuracy of the query results.

[0138] Specifically, the data is periodically re-evaluated, and changes in query types and data sensitivity are analyzed. Noise parameters are then adjusted based on the evaluation results.

[0139] Understandably, query types and data sensitivity may change over time and with data variations. Therefore, a mechanism for dynamically adjusting noise parameters is needed.

[0140] For example, data can be re-evaluated periodically to analyze changes in query types and data sensitivity, and noise parameters can be adjusted based on the evaluation results. For instance, data can be analyzed weekly or monthly, and noise parameter settings can be adjusted based on the analysis results.

[0141] Specifically, the noise parameters are adjusted according to the application scenario.

[0142] Optionally, consider the actual application scenarios of corporate credit data, such as credit assessment, risk analysis, and regulatory reporting. Different application scenarios have different requirements for the accuracy and privacy protection of the data.

[0143] When setting noise parameters, the requirements of the application scenario need to be considered comprehensively. For example, when conducting credit assessments, a higher accuracy of query results may be required, so the noise parameter can be appropriately reduced; while when submitting reports to regulatory agencies, greater emphasis may be placed on data privacy protection, in which case the noise parameter can be increased.

[0144] 140. Verify the noise parameter settings and continuously optimize the noise parameters.

[0145] Specifically, a set number of enterprise credit data are selected as samples to simulate different types of queries, and noise is added according to the set noise parameters; the difference between the query results after adding noise and the actual results is compared to evaluate the impact of noise parameters on the accuracy of query results; if the difference exceeds the set threshold, the noise parameters are adjusted until a reasonable balance is reached.

[0146] Specifically, we will continuously optimize the setting of noise parameters; establish a feedback mechanism to collect user feedback on the accuracy and privacy protection of query results; and continuously adjust the setting strategy of noise parameters based on feedback and actual application to improve the effectiveness of differential privacy algorithms in protecting corporate credit data privacy.

[0147] Understandably, flexibly setting noise parameters for different types of queries or data sensitivity requires comprehensive consideration of factors such as query type, data sensitivity, privacy budget allocation, query sensitivity measurement, and data distribution characteristics, and adjustments and optimizations should be made in conjunction with actual application scenarios. By setting noise parameters reasonably, the impact on the accuracy of query results can be minimized while protecting the privacy of corporate credit data.

[0148] As described above, the embodiments of this application can significantly improve the accuracy of results while protecting privacy by flexibly adjusting the noise level according to the specific circumstances of the query.

[0149] 200. Based on an improved noise mechanism, noise is selectively increased.

[0150] Specifically, reducing noise in some insensitive or low-sensitivity data queries, while increasing noise in high-sensitivity queries, can improve data usability while maintaining privacy.

[0151] 210. Classify and standardize enterprise credit data.

[0152] Optionally, enterprise credit data can be categorized based on its sensitivity and query type. For example, data can be classified as highly sensitive, moderately sensitive, and low sensitive, or as count query data, range query data, and summation query data.

[0153] Optionally, the categorized data can be standardized to ensure that different types of data have the same scale and distribution. For example, this can be achieved through data normalization or standardization, where standardization can improve the effectiveness of noise addition and data usability.

[0154] 220. Select the desired noise type to add and determine the noise parameters.

[0155] Optionally, you can select a noise type to add. For example, common noise types include Gaussian noise and Laplace noise. Gaussian noise is suitable for continuous data, while Laplace noise is suitable for discrete data. Choose the appropriate noise type according to the type and characteristics of the data.

[0156] Optionally, the noise parameters can be determined based on the data sensitivity and query type; these parameters include the mean, variance, and standard deviation of the noise. The selection of these noise parameters should be based on the data sensitivity and query type. For example, a smaller noise parameter can be selected for highly sensitive data and high-precision queries, while a larger noise parameter can be selected for low-sensitivity data and low-precision queries.

[0157] 230. Determine the magnitude of noise based on data sensitivity and query type.

[0158] Optionally, a data sensitivity-based strategy can be used to add more noise to highly sensitive data in order to improve the level of data privacy protection. The size of the noise can be determined according to the sensitivity level of the data. For example, for the most sensitive data, noise with a standard deviation equal to a certain percentage of the data range can be added.

[0159] For moderately sensitive data, add appropriate noise to minimize the impact on data availability while protecting privacy. The noise level can be determined based on the data sensitivity and query type. For example, for moderately sensitive data, noise with a standard deviation equal to a small percentage of the data range can be added.

[0160] For low-sensitivity data, adding small noise or no noise can improve data availability. The decision to add noise can be made based on the sensitivity of the data and the type of query. For example, for the least sensitive data, no noise or very little noise can be added.

[0161] Optionally, based on the query type, a smaller noise can be added for count queries, since count queries are relatively insensitive to changes in the data. The size of the noise can be determined according to the accuracy requirements of the query and the total amount of data. For example, for count queries, noise with a standard deviation equal to a certain percentage of the total amount of data can be added.

[0162] For range queries, you can add moderate noise because range queries are sensitive to the distribution of data. The amount of noise can be determined based on the range of the query and the distribution of the data. For example, for range queries, you can add noise with a standard deviation that is a certain percentage of the query range.

[0163] For summation queries, larger noise can be added because summation queries are sensitive to changes in the data. The size of the noise can be determined based on the accuracy requirements of the query and the total amount of data. For example, for summation queries, noise with a standard deviation that is a large proportion of the total data can be added.

[0164] 240. Generate random noise and add it to the data, then update the data.

[0165] Optionally, a random number generator can be used to generate random noise, which can then be added to the data. A pseudo-random number generator or a hardware random number generator can be used to ensure the randomness and unpredictability of the noise.

[0166] Optionally, the data with added random noise can be updated in the database, replacing the original data. When updating the data, it is necessary to ensure data integrity and consistency to avoid data loss or errors.

[0167] 250. Verify and evaluate the data after adding noise.

[0168] Optionally, data validation tools can be used to validate the data after adding random noise. Validating the data with added noise ensures its validity and usability; this can be done using data validation tools or by manually checking the accuracy and completeness of the data.

[0169] Optionally, a privacy assessment can be performed on the data after adding random noise to ensure that data privacy is effectively protected. This can be done using privacy assessment tools or by conducting simulated attacks to evaluate the degree of privacy protection.

[0170] Optionally, a performance evaluation can be performed on the data after adding random noise to assess the impact of noise addition on data query performance and data availability. This evaluation can be conducted using performance testing tools or real-world application scenarios.

[0171] As described above, the embodiments of this application reduce unnecessary noise interference and improve the performance of data query and analysis by selectively adding noise.

[0172] 300. Merge multiple data queries into one data query.

[0173] Specifically, by using a joint analysis mechanism, multiple queries are merged into one query, reducing the frequency of noise introduction. As a result, the overall noise level can be reduced, thereby improving the accuracy of each individual query.

[0174] 310. Analyze query requirements, determine query relevance, and merge multiple data queries into one data query.

[0175] Optionally, it's necessary to understand the use cases and business needs of the enterprise credit data to determine query relevance. Specifically, this involves gaining a deep understanding of the use cases and business needs of the enterprise credit data to identify which queries are frequently used together and the relationships between them. For example, in credit assessment, it might be necessary to simultaneously query information such as the enterprise's financial status, credit rating, and default records.

[0176] Optionally, analyze the relevance between different queries; if multiple queries involve the same data fields or the similarity of query conditions exceeds a set threshold, merge them into one query. For example, querying a company's sales revenue within a specific time period and querying its profits within the same time period may be merged into a single query that queries the company's financial situation within that time period.

[0177] 320. Select a merge strategy to optimize the query statement.

[0178] 321. Based on the relevance of the query and business requirements, select a merging strategy, which includes: condition merging, result merging, and function merging.

[0179] For example, condition merging: combining the conditions of multiple queries into a single complex query condition. For instance, merging queries for companies with a credit rating of A and queries for companies with sales exceeding 1 million into a query for companies with both a credit rating of A and sales exceeding 1 million.

[0180] For example, result merging involves executing multiple queries separately and then combining the results. For instance, querying a company's financial condition and its credit rating yields financial statements and a credit rating report respectively; these are then combined into a single, comprehensive company information report.

[0181] For example, function merging: The results of multiple queries are taken as input, processed by a single function, and the final query result is obtained. For instance, querying a company's sales revenue and costs, and then calculating the company's profit using a profit function.

[0182] 322. When designing merge queries, optimize the query statements. Methods for optimizing query statements include avoiding duplicate queries, using indexes, and limiting the query result set.

[0183] For example, avoid duplicate queries: If multiple queries involve the same data fields, try to avoid querying those fields repeatedly. For instance, you can use subqueries or join queries to reduce the number of times data is read.

[0184] For example, using indexes: Creating indexes for frequently queried fields can speed up queries and reduce query time and noise.

[0185] For example, limiting the query result set: If you only need to query a portion of the data, you can use the LIMIT statement to limit the size of the query result set, reducing data reading and processing time.

[0186] 330. Modify the query interface, process the query results, and monitor query performance.

[0187] 331. Based on the designed merge query, modify the data query interface so that the query interface can accept the merged query request.

[0188] For example, if using a database query language, you may need to modify the SQL statement or stored procedure; if using an API to perform queries, you may need to modify the API parameters and return values.

[0189] 332. After executing the merge query, process the query results; if a result merging strategy is used, merge and organize multiple query results; if a function merging strategy is used, call the corresponding function for processing.

[0190] 333. After implementing the merge query, use database performance monitoring tools or log analysis tools to monitor query performance, including query execution time, CPU utilization, and memory utilization.

[0191] Specifically, after implementing a merge query, it's necessary to monitor query performance to ensure it doesn't significantly impact system performance. Database performance monitoring tools or log analysis tools can be used to monitor metrics such as query execution time, CPU usage, and memory usage. If a performance degradation is detected, the query statement can be further optimized or the merge strategy adjusted.

[0192] As described above, the embodiments of this application merge multiple queries into one query, reducing the frequency of noise introduction, improving the comparability between different query results, and allowing for more accurate cross-query analysis, thereby helping enterprises make better decisions.

[0193] 400. Preprocess the enterprise credit data after noise processing to protect the privacy of the enterprise credit data through a differential privacy algorithm that introduces noise.

[0194] 410. Clarify the source of enterprise credit data and conduct a preliminary review of enterprise credit data.

[0195] 411. Determine the source channels of enterprise credit data and understand the data collection methods and time range.

[0196] Specifically, identify the sources of corporate credit data, such as internal databases, external data providers, and surveys. Understand the methods and timeframes of data collection to better comprehend the characteristics of the data and potential issues.

[0197] 412. Conduct a preliminary review of the collected data to understand its format, field meanings, and data types.

[0198] Specifically, a preliminary review is conducted on the collected data to understand its format, field meanings, data types, etc., and to check whether the data is complete and whether there are any missing or duplicate values.

[0199] 420. Perform outlier detection on enterprise credit data.

[0200] 421. Calculate the basic statistics of the data. Through these basic statistics, we can understand the distribution of the data and identify any outliers.

[0201] Specifically, calculating basic statistics for the data, such as the mean, median, standard deviation, and quartiles, allows us to gain a preliminary understanding of the data distribution and identify potential outliers. For example, if a data point deviates from the mean by more than a certain multiple of the standard deviation, or is located at an extreme point in the data distribution, it may be considered an outlier.

[0202] 422. Use visualization tools to show the distribution of data.

[0203] Specifically, visualization tools such as histograms, box plots, and scatter plots can be used to visually display the distribution of data. Through visualization analysis, outliers and patterns in the data can be more easily identified. For example, box plots can display the median, quartiles, and outlier ranges of the data, thus enabling the rapid identification of potential outliers.

[0204] 423. Use model-based methods to detect outliers, including cluster analysis and outlier detection algorithms.

[0205] Specifically, model-based methods can be used to detect outliers, such as cluster analysis and outlier detection algorithms. These methods identify outliers that differ from the majority of data points by analyzing the inherent structure and patterns of the data. For example, cluster analysis can divide data points into different groups, and outlier detection algorithms can identify outliers that are far from other data points.

[0206] 430. Remove noise from enterprise credit data.

[0207] 431. Use filtering methods to remove noise from enterprise credit data, including mean filtering, median filtering, and Gaussian filtering.

[0208] Specifically, filtering methods can be used to remove noise from data. Common filtering methods include mean filtering, median filtering, and Gaussian filtering. Mean filtering replaces the current data point with the average of the surrounding data points, median filtering replaces the current data point with the median of the surrounding data points, and Gaussian filtering calculates a weighted average of the surrounding data points based on a Gaussian distribution.

[0209] 432. If the data contains noise that is correlated with other variables, use regression analysis to remove the noise;

[0210] Specifically, if the data contains noise that is correlated with other variables, regression analysis can be used to remove the noise. By establishing a regression model between the data and other relevant variables, the true value of the data can be predicted, thereby eliminating the impact of noise. For example, if a certain indicator in corporate credit data is correlated with factors such as the company's size and industry, a regression model can be established to predict the true value of that indicator, removing noise that is unrelated to these factors.

[0211] 433. For time series data or data with continuous changing trends, use data smoothing methods to remove noise. Data smoothing methods include moving average and exponential smoothing.

[0212] Specifically, for time series data or data with continuous trends, data smoothing methods can be used to remove noise. Common data smoothing methods include moving average and exponential smoothing. Moving average uses the average value of data over a certain period to replace the current data point, while exponential smoothing uses a weighted average of historical data, giving more weight to recent data.

[0213] As described above, the embodiments of this application remove outliers and noise through data cleaning and preprocessing, thereby improving the quality of the original data and directly affecting the accuracy and effectiveness of the final analysis results, thus improving support for business decisions.

[0214] Therefore, this application embodiment analyzes the query types of enterprise credit data and adaptively adjusts noise parameters; selectively adds noise based on an improved noise mechanism; merges multiple data queries into one data query; preprocesses the noise-processed enterprise credit data to protect the privacy of enterprise credit data through a differential privacy algorithm introduced by noise; and flexibly adjusts the noise level according to the specific circumstances of the query, which can significantly improve the accuracy of the results while protecting privacy. By selectively adding noise, unnecessary noise interference is reduced, improving the performance of data query and analysis. At the same time, merging multiple queries into one query reduces the frequency of noise introduction and improves the comparability between different query results.

[0215] This application's embodiments improve data accuracy: By dynamically adjusting noise levels according to the specific query, the accuracy of results can be significantly improved while protecting privacy. For example, noise can be increased for high-sensitivity queries and decreased for low-sensitivity queries, enabling users to obtain more reliable data results.

[0216] Performance optimization involves selectively adding noise through an improved noise mechanism to reduce unnecessary noise interference and improve the performance of data querying and analysis. This significantly improves the speed and efficiency of data processing, especially in scenarios requiring real-time analysis.

[0217] Enhanced comparability: By combining multiple queries into a single query through joint analysis, the frequency of noise introduction is reduced, which can improve the comparability between different query results and allow for more accurate cross-query analysis, thereby helping enterprises make better decisions.

[0218] Improving data quality, through data cleaning and preprocessing to remove outliers and noise, enhances the quality of raw data, directly impacting the accuracy and effectiveness of the final analysis results, thereby improving support for business decisions.

[0219] The steps described above are not strictly performed in the order of their numbers; they should be understood as a whole.

[0220] Secondly, based on the above embodiments, Figure 2 This is a schematic diagram of a differential privacy algorithm optimization device for protecting enterprise credit data privacy, provided as an embodiment of this application. (Reference) Figure 2 The differential privacy algorithm optimization device for protecting enterprise credit data privacy provided in this embodiment specifically includes: a query analysis module 201, a noise adjustment module 202, a query merging module 203, and a data processing module 204.

[0221] The query analysis module 201 is used to analyze the query types of enterprise credit data and adaptively adjust the noise parameters; the noise adjustment module 202 is used to selectively increase noise based on the improved noise mechanism; the query merging module 203 is used to merge multiple data queries into one data query; and the data processing module 204 is used to preprocess the noise-processed enterprise credit data to protect the privacy of enterprise credit data through a differential privacy algorithm introduced by noise.

[0222] As described above, the embodiments of this application analyze the query types of enterprise credit data and adaptively adjust noise parameters; selectively increase noise based on an improved noise mechanism; merge multiple data queries into one data query; preprocess the noise-processed enterprise credit data to protect the privacy of enterprise credit data through a differential privacy algorithm introduced by noise; flexibly adjust the noise level according to the specific situation of the query, which can significantly improve the accuracy of the results while protecting privacy; reduce unnecessary noise interference by selectively adding noise, improve the performance of data query and analysis; and merge multiple queries into one query to reduce the frequency of noise introduction, thereby improving the comparability between different query results.

[0223] The differential privacy algorithm optimization device for enterprise credit data privacy protection provided in this application embodiment can be used to execute the differential privacy algorithm optimization method for enterprise credit data privacy protection provided in the above embodiment, and has corresponding functions and beneficial effects.

[0224] Thirdly, embodiments of this application also provide an electronic device that can integrate the differential privacy algorithm optimization device for enterprise credit data privacy protection provided in embodiments of this application. Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. (Reference) Figure 3 The electronic device includes: an input device 33, an output device 34, a memory 32, and one or more processors 31; the memory 32 is used to store one or more programs; when the one or more programs are executed by the one or more processors 31, the one or more processors 31 implement the differential privacy algorithm optimization method for protecting enterprise credit data privacy as provided in the above embodiments. The input device 33, output device 34, memory 32, and processors 31 can be connected via a bus or other means. Figure 3 Taking the example of a connection between China and Israel via a bus.

[0225] The processor 31 executes various functional applications and data processing of the device by running software programs, instructions and modules stored in the memory 32, thereby realizing the differential privacy algorithm optimization method for protecting enterprise credit data privacy as described above.

[0226] The electronic device provided above can be used to execute the differential privacy algorithm optimization method for protecting enterprise credit data privacy provided in the above embodiments, and has corresponding functions and beneficial effects.

[0227] Fourthly, embodiments of this application also provide a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, it controls the device where the computer-readable storage medium is located to execute the differential privacy algorithm optimization method for protecting enterprise credit data privacy as described above, and can achieve the same beneficial effects.

[0228] Of course, the computer-executable instructions provided in the embodiments of this application are not limited to the differential privacy algorithm optimization method for protecting enterprise credit data privacy as described above, but can also execute related operations in the differential privacy algorithm optimization method for protecting enterprise credit data privacy provided in any embodiment of this application.

[0229] Fifthly, embodiments of this application also provide a computer program product. The methods described in the various embodiments of this application can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the various embodiments of this application are executed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, network equipment, user equipment, core network equipment, OAM (Open Application Model), or other programmable devices.

[0230] The computer program or instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions may be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium may be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; or an optical medium, such as a digital video optical disc; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.

[0231] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0232] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0233] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0234] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0235] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0236] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

Claims

1. A differential privacy algorithm optimization method for protecting enterprise credit data privacy, characterized in that, The method includes: Analyze the query types of enterprise credit data and adaptively adjust noise parameters; Based on an improved noise mechanism, noise is selectively increased; Merge multiple data queries into one data query; this involves analyzing query requirements, determining query relevance, merging multiple data queries into one data query, selecting a merging strategy, optimizing the query statement, modifying the query interface, processing the query results, and monitoring query performance. Preprocess the noise-processed corporate credit data to protect the privacy of the corporate credit data through a differential privacy algorithm that introduces noise. The modification of the query interface, processing of query results, and monitoring of query performance include: Based on the designed merge query, modify the data query interface so that the query interface can accept the merged query request; After executing the merge query, the query results are processed. If a result merging strategy is used, the results of multiple queries will be merged and organized. If a function merging strategy is used, the corresponding function is called to process it; After implementing the merge query, use database performance monitoring tools or log analysis tools to monitor query performance, including query execution time, CPU utilization, and memory utilization. If a decrease in query performance is found, further optimize the query statement or adjust the merge strategy. Methods to optimize query statements include avoiding duplicate queries, using indexes, and limiting the query result set.

2. The differential privacy algorithm optimization method for protecting enterprise credit data privacy according to claim 1, characterized in that, The query types for analyzing enterprise credit data adaptively adjust noise parameters, including: Analyze the query types of corporate credit data and assess the sensitivity of corporate credit data; Identify the factors influencing the noise parameters, including the allocated privacy budget, the sensitivity measure of the query, and the distribution characteristics of the data; Set noise parameters, dynamically adjust noise parameters, and adjust noise parameters according to application scenarios; Verify the noise parameter settings and continuously optimize them.

3. The differential privacy algorithm optimization method for protecting enterprise credit data privacy according to claim 1, characterized in that, The improved noise mechanism selectively increases noise, including: Classify and standardize enterprise credit data; Select the desired noise type to add, and specify the noise parameters; The magnitude of noise is determined based on data sensitivity and query type; Generate random noise and add it to the data to update the data; The data after adding noise was validated and evaluated.

4. The differential privacy algorithm optimization method for protecting enterprise credit data privacy according to claim 1, characterized in that, The preprocessing of the noise-processed enterprise credit data includes: Identify the source of enterprise credit data and conduct a preliminary review of the enterprise credit data; Perform outlier detection on enterprise credit data; Remove noise from corporate credit data.

5. The differential privacy algorithm optimization method for protecting enterprise credit data privacy according to claim 2, characterized in that, The analysis of enterprise credit data query types and the assessment of the sensitivity of enterprise credit data include: Identify the types of queries present in enterprise credit data, where different query types have different data sensitivity and privacy requirements; The sensitivity of corporate credit data is assessed, including an evaluation of the company's financial condition, default history, and credit rating. The factors influencing the determination of noise parameters include: Allocate a privacy budget to control the extent of added noise; Different privacy budgets are allocated based on the type of query and data sensitivity. Establish a metric for query sensitivity and determine the sensitivity of a query by setting a method. The higher the sensitivity of a query, the larger the noise parameter should be added. Establish the distribution characteristics of enterprise credit data and analyze the data distribution using statistical methods; The step of setting noise parameters, dynamically adjusting noise parameters, and adjusting noise parameters according to the application scenario includes: Different noise parameter setting strategies should be developed for different types of queries; Specifically, for range queries, the noise parameter is determined based on the size of the query range and the distribution of the data; for count queries, the noise parameter is set based on the total number of companies in the dataset and the precision requirements of the query; and for summation queries, the noise parameter is set based on the numerical range of the data and the importance of the query. Different noise parameters are set according to the sensitivity level of the data, where the sensitivity levels include high sensitivity level, medium sensitivity level and low sensitivity level; Among them, the noise parameter set for the high sensitivity level is greater than that set for the medium sensitivity level, and the noise parameter set for the medium sensitivity level is greater than that set for the low sensitivity level. Regularly re-evaluate the data, analyze changes in query types and data sensitivity, and adjust noise parameters based on the evaluation results; Adjust noise parameters according to the application scenario; The process of verifying and continuously optimizing the noise parameters includes: Select a set number of enterprise credit data as samples, simulate different types of queries, and add noise according to the set noise parameters; Compare the results of the query with added noise to the actual results, and evaluate the impact of noise parameters on the accuracy of the query results. If the difference exceeds the set threshold, adjust the noise parameters; Continuously optimize noise parameter settings; Establish a feedback mechanism to collect user feedback on the accuracy of query results and the degree of privacy protection; Based on feedback and actual application, the noise parameter setting strategy is continuously adjusted.

6. The differential privacy algorithm optimization method for protecting enterprise credit data privacy according to claim 3, characterized in that, The classification and standardization of enterprise credit data includes: Classify enterprise credit data by categorizing it into different types based on data sensitivity and query type; Standardize the categorized data; The step of selecting and adding a set noise type and determining noise parameters includes: Select the desired noise type to add; The noise parameters are determined based on the sensitivity of the data and the type of query; these parameters include the mean, variance, and standard deviation of the noise. The step of generating random noise and adding it to the data to update the data includes: Use a random number generator to generate random noise, and add the generated random noise to the data; Update the database with the data after adding random noise, replacing the original data; The verification and evaluation of the data after adding noise includes: Use data validation tools to validate the data after adding random noise; Perform a privacy assessment on the data after adding random noise; Performance evaluation was performed on the data after adding random noise.

7. The differential privacy algorithm optimization method for protecting enterprise credit data privacy according to claim 1, characterized in that, The analysis of query requirements, determination of query relevance, and merging of multiple data queries into a single data query include: Identify the use cases and business needs for obtaining enterprise credit data, and determine the relevance of the query. Analyze the relevance between different queries; If multiple queries involve the same data fields or the similarity of the query conditions exceeds a set threshold, they will be merged into one query. The selection of a merging strategy to optimize the query statement includes: Based on the relevance of the query and business requirements, select a merging strategy, which includes: condition merging, result merging, and function merging. When designing merge queries, optimize the query statements.

8. The differential privacy algorithm optimization method for protecting enterprise credit data privacy according to claim 4, characterized in that, The source of the enterprise credit data is clearly identified, and the enterprise credit data is preliminarily reviewed, including: Determine the source channels of corporate credit data and understand the data collection methods and timeframes; Conduct a preliminary review of the collected data to understand its format, field meanings, and data types; The outlier detection of enterprise credit data includes: Calculate the basic statistics of the data to understand the distribution of the data and identify outliers; Use visualization tools to show the distribution of the data; Outlier detection is achieved using model-based methods, which include cluster analysis and outlier detection algorithms. The removal of noise from enterprise credit data includes: Filtering methods are used to remove noise from enterprise credit data. These methods include mean filtering, median filtering, and Gaussian filtering. If the data contains noise that is correlated with other variables, use regression analysis to remove the noise. For time series data or data with continuous changing trends, data smoothing methods are used to remove noise. These methods include moving average and exponential smoothing.

9. A differential privacy algorithm optimization device for protecting enterprise credit data privacy, characterized in that, include: The query analysis module is used to analyze the query types of enterprise credit data and adaptively adjust noise parameters. A noise conditioning module for selectively increasing noise based on an improved noise mechanism; The query merging module is used to merge multiple data queries into a single data query; The query merging module is specifically used to analyze query requirements, determine query relevance, merge multiple data queries into one data query, select a merging strategy, and optimize the query statement. Modify the query interface, process query results, and monitor query performance; The data processing module is used to preprocess the noise-processed corporate credit data in order to protect the privacy of the corporate credit data through a differential privacy algorithm that introduces noise. The query merging module is specifically used for: Based on the designed merge query, modify the data query interface so that the query interface can accept the merged query request; After executing the merge query, the query results are processed. If a result merging strategy is used, the results of multiple queries will be merged and organized. If a function merging strategy is used, the corresponding function is called to process it; After implementing the merge query, use database performance monitoring tools or log analysis tools to monitor query performance, including query execution time, CPU utilization, and memory utilization. If a decrease in query performance is found, further optimize the query statement or adjust the merge strategy. Methods to optimize query statements include avoiding duplicate queries, using indexes, and limiting the query result set.

Citation Information

Patent Citations

  • Privacy budget allocating and data publishing method and privacy budget allocating and data publishing system for protecting data query privacy

    CN108537055A

  • Data sharing method and device based on differential privacy protection

    CN112613065A