Data sharing privacy protection platform responding to tax digitization requirements
By constructing a data sharing privacy protection platform and employing technologies such as dynamic differential privacy mechanisms and fusion neural networks, the problem of blind spots in privacy risk identification in multi-source tax data sharing has been solved, realizing a closed loop of privacy protection throughout the entire process and improving the security and efficiency of tax data sharing.
Patent Information
- Application Number
- CN202511374740.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Existing tax data privacy protection technologies are ill-suited to the dynamic and ever-changing multi-source data sharing scenarios. Traditional methods cannot effectively integrate multi-source data, leaving blind spots in privacy risk identification and lacking a complete privacy protection loop, resulting in a high risk of privacy leaks.
By employing a dynamic differential privacy mechanism, a privacy feature extraction model, a security-aware array, a fusion neural network, and a multi-constraint optimization algorithm, a data sharing privacy protection platform is constructed to achieve real-time risk assessment and dynamic policy adjustment of multi-source tax information.
It enables accurate privacy risk assessment and dynamic protection of multi-source tax data, improves the security and efficiency of data sharing, forms a closed-loop privacy mechanism throughout the entire process, enhances the security and efficiency of tax data, and strengthens the reliability and efficiency of tax data sharing.
Smart Images

Figure CN120951387B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of tax data security, in particular to a data sharing privacy protection platform responding to the demand of tax digitization. BACKGROUND
[0002] With the in-depth development of digital economy, tax management is accelerating the transformation to digitization and intelligentization, and the sharing and integration of multi-source tax data has become an important way to improve the efficiency of tax administration and optimize the service experience. However, tax data contains a large amount of sensitive information such as enterprise operation information and personal income details, and there is always a risk of privacy leakage in the process of cross-department and cross-system data sharing.
[0003] Currently, tax data privacy protection relies on static encryption technology or single risk assessment method, which is difficult to adapt to dynamically changing data sources and complex sharing scenarios. The data formats and sensitivity levels of different tax information systems are different, and traditional privacy-aware methods cannot effectively integrate multi-source data, resulting in blind spots in privacy risk identification; the internal correlation of data changes dynamically over time during transmission and integration, and existing risk assessment models lack consideration of time sequence characteristics, making it difficult to accurately generate privacy distribution maps, thereby affecting the scientificity of sharing decisions. In addition, the adaptability between privacy protection strategies and actual sharing needs is insufficient, often resulting in excessive protection limiting data circulation efficiency, or insufficient protection leading to privacy leakage, and lacking effective feedback mechanisms for dynamic adjustment of strategies, unable to form a full-process privacy protection closed loop. The existence of these problems not only restricts the in-depth promotion of tax digitization, but also poses a potential threat to data security and the rights and interests of individuals and enterprises. SUMMARY
[0004] The purpose of the present application is to provide a data sharing privacy protection platform responding to the demand of tax digitization to solve the problems raised in the background art.
[0005] To achieve the above purpose, the present application provides a data sharing privacy protection platform responding to the demand of tax digitization, which comprises:
[0006] A tax data privacy-aware layer obtains multi-source tax information based on a dynamic differential privacy mechanism, applies a privacy feature extraction model to generate an initial privacy risk heat map, and performs sensitive area segmentation on the initial privacy risk heat map;
[0007] A sensitive information positioning and mapping layer configures a security-aware array according to the segmented sensitive areas, performs an encrypted data detection operation, calculates data internal correlation signals through a time series analysis algorithm, and generates a data internal privacy distribution stereogram;
[0008] The cross-source data security fusion layer establishes a spatiotemporal mapping relationship between tax data and privacy distribution, aligns time sources using a time synchronization protocol, and integrates the initial privacy risk heatmap and the three-dimensional map of internal data privacy distribution through a fusion neural network to generate a privacy leakage risk score.
[0009] The compliance-sharing decision-making layer uses a multi-constraint optimization algorithm to convert privacy leakage risk scores into executable sharing instructions, and sends the executable sharing instructions to the tax data processing system in real time through a secure transmission protocol.
[0010] The execution status feedback layer monitors the data change status after the tax data processing system executes executable shared instructions in real time, calculates the difference between the data change status and the preset security change threshold, generates privacy protection strategy evaluation indicators, and dynamically adjusts the privacy protection strategy until the privacy protection strategy evaluation indicators meet the optimization conditions.
[0011] Preferably, the method for obtaining multi-source tax information based on a dynamic differential privacy mechanism includes:
[0012] Multiple data collection modules are deployed in the tax data privacy awareness layer, including a tax declaration interface module, a corporate financial database module, and an economic activity monitoring module, which collect different types of multi-source tax information, including taxpayer identity data, corporate revenue data, and market transaction data.
[0013] Simulate a privacy protection mechanism, dynamically adjust the data sampling intensity, and define areas involving personal identity information as high-sensitivity areas and areas involving economic statistics as low-sensitivity areas based on a preset privacy risk model. Dynamically define a sensitivity function to divide high-sensitivity and low-sensitivity areas.
[0014] The privacy feature recognition model analyzes the privacy risk features in multi-source tax information in real time. The input of the privacy feature recognition model is multi-source tax information, and the output is the privacy risk feature region.
[0015] The boundaries of sensitive areas are dynamically updated based on privacy risk feature regions. If a new privacy risk feature region is identified, it is included in the high-sensitivity region; if no privacy risk features are detected in the high-sensitivity region, it is downgraded to a low-sensitivity region.
[0016] High-sensitivity areas employ high-intensity encrypted sampling, while low-sensitivity areas use low-intensity conventional sampling. By dynamically adjusting the parameter settings of the data acquisition module through control signals, the final multi-source tax information is obtained.
[0017] Preferably, the method for generating an initial privacy risk heatmap by applying a privacy feature extraction model to multi-source tax information includes:
[0018] The multi-source tax information is standardized and normalized, and the processed multi-source tax information is combined into a multi-dimensional tensor according to the data dimensions. The dimensions of the multi-dimensional tensor include time series, spatial location, data category and number of channels.
[0019] Construct a privacy feature extraction model structure, which includes a bottom-up feature extraction path, a top-down resolution recovery path, and cross-feature connections.
[0020] Multi-source tax information is input into the privacy feature extraction model structure. In the bottom-up feature extraction path, the convolutional neural network module is used to extract the feature representation of multi-source tax information, generate feature maps of different granularities, and gradually extract high-level semantic information.
[0021] In the top-down resolution recovery path, high-level semantic information is gradually restored to spatial resolution through upsampling operations;
[0022] Establish cross-feature connections between the bottom-up feature extraction path and the top-down resolution recovery path, and fuse feature maps of the same granularity;
[0023] Point convolution is applied to each output layer of the privacy feature extraction model structure to adjust the feature channels, integrate and fuse feature maps of different granularities, and generate an initial privacy risk heatmap.
[0024] Preferably, the method for segmenting sensitive regions in the initial privacy risk heatmap includes:
[0025] The privacy risk score for each region in the initial privacy risk heatmap is calculated by weighted aggregation of taxpayer identity data, enterprise revenue data, and market transaction data included in multi-source tax information.
[0026] Preset privacy risk scoring threshold one and privacy risk scoring threshold two, and compare the privacy risk score with privacy risk scoring threshold one and privacy risk scoring threshold two respectively;
[0027] If the privacy risk score is lower than the privacy risk score threshold, the area corresponding to the privacy risk score will be marked as a safe zone.
[0028] If the privacy risk score is higher than the privacy risk score threshold one but lower than the privacy risk score threshold two, then the area corresponding to the privacy risk score will be marked as a warning area.
[0029] If the privacy risk score is higher than the privacy risk score threshold two, then the area corresponding to the privacy risk score will be marked as a danger zone;
[0030] Different marking symbols are used to segment sensitive areas in the initial privacy risk heatmap. Different marking symbols represent different sensitive areas. The safe zone is defined as a low-risk sensitive area, the warning zone as a medium-risk sensitive area, and the danger zone as a high-risk sensitive area.
[0031] Preferably, the method for generating a three-dimensional map of the privacy distribution within the data includes:
[0032] Based on the segmented sensitive areas, security sensing arrays of different densities are configured on the surface of the tax data processing system;
[0033] The security sensing array includes encrypted detection units that employ differentiated detection strategies for different sensitive areas. The array sends encrypted detection signals into the data to perform an initial global scan.
[0034] The detection angle is dynamically adjusted based on the sensitive areas of the initial privacy risk heatmap. The detection angle interval for high-risk sensitive areas is smaller than that for medium-risk sensitive areas, and the detection angle interval for medium-risk sensitive areas is smaller than that for low-risk sensitive areas.
[0035] The application angle optimization algorithm dynamically adjusts the detection angle, introduces a multi-parameter coupled data model, and continuously corrects the parameter settings in the data model through an iterative time series analysis algorithm until the preset number of iterations is reached and then stops.
[0036] A data propagation matrix is constructed based on the path tracing algorithm, and the privacy distribution is solved by the regularization optimization method, finally generating a three-dimensional map of the internal privacy distribution of the data with sensitive labels.
[0037] Preferably, the method for establishing the spatiotemporal mapping relationship between tax data and privacy distribution includes:
[0038] A unified global coordinate system is defined with any point on the surface of the tax data processing system as the reference origin, the horizontal and vertical axes parallel to the surface of the tax data processing system, and the depth axis perpendicular to the surface of the tax data processing system.
[0039] Use calibration tools to calibrate the data acquisition module and obtain its internal and external parameters;
[0040] The data coordinates of the initial privacy risk heatmap are transformed to the acquisition coordinate system using the internal parameters of the data acquisition module, and then the acquisition coordinate system is transformed to the global coordinate system using the external parameters of the data acquisition module, thus obtaining the initial privacy risk heatmap in the global coordinate system.
[0041] Using the origin of the global coordinate system as a reference point, the security sensing array is calibrated to obtain the external parameters of the security sensing array;
[0042] By transforming the internal privacy distribution map of the data into the global coordinate system through the external parameters of the security-aware array, a three-dimensional map of the internal privacy distribution of the data in the global coordinate system is obtained.
[0043] In the global coordinate system, the initial privacy risk heatmap and the three-dimensional map of privacy distribution within the data are spatially aligned to establish a spatial mapping between tax data and privacy distribution;
[0044] The time source connected to the data acquisition module is set as the master time source, and the time source connected to the security sensing array is set as the slave time source. Time source synchronization is performed through a time synchronization protocol to establish a time mapping between tax data and privacy distribution.
[0045] Preferably, the method of integrating the initial privacy risk heatmap and the three-dimensional map of internal privacy distribution of data through a fusion neural network includes:
[0046] Based on the initial privacy risk heatmap and the three-dimensional map of privacy distribution within the data in the global coordinate system, an undirected network structure is constructed.
[0047] A multi-layer network architecture based on a multi-head attention mechanism is used to build a fusion neural network to fuse the initial privacy risk heatmap in the global coordinate system and the three-dimensional map of privacy distribution within the data.
[0048] A fusion neural network includes an input processing layer, a feature projection layer, a network attention layer, a cross-source interaction layer, and an output scoring layer;
[0049] An undirected network structure is used as the input to the input processing layer of a fusion neural network, and a privacy leakage risk score is generated through the output scoring layer.
[0050] Preferably, the method for constructing an undirected network structure includes:
[0051] Each sensitive region in the initial privacy risk heatmap in the global coordinate system is taken as a data node, and the feature vector of each sensitive region is extracted as the feature of the data node.
[0052] Each privacy-focused region in the three-dimensional map of the data's internal privacy distribution in the global coordinate system is taken as a distribution node, and the spatial feature vector of each privacy-focused region is extracted as the feature of the distribution node.
[0053] Collect all data nodes and distributed nodes to form a node set;
[0054] Traverse all data nodes and calculate the spatial distance between any two data nodes in the global coordinate system;
[0055] A preset data distance threshold is set. If the spatial distance between any two data nodes in the global coordinate system is less than the data distance threshold, a bidirectional connection is added between the corresponding two data nodes.
[0056] If the spatial distance between any two data nodes in the global coordinate system is greater than or equal to the data distance threshold, then no connection is added;
[0057] Traverse all distributed nodes and calculate the spatial distance between any two distributed nodes in the global coordinate system;
[0058] A preset distribution distance threshold is set. If the spatial distance between any two distribution nodes in the global coordinate system is less than the distribution distance threshold, a bidirectional connection is added between the corresponding two distribution nodes.
[0059] If the spatial distance between any two distributed nodes in the global coordinate system is greater than or equal to the distribution distance threshold, then no connection is added;
[0060] By performing a proximity search operation, the nearest distribution node is found for each data node, and a bidirectional connection is added between the data node and the corresponding nearest distribution node.
[0061] Collect all bidirectional connections to form a connection set;
[0062] Construct an undirected network structure based on the set of nodes and the set of connections.
[0063] Preferably, the method for converting privacy leakage risk scores into executable sharing instructions based on multi-constraint optimization algorithms includes:
[0064] The internal space of tax data is divided into small cubic units, each of which records the current privacy distribution value and privacy leakage risk score.
[0065] It also lists all adjustable tax-sharing parameters, including the adjustment range for each parameter, the cost of privacy impact, and the strength of historical impact on privacy breaches;
[0066] Three optimization objectives are established, including privacy and security objective one, privacy and security objective two, and economic efficiency objective;
[0067] Two types of constraints are set: hard constraints and flexible constraints.
[0068] Using a multi-constraint optimization algorithm, multiple sets of optimization schemes for shared tax parameters are randomly generated. The execution results of the three optimization objectives are evaluated for each scheme, and the execution result scores of the three optimization objectives are obtained. The execution result scores of the three optimization objectives are aggregated and calculated to obtain a comprehensive evaluation value.
[0069] Select the solution with the highest comprehensive evaluation value from multiple solutions as the final optimized solution;
[0070] The selected final optimization scheme is converted into actual executable shared instructions;
[0071] Executable sharing instructions include data encryption instructions, access control instructions, and privacy monitoring instructions;
[0072] The methods for obtaining the performance scores of the three optimization objectives include:
[0073] The percentage reduction of the maximum privacy breach value in tax data is used as the performance score for the first privacy and security objective.
[0074] The dispersion of privacy breach risk scores in different regions is calculated as the execution result score for privacy security objective two.
[0075] The total resource consumption resulting from adjustments to all tax-shared parameters is used as the performance score for the economic efficiency objective.
[0076] Preferably, the method for generating privacy protection strategy evaluation metrics includes:
[0077] Collect executable shared instruction parameters executed by the tax data processing system, synchronously monitor data change status, calculate the deviation between the data change status and the preset security change threshold, and generate normalized privacy protection strategy evaluation indicators.
[0078] Preset normalized privacy protection strategy evaluation index threshold value one and normalized privacy protection strategy evaluation index threshold value two;
[0079] If the normalized privacy protection strategy evaluation index is higher than the critical value of the normalized privacy protection strategy evaluation index, then the privacy protection strategy is judged to be effective.
[0080] If the normalized privacy protection strategy evaluation index is between the first and second threshold values of the normalized privacy protection strategy evaluation index, then it is determined that the privacy protection strategy needs to be continuously observed.
[0081] If the normalized privacy protection strategy evaluation index is lower than the normalized privacy protection strategy evaluation index threshold of 1, the privacy protection strategy is invalid, the privacy protection strategy should be optimized, and an early warning notification should be generated immediately.
[0082] Compared with the prior art, the beneficial effects of the present invention are:
[0083] Through a multi-layered collaborative architecture, intelligent processing of the entire process for tax data privacy protection is achieved. The tax data privacy perception layer adopts a dynamic differential privacy mechanism, which can adjust the privacy protection strength in real time according to data characteristics when acquiring multi-source tax information. Combined with the initial privacy risk heatmap generated by the privacy feature extraction model and sensitive area segmentation technology, it can accurately locate high-risk privacy areas in different data sources, solving the problem of insufficient adaptability of traditional static perception methods to multi-source data.
[0084] The security awareness array configured in the sensitive information location mapping layer can deeply mine the internal characteristics of data without leaking the original data through encrypted data detection operations. Combined with time series analysis algorithms, it can capture the dynamic changes of data correlation signals and generate a three-dimensional map of the internal privacy distribution of data. This provides a multi-dimensional and time-series reference for subsequent risk assessment, overcoming the shortcomings of traditional assessment models that do not adequately consider the dynamic correlation of data.
[0085] The spatiotemporal mapping relationship established by the cross-source data security fusion layer, combined with the time synchronization protocol, ensures the consistency of data from different sources and time dimensions during the fusion process. The integration of the initial privacy risk heat map and privacy distribution three-dimensional map by the fusion neural network can generate a comprehensive privacy leakage risk score by comprehensively considering multiple risk factors, making the risk assessment results more in line with actual sharing scenarios.
[0086] The compliance-sharing decision-making layer transforms risk scores into executable instructions based on a multi-constraint optimization algorithm. This satisfies both the business needs of data sharing and the compliance requirements of privacy protection, avoiding over-protection or under-protection. The execution status feedback layer monitors data changes in real time, calculates the difference from preset security thresholds, and generates privacy protection policy evaluation indicators. This drives dynamic adjustments to the policies, forming a closed-loop mechanism of "perception-evaluation-decision-feedback." This ensures that privacy protection measures remain adapted to changes in the data sharing scenario, improving the platform's continuous effectiveness in complex environments. Attached Figure Description
[0087] Figure 1 This is a schematic diagram illustrating the working principle of the data sharing and privacy protection platform for responding to the needs of tax digitization as described in this invention.
[0088] Figure 2 A flowchart for generating an initial privacy risk heatmap for a privacy feature extraction model;
[0089] Figure 3 A flowchart for generating a three-dimensional map of the internal privacy distribution of data;
[0090] Figure 4 A flowchart for establishing spatiotemporal mapping relationships; Detailed Implementation
[0091] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0092] Please see Figure 1 This invention provides a data sharing and privacy protection platform in response to the needs of tax digitization, the platform comprising:
[0093] This data sharing and privacy protection platform is designed based on the needs of tax digitization. It achieves privacy protection and secure sharing of multi-source tax data through a layered architecture. The tax data privacy awareness layer uses a dynamic differential privacy mechanism to collect multi-source tax information, generates an initial privacy risk heatmap through a privacy feature extraction model, and segments sensitive areas within the heatmap. The sensitive information location and mapping layer configures a security awareness array based on the segmentation results, performs encrypted data detection operations, and generates a three-dimensional map of the internal privacy distribution of the data using time-series analysis algorithms. The cross-source data security fusion layer establishes a spatiotemporal mapping relationship between tax data and privacy distribution, integrates the initial privacy risk heatmap and the three-dimensional privacy distribution map through a fusion neural network, and outputs a privacy leakage risk score. The compliance sharing decision layer transforms the risk score into executable sharing instructions based on a multi-constraint optimization algorithm and sends them to the tax data processing system via a secure transmission protocol. The execution status feedback layer monitors the data changes after the system executes the instructions in real time, calculates the difference from the preset security threshold, generates privacy protection strategy evaluation indicators, and dynamically adjusts the strategy until the optimization conditions are met.
[0094] Example 1: See Figure 2 The tax data privacy awareness layer achieves dynamic acquisition of multi-source tax information by deploying multiple data collection modules. The tax declaration interface module connects to the taxpayer declaration system, collecting structured data in real time, including personal identity information, income details, and tax records. The enterprise financial database module connects to the enterprise ERP system, extracting operational data such as financial statements, transaction records, and cost expenditures. The economic activity monitoring module integrates third-party market data, covering commodity transaction records, industry statistical indicators, and macroeconomic trends. These three modules operate in parallel, forming a multi-dimensional data collection network covering both micro-level individuals and the macro-level market.
[0095] The dynamic differential privacy mechanism implements differentiated privacy protection strategies during data collection. The privacy risk model divides data attributes into two sensitive areas: fields containing direct identifiers such as ID numbers and bank accounts are defined as high-sensitivity areas, while fields containing statistical information such as industry averages and regional economic output are classified as low-sensitivity areas. The sensitivity function dynamically calculates privacy weights based on field type, with the weight value increasing as the data granularity refines. For high-sensitivity areas, a strong perturbation algorithm based on Laplace noise is used, with noise intensity inversely proportional to the privacy budget; for low-sensitivity areas, Gaussian noise is applied for mild perturbation, reducing the risk of privacy leakage while ensuring statistical usability.
[0096] The privacy feature identification model employs a three-stage convolutional neural network architecture to process multi-source data. The first-stage network processes taxpayer identity data, capturing spatial correlations between fields such as ID numbers and addresses through spatial pyramid pooling. The second-stage network analyzes corporate financial data, using temporal convolutional layers to identify the periodic characteristics of indicators such as revenue and costs. The third-stage network integrates market transaction data, using a graph convolutional structure to model the topological relationships between goods, prices, and trading parties. The outputs of the three networks are fused through a feature concatenation layer and then input into a bidirectional long short-term memory network for temporal modeling, ultimately outputting a probability distribution map of privacy risk feature regions.
[0097] The sensitive area boundary update mechanism is based on a sliding time window. The feature recognition model performs a full scan of the input data every five minutes. When a new data field (such as newly added biometric information) is detected and its privacy risk probability exceeds a threshold, the high-sensitivity area boundary is automatically expanded. For high-sensitivity fields that do not trigger risk warnings for three consecutive cycles, the system gradually reduces their protection strength and moves them into the low-sensitivity area. The boundary adjustment process is fed back to the data acquisition module in real time via control signals, triggering dynamic switching of the sampling strategy. The high-sensitivity area uses an encrypted transmission channel, employing the AES-256 algorithm to encrypt the original data before perturbation processing; the low-sensitivity area maintains regular HTTPS transmission, with noise added only to key statistics.
[0098] The resolution restoration path employs transposed convolutions for upsampling, doubling the feature map size at each stage. The transposed convolution kernels are initialized with bilinear interpolation weights, and the optimal upsampling parameters are learned progressively during training. To avoid artifacts, a 3×3 thinning convolutional layer is applied after each upsampling stage for feature smoothing. Cross-feature connections fuse features of the same scale using element-wise addition, with the number of channels unified via a 1×1 convolution before fusion. The feature integration stage uses a multi-scale feature pyramid structure, unifying feature maps of different granularities to the highest resolution through nearest-neighbor interpolation before channel concatenation.
[0099] The final output initial privacy risk heatmap uses three-channel RGB encoding to represent the risk level. The intensity of the red channel is proportional to the probability of high-risk features, the green channel reflects the distribution of medium-risk features, and the blue channel represents the coverage of basic data. After the heatmap is generated, it undergoes non-maximum suppression processing to eliminate false risk points caused by local noise, and morphological closing operations are used to fill small holes, forming continuous risk area divisions. The heatmap coordinate system maintains a strict mapping relationship with the original tax data, with each pixel corresponding to the spatial location and time node of the actual data record, providing accurate input for subsequent sensitive area segmentation.
[0100] The sensitive region segmentation algorithm first calculates the comprehensive risk value of each pixel in the heatmap by weighted summation of the RGB three-channel values. The weighting coefficients are dynamically adjusted based on the tax data type; the weight of the red channel is increased for individual income tax-related areas, while the influence of the green channel is enhanced for corporate tax-related areas. The risk scoring thresholds are adaptively determined, using the 95th quantile of the heatmap risk value distribution as threshold two and the 85th quantile as threshold one. The segmentation process employs a region growing algorithm, merging adjacent similar risk regions starting from the seed pixel, combined with a watershed algorithm to handle regions with varying risk gradients, ultimately outputting a three-level sensitive region segmentation result with clearly defined boundaries.
[0101] The system maintains dynamic sampling logs during operation, recording the parameter adjustment history of each data acquisition module. Log information includes fields such as sampling time, data source, sensitive area type, noise intensity, and encryption method, and is protected by blockchain technology to ensure log immutability. The sampling strategy optimization module periodically analyzes log data to identify frequently accessed data fields and rare query patterns, adjusting the privacy budget allocation scheme accordingly. For fields that remain consistently highly sensitive, the system automatically triggers a data anonymization upgrade process, employing k-anonymization or l-diversity techniques to enhance protection; long-term idle low-sensitivity fields gradually release their privacy budget, improving data availability.
[0102] Data quality control is integrated throughout the entire privacy-aware process. After noise is added, statistical consistency checks are performed to ensure that the deviations of the perturbed mean, variance, and other statistical measures from the original data are controlled within preset ranges. Encrypted data undergoes a decryption-re-encryption loop to ensure algorithm correctness. An outlier detection mechanism is implemented in the feature extraction model output, triggering manual review for heatmap values exceeding reasonable limits. The system establishes a complete version control system; all changes to model parameters, threshold settings, and algorithm selections are recorded with version numbers and effective dates, supporting rapid rollback to any historical state.
[0103] Example 2: See Figure 3 The initial privacy risk heatmap segmentation of sensitive areas is achieved based on a multi-dimensional assessment system. The segmentation algorithm first grids the heatmap, with each grid cell corresponding to a specific set of tax data within a specific spatiotemporal range. When scoring the privacy risk within an assessment cell, three dimensions are comprehensively considered: data attribute characteristics, access frequency, and correlation. The data attribute dimension analyzes field types, including different combinations of weights for direct identifiers, quasi-identifiers, and sensitive attributes. The access frequency dimension counts historical query counts, with recently accessed high-frequency areas receiving a risk bonus. The correlation dimension detects the possibility of cross-dataset connections; combinations of fields that can be linked to reconstruct an individual's identity trigger risk escalation. The scoring calculation uses a weighted cumulative method, with the weights of each dimension dynamically configured according to the tax scenario. For corporate income tax declaration scenarios, the correlation dimension is emphasized, while for individual income tax scenarios, the influence of the data attribute dimension is strengthened.
[0104] Sensitive area classification employs a dynamic threshold mechanism. The system maintains a historical score distribution database, recording typical risk value ranges under different business scenarios. The initial values of threshold one and threshold two are set based on the statistical quantiles of this database and are continuously updated during operation. A lag interval design is introduced into the classification process: when a risk score crosses a threshold from low to high, it must exceed the threshold for three consecutive cycles before triggering an area upgrade; when crossing from high to low, it is immediately downgraded. This design avoids frequent switching of sensitive states due to short-term data fluctuations. The boundary of the danger zone is expanded outwards by a morphological dilation algorithm to a certain buffer range, preventing the omission of high-risk data due to positioning errors.
[0105] The security awareness array employs an adaptive density distribution strategy. The array consists of multiple programmable detection units, each containing three functional modules: data scanning, encryption computation, and communication. The array density exhibits a non-linear relationship with the sensitivity level of the area, with the number of detection units deployed in high-risk sensitive areas increasing quadratically. Communication links are established between units via a self-organizing network protocol, forming a multi-hop mesh topology. The array controller generates a deployment plan based on heatmap segmentation results and distributes task parameters to each detection unit via wireless configuration signals. After deployment, a self-test process is executed to verify the time synchronization accuracy and signal coverage continuity between units, and to initiate supplementary deployment procedures for detected coverage blind spots.
[0106] The encrypted data probing operation employs a multi-mode scanning mechanism. The basic scanning mode performs low-intensity, broad-spectrum probing across all data areas to obtain macroscopic characteristics of the data distribution. The targeted scanning mode targets sensitive areas, employing differentiated probing depth and precision based on the area's risk level. High-risk sensitive areas undergo full-field, bit-by-bit scanning, medium-risk areas are sampled, and low-risk areas only undergo metadata checks. The probing signals are protected using elliptic curve cryptography, with each probing packet carrying a timestamp and area identifier to prevent replay attacks and unauthorized probing. During the scanning process, data response characteristics are monitored in real-time, and abnormal response patterns (such as excessively long delays or verification errors) automatically trigger a secondary verification process.
[0107] Time series analysis algorithms process the raw signals acquired by the probes. The algorithm inputs spatiotemporal sequence data collected by multiple probe units, first performing temporal alignment and outlier removal. Temporal alignment is based on a high-precision clock synchronization protocol, mapping data collected by different units to a unified time axis. Outlier removal employs a sliding window outlier detection method, with the window size adaptively adjusted according to the data sampling rate. In the signal analysis stage, a multi-order autoregressive model is constructed to capture the temporal dependencies within the data. Robust fitting methods are used for model parameter estimation to reduce the impact of individual outlier data points. The analysis results are converted into correlation strength indices, reflecting the potential relationships between different data fields.
[0108] The generation of the internal privacy distribution stereo map involves multiple processing stages. After preprocessing, the raw detection data is used to construct a 3D voxel mesh model. The XY plane of the mesh corresponds to the data's logical structure, and the Z-axis represents the time dimension. Each voxel stores an association strength value and a risk marker, with the risk marker inherited from the segmentation results of the initial heatmap. Stereo map rendering employs a ray casting algorithm, calculating the cumulative projection of voxel data from different viewpoints. High-risk areas are highlighted in red, medium-risk areas are presented with a yellow gradient, and low-risk areas maintain a blue tone. Dynamic detail level adjustments are implemented during viewpoint changes; high-risk areas maintain high-resolution rendering throughout, while other areas automatically reduce rendering precision according to the display scale.
[0109] The detection angle optimization algorithm is implemented based on the feedback control principle. The algorithm maintains a detection efficiency evaluation matrix, recording the effective data acquisition rate for each angle combination. The matrix value is updated after each detection task, and subsequent tasks prioritize angle combinations with high historical efficiency. Angle adjustments follow a gradual principle, with each adjustment not exceeding a preset upper limit to avoid data gaps caused by sudden angle changes. Angle optimization is performed more frequently in high-risk sensitive areas than in other areas to ensure timely data updates. During optimization, the relationship between signal penetration depth and signal-to-noise ratio (SNR) is monitored; when the SNR falls below a threshold, the algorithm automatically switches to a low-penetration, high-precision mode.
[0110] The path tracing algorithm addresses the propagation characteristics of data in multi-hop probe paths. The algorithm treats each probe unit as a graph node, with data transmission paths between nodes as edges, constructing a complete propagation topology. Edge weights are calculated by considering factors such as transmission delay, packet loss rate, and encryption strength. Privacy distribution calculation is transformed into the problem of influence propagation among nodes in the graph, using a random walk model to simulate the diffusion process of data relationships. The algorithm outputs the potential impact range of each data field and marks sensitive paths that may trigger privacy cascading effects. These paths are highlighted in a 3D graph with a pulsed light effect, highlighting the need for secure isolation of related data.
[0111] The regularization optimization process balances the accuracy and completeness of the probe data. The optimization objective function includes three terms: a data fitting term measures the degree of fit between the reconstructed distribution and the original signal, a sparse constraint term promotes the suppression of irrelevant noise, and a smoothing term ensures the spatial continuity of the distribution. The optimization process employs an iterative solution using the alternating direction multiplier method, updating the distribution estimate and Lagrange multipliers in each iteration. The iteration termination condition comprehensively considers the convergence of the objective function, computational time budget, and result stability. The final output privacy distribution 3D map undergoes post-processing filtering to eliminate high-frequency noise introduced during the optimization process while preserving the true distribution details.
[0112] The sensitivity labeling system for the 3D map implements hierarchical management. The basic labeling layer inherits the segmentation results from the initial heatmap, the dynamic labeling layer reflects the latest detected risk changes, and the predictive labeling layer predicts future risk trends based on time-series analysis. Each layer of labels uses different visual encoding methods: basic labels use solid line boundaries, dynamic labels exhibit a flashing effect, and predictive labels are displayed as semi-transparent outlines. Label conflict handling follows a conservative principle; any high-risk area identified by any labeling layer is ultimately classified as a danger zone. The labeling management system records operation logs for each update and supports replaying the label evolution history along a timeline to aid in the analysis of risk propagation patterns.
[0113] Example 3: See Figure 4 The establishment of the spatiotemporal mapping relationship begins with the construction of the global coordinate system. Using an arbitrarily selected point on the surface of the tax data processing system as the origin O, a right-handed coordinate system is established: the X-axis extends horizontally to the right along the system surface, the Y-axis extends vertically upwards, and the Z-axis extends outwards perpendicular to the system surface. The coordinate system scale is normalized, mapping the system's maximum physical size to a unit length of 1. All data spatial positions are converted to relative coordinates within this coordinate system, eliminating the computational complexity caused by differences in the physical dimensions of different devices. The calibration process of the data acquisition module uses a precision optical positioning device to obtain the offset vector between the module's internal optical center point C and the global origin O. Δx represents the horizontal deviation, Δy represents the vertical deviation, and Δz represents the depth deviation. The rotation matrix R of the measurement module includes three Euler angle parameters: pitch angle α, yaw angle β, and roll angle γ.
[0114] The initial privacy risk heatmap data coordinate transformation involves two stages. The first stage transforms the heatmap pixel coordinates... Transform to the coordinate system of the acquisition module. The transformation relationship is determined by the module's internal parameters, including focal length. Principal point coordinates and distortion coefficient The conversion process first corrects radial and tangential distortion, then maps 2D pixels to the 3D acquisition space through perspective transformation. The second stage utilizes the extrinsic parameters obtained from calibration to map the 3D points in the acquisition coordinate system. Transform to global coordinate system, the transformation formula is as follows: The converted heatmap data includes global spatial coordinates. and timestamp This forms four-dimensional spatiotemporal data points.
[0115] The calibration of the security sensing array is performed using distributed calibration targets. Each detection unit in the array is equipped with a programmable LED marker, which flashes according to a preset pattern during calibration. A high-speed infrared camera captures the position of the marker, and triangulation is used to calculate the precise position of each unit in the global coordinate system. and orientation vector Calibration data is stored in the array control center, establishing a mapping table between unit identifiers (IDs) and spatial locations. Coordinate transformation of the privacy distribution 3D map is achieved by querying this mapping table, associating each probe data point with its corresponding global coordinates. Signal transmission delay is compensated during the transformation process, based on the fixed distance between the probe unit and the controller. and signal propagation speed Calculate the time compensation amount Adjust the data timestamp to .
[0116] The spatiotemporal alignment algorithm processes the transformed heatmap and 3D image data. Spatial alignment employs an iterative nearest-point algorithm, selecting at least four matching point pairs in the overlapping region and calculating the optimal rigid body transformation to minimize the mean square error. Temporal alignment is based on sliding window correlation analysis, with the window size... The frequency is adaptively adjusted based on the data update frequency, ranging from 100ms to 5s. Clock drift is detected during alignment; if the master-slave time source deviation exceeds a threshold... When this occurs, the time synchronization protocol is triggered to recalibrate. The synchronization protocol uses an improved PTPv2 standard, adding message types and verification fields for tax data.
[0117] The establishment of connections is based on a hybrid similarity metric, specifically spatial similarity. Calculate the Euclidean distance between nodes Convert to Gaussian kernel function Feature similarity Node representation is calculated using cosine similarity. and The cosine of the angle between them. Time similarity. The overlap of node activity time windows is evaluated. The overall similarity score is calculated as follows:
[0118] ;
[0119] in, , , For adjustable weight parameters, satisfying Constraints. Connection threshold. The value is dynamically adjusted based on network sparsity requirements, with a default value of 0.65. For node pairs with similarity exceeding a threshold, a bidirectional connection edge is established. edge weight It is proportional to the similarity score. Cross-source connections force each data node to establish edges with its three nearest distributed nodes, ensuring information exchange between the two types of nodes.
[0120] The network training process employs a two-stage strategy. The pre-training stage uses historical labeled data to minimize the mean squared error of score prediction. ,in The actual risk values are labeled by experts. Adversarial training is introduced during the fine-tuning phase, using a discriminator network to distinguish between predicted and actual score distributions, and an adversarial term is added to the loss function. The optimizer uses a variant of AdamW, with an initial learning rate of [value missing]. The performance decays by 15% every 20 epochs. The training data is divided into training, validation, and test sets in a 7:2:1 ratio. The validation set is used for the early stopping mechanism, and the test set is used only for the final performance evaluation.
[0121] The spatiotemporal mapping visualization monitoring interface offers multi-dimensional displays. The 3D spatial view displays the fusion of heatmaps and stereoscopic images, supporting transparency adjustment and layered display. The timeline view presents the evolution trend of risk scores and can mark key event points. The network view displays the connection strength between nodes, using side width to represent similarity scores. The interface integrates interactive analysis tools, supporting the generation of statistical reports by selecting specific spatiotemporal regions, including regional risk indicators, the number of associated nodes, and feature distribution histograms. All views are interconnected; selections in any view automatically update the displayed content of other views.
[0122] The system maintains version management of mapping relationships during runtime. Each coordinate transformation or alignment operation generates a version snapshot, recording operation parameters and result verification values. The version rollback function allows restoration to any historical state for system recovery in abnormal situations. A version difference analysis tool visualizes the areas of change between different versions, aiding in locating the root cause of problems. The mapping relationship database implements an incremental backup strategy, performing a full backup of the basic coordinate system parameters daily and incremental update data backups hourly. Backup data is stored in an off-site disaster recovery center after AES-256 encryption to ensure system disaster recovery capabilities.
[0123] Example 4: The optimization scheme for tax-sharing parameters is generated using a genetic algorithm framework. The initial population contains 50 sets of random parameter combinations, each containing 15 adjustable parameters. An example of an excellent scheme parameter from a certain iteration is as follows: data encryption level 4, access frequency limited to 6 times / hour, retention period of 90 days, field anonymization ratio of 75%, audit log level 3, and data watermark strength of 2. Scheme evaluation shows that this combination increases the privacy leakage rate of unit E-37 by 42%, optimizes the risk score dispersion index to 0.19, and controls resource consumption to 900MB. After 200 generations of evolution, the algorithm outputs the optimal solution set on the Pareto front, from which the decision engine selects a compromise scheme that balances security and efficiency.
[0124] The hard constraints are set based on tax regulations. All solutions must meet the following requirements: individual income tax data encryption level ≥3, enterprise financial data retention period ≥180 days, and access to sensitive fields must be two-factor authentication. Flexible constraints allow for adjustable parameters including: resource consumption limits can be appropriately relaxed provided that privacy security objective one (reduction in leakage value) is not less than 35%; when the economic efficiency objective (resource consumption) exceeds the threshold, privacy security objective two (scoring dispersion) is allowed to fluctuate by ±0.05. During one optimization process, the system detected that the retention period of solution A-207 was set to 75 days (lower than the hard constraint on enterprise data), automatically triggering the constraint violation handling procedure and adjusting it to the minimum allowable value of 180 days.
[0125] Command transmission employs a fragmented encryption mechanism. Each command is split into multiple data packets, each individually encrypted and digitally signed. End-to-end verification is implemented during transmission, with the receiver returning a verification result for each data packet. A transmission log shows that the command sequence contained 28 data packets, each strictly limited to 1KB in size and encrypted using different temporary session keys. The third data packet failed verification due to network jitter; the system automatically retransmitted it instead of continuing subsequent packet transmissions, ensuring command integrity. All successfully received packets are reassembled in memory, undergo integrity and authorization checks, and then submitted to the execution engine.
[0126] The monitoring module tracks the effects of commands. The system maintains a command execution status table, recording the parameter change history and current status of each cube unit. Data from a certain monitoring period shows that after unit E-42 executed the encryption command, the actual encryption strength reached the expected level of 4.1 (target 4.0); the access frequency was limited to fluctuations of 6.2 times / hour (set value 6 times); privacy leakage monitoring showed a reduction rate that increased from 32% to 39%. Anomaly detection found that the watermark command execution effect of unit E-55 deviated from the expected value by 15%, triggering the automatic compensation mechanism, dynamically adjusting the watermark embedding parameters to a strength level of 2.3 (original setting 2.0).
[0127] The sharing strategy is continuously improved through dynamic optimization cycles. Every 24 hours, the system reassesses the status of all cube units, initiating local re-optimization for units whose performance declines by more than 10%. One periodic assessment revealed a resurgence in privacy leaks in units E-19 to E-27 in the commercial area. Analysis indicated this was due to increased data association complexity caused by new enterprise access. The optimization algorithm generated enhanced instructions for this area: encryption level increased to 4.5, access frequency reduced to 4 times / hour, and differential privacy protection added (instruction code DIFFP-2). These new instructions are rolled out during off-peak hours to avoid the impact of concentrated updates on system performance.
[0128] The visualization analysis interface provides a panoramic view of the optimization process. The 3D map uses colored blocks to represent cubic units, with color depth corresponding to risk levels and border thickness indicating optimization priority. The parameter adjustment panel lists all adjustable parameter sliders for the currently selected unit, displaying real-time impact curves on the three optimization objectives. Historical trend charts compare and display unit performance metrics under different strategies, supporting filtering by time range. Decision-makers can directly modify optimization weights through drag-and-drop interaction; the system immediately recalculates and displays the new Pareto front solution set.
[0129] An anomaly handling mechanism ensures stable operation of the optimization process. When three consecutive unit optimization failures are detected, the system automatically switches to conservative mode: maintaining existing parameters and only performing monitoring and data collection. If a hardware failure causes the optimization engine to malfunction, the system records the last valid state before the failure and resumes execution from the point of interruption after service recovery. Detailed diagnostic reports are generated for all anomalies, including environmental parameters, resource usage, and operation logs, for the technical team to analyze the root cause. A periodic maintenance window is used for application optimization algorithm updates; new algorithm versions are first tested on 5% of units, and the scope is gradually expanded after verifying their effectiveness.
[0130] Example 5: The monitoring system for the execution status feedback layer adopts a multi-channel data acquisition architecture to capture the instruction execution status of the tax data processing system in real time. The monitoring channels include a system log parser, a network traffic sniffer, and a memory snapshot agent, each collecting runtime data from different dimensions. The log parser extracts call records of executable shared instructions, including structured fields such as instruction type, execution timestamp, and return code. The network traffic sniffer listens to data transmission ports, recording the source and destination addresses, transmission protocols, and payload characteristics of data packets. The memory snapshot agent collects stack information of key processes at fixed intervals, monitoring buffer usage and object reference relationships. The timestamps of the three types of collectors are synchronized using a high-precision clock to ensure the temporal consistency of data across channels.
[0131] The quantitative calculation of data change states is implemented based on a difference detection algorithm. The algorithm maintains a baseline state database, storing snapshots of system data features before the execution of instructions. Each time new data is detected, a feature vector is extracted and compared with the baseline snapshot to calculate a similarity score. The feature vector contains twenty dimensions, covering core indicators such as data structure integrity, field value distribution, and complexity of relationships. Similarity calculation uses a weighted Hamming distance, assigning higher weights to sensitive fields. Difference values are normalized to the [0,1] interval using a sigmoid function; larger values indicate greater deviation from the baseline state. A sliding window smoothing process is introduced during the calculation to eliminate misjudgments caused by instantaneous fluctuations.
[0132] The security change threshold is set using a dynamic adjustment strategy. The system initially loads an industry-standard threshold configuration file, containing suggested security ranges for different instruction types. During operation, the thresholds are dynamically optimized based on actual business scenarios. For frequently executed instruction types, the system automatically records historical security difference values and uses the 95th percentile as the new threshold upper limit. When environmental parameters change significantly (such as system upgrades or increased data volume), a threshold recalibration process is triggered. The calibration process uses an isolated testing mode, simulating the execution of typical instruction sequences in a shadow system to observe the changing trends of the security difference range.
[0133] The calculation of the normalized privacy protection strategy evaluation index integrates multiple influencing factors. Core factors include instruction execution completeness, data discrepancy deviation, and risk control coverage. Execution completeness assesses the actual completion of each step of the instruction, imposing penalties on sub-commands that are not executed or partially executed. Data discrepancy deviation measures the gap between the current state and the ideal security state, considering both the absolute value and the rate of change. Risk control coverage calculates the proportion of protected data fields to total sensitive fields, weighting the exposure risk of uncovered fields according to their sensitivity level. These three factors are aggregated using a geometric mean to form the final evaluation index, avoiding the dominance of a single factor in the evaluation results.
[0134] The generation of early warning notifications is tightly integrated with the tiered response mechanism. When the evaluation metric falls below threshold one, the system determines that the current strategy has completely failed and immediately initiates a level three response: terminating the execution of relevant instructions, rolling back to the previous security state, and generating a red early warning notification. The notification includes the event number, occurrence time, scope of impact, and emergency remediation suggestions, and is sent to the security management team through a dedicated channel. When the metric is between threshold one and two, a level two response is triggered: instructions continue to execute but the processing speed is reduced, the monitoring frequency is increased to once per minute, and a yellow observation notification is generated. The notification includes detailed monitoring data, requiring administrators to manually confirm subsequent actions. When the metric exceeds threshold two, the system only records green normal logs, maintaining the regular monitoring rhythm.
[0135] The dynamic adjustment process employs an incremental optimization method. Each adjustment performs a local search within the parameter space of the current strategy, avoiding the computational overhead of global re-optimization. The direction of adjustment is determined by the changing trend of the evaluation metrics: if the metrics continue to decline, the privacy protection strength parameter is strengthened; if the metrics fluctuate steadily, resource allocation efficiency is optimized; if the metrics rebound, non-critical constraints are gradually relaxed. The magnitude of parameter adjustments is proportional to the degree of metric deviation, and an S-curve is used to control the adjustment speed to prevent parameter oscillations. An observation period is set after each adjustment, during which further modifications are frozen to ensure system stability before evaluating the adjustment effect.
[0136] The closed-loop optimization mechanism achieves continuous improvement through feedback loops. Monitoring data is input into the strategy analysis engine, which identifies effective and ineffective strategy components. Effective components are extracted as pattern rules and added to the excellent strategy knowledge base; ineffective components generate negative examples for training the anomaly detection model. Rules in the knowledge base are stored according to application scenarios, supporting strategy recommendations based on similarity matching. When a new scenario emerges, the system retrieves the most relevant historical rules as the initial strategy, significantly shortening the optimization convergence time. The feedback loop cycle is dynamically adjusted according to system load, extending to once every 2 hours during peak business periods and shortening to once every 15 minutes during off-peak periods.
[0137] The strategy evaluation report generation module provides decision support. The report is automatically generated according to a preset template and includes four main parts: an overview of the execution instructions, a summary of monitoring data, trends in evaluation indicators, and adjustment recommendations. The overview section uses a visual flowchart to show the instruction execution path, marking completed and incomplete nodes. The data summary section selects key monitoring indicators and displays expected and actual values using time-series curves. The trend analysis section calculates the moving average and standard deviation of the evaluation indicators to identify long-term patterns of change. The recommendations section lists three optional adjustment schemes, focusing on safety, efficiency, and balance respectively, and includes a projected impact analysis for each scheme. The report is stored in encrypted PDF format and includes a digital signature to ensure integrity.
[0138] The version control system manages the complete evolution history of privacy protection policies. Each policy adjustment generates a new version number, following the naming convention of major version.minor version.revision number. The version repository stores policy configuration files, evaluation metric snapshots, and explanations of the reasons for adjustments—a unified data set. The repository supports filtering versions by time range, instruction type, and evaluation results, quickly identifying the best policy for specific scenarios. A difference comparison tool visualizes parameter changes between adjacent versions, highlighting significant adjustments with color. The rollback function allows loading any historical version and restoring to the policy configuration of that version. Version control is integrated with the change approval process; major policy adjustments require multi-level authorization confirmation.
[0139] The system self-test module periodically verifies the reliability of the monitoring link. The self-test process simulates real-world business scenarios to generate test commands, which are then injected into each stage of the monitoring system. These test commands include data changes in preset patterns to verify monitoring sensitivity and the accuracy of difference calculations. Network latency testing checks the clock synchronization status of each data collector, recalibrating nodes with deviations exceeding 50ms. Storage integrity testing verifies the consistency of writing and reading monitoring data, using checksums to detect silent errors. The self-test results generate a health score; components below the threshold are automatically isolated and alarms are triggered. Routine self-tests are performed daily, with additional special checks added after system upgrades or configuration changes.
[0140] The visual monitoring interface provides a real-time view of the execution status. The main dashboard displays five core evaluation metrics using a radar chart, with the outer ring showing the safety change threshold boundaries. The instruction execution view presents the data processing pipeline using a topology diagram, with node size indicating load intensity and color indicating health status. The difference analysis view displays the baseline and current data feature distribution side-by-side, highlighting significant deviations. The strategy adjustment history is presented in timeline format, with clickable parameters for each adjustment. The interface supports multi-screen interaction; filtering operations performed by the administrator in any view will synchronously update the data range in other views. The alarm panel centrally displays active alert notifications, sorted by urgency, and supports one-click confirmation and handling.
[0141] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0142] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A data sharing and privacy protection platform in response to the needs of tax digitization, characterized in that, include: The tax data privacy awareness layer acquires multi-source tax information based on a dynamic differential privacy mechanism, applies a privacy feature extraction model to the multi-source tax information to generate an initial privacy risk heatmap, and performs sensitive area segmentation on the initial privacy risk heatmap. The sensitive information location mapping layer configures a security sensing array based on the segmented sensitive areas, performs encrypted data detection operations, calculates internal correlation signals within the data using time series analysis algorithms, and generates a three-dimensional map of the internal privacy distribution of the data, including: The security sensing array configured in the sensitive information location mapping layer, through encrypted data detection operations, deeply mines the internal features of the data without disclosing the original data, and combines time series analysis algorithms to capture the dynamic changes of data correlation signals, generating a three-dimensional map of the privacy distribution within the data. The cross-source data security fusion layer establishes a spatiotemporal mapping relationship between tax data and privacy distribution, aligns time sources using a time synchronization protocol, and integrates the initial privacy risk heatmap and the three-dimensional map of internal data privacy distribution through a fusion neural network to generate a privacy leakage risk score. The compliance-sharing decision-making layer uses a multi-constraint optimization algorithm to convert privacy leakage risk scores into executable sharing instructions, and sends the executable sharing instructions to the tax data processing system in real time through a secure transmission protocol. The execution status feedback layer monitors the data change status of the tax data processing system after executing executable shared instructions in real time, calculates the difference between the data change status and the preset security change threshold, generates privacy protection strategy evaluation indicators, and dynamically adjusts the privacy protection strategy until the privacy protection strategy evaluation indicators meet the optimization conditions. The method for generating a three-dimensional map of the internal privacy distribution of the data includes: Based on the segmented sensitive areas, security sensing arrays of different densities are configured on the surface of the tax data processing system; The security sensing array includes encrypted detection units that employ differentiated detection strategies for different sensitive areas. The array sends encrypted detection signals into the data to perform an initial global scan. The detection angle is dynamically adjusted based on the sensitive areas of the initial privacy risk heatmap. The detection angle interval for high-risk sensitive areas is smaller than that for medium-risk sensitive areas, and the detection angle interval for medium-risk sensitive areas is smaller than that for low-risk sensitive areas. The application angle optimization algorithm dynamically adjusts the detection angle, introduces a multi-parameter coupled data model, and continuously corrects the parameter settings in the data model through an iterative time series analysis algorithm until the preset number of iterations is reached and then stops. A data propagation matrix is constructed based on the path tracing algorithm, and the privacy distribution is solved by the regularization optimization method. Finally, a three-dimensional map of the internal privacy distribution of the data with sensitive labels is generated. The method for establishing the spatiotemporal mapping relationship between tax data and privacy distribution includes: A unified global coordinate system is defined with any point on the surface of the tax data processing system as the reference origin, the horizontal and vertical axes parallel to the surface of the tax data processing system, and the depth axis perpendicular to the surface of the tax data processing system. Use calibration tools to calibrate the data acquisition module and obtain its internal and external parameters; The data coordinates of the initial privacy risk heatmap are transformed to the acquisition coordinate system using the internal parameters of the data acquisition module, and then the acquisition coordinate system is transformed to the global coordinate system using the external parameters of the data acquisition module, thus obtaining the initial privacy risk heatmap in the global coordinate system. Using the origin of the global coordinate system as a reference point, the security sensing array is calibrated to obtain the external parameters of the security sensing array; By transforming the internal privacy distribution map of the data into the global coordinate system through the external parameters of the security-aware array, a three-dimensional map of the internal privacy distribution of the data in the global coordinate system is obtained. In the global coordinate system, the initial privacy risk heatmap and the three-dimensional map of privacy distribution within the data are spatially aligned to establish a spatial mapping between tax data and privacy distribution; Set the time source connected to the data acquisition module as the master time source and the time source connected to the security sensing array as the slave time source. Synchronize the time sources through a time synchronization protocol to establish a time mapping between tax data and privacy distribution. The method for converting privacy leakage risk scores into executable sharing instructions based on multi-constraint optimization algorithms includes: The internal space of tax data is divided into small cubic units, each of which records the current privacy distribution value and privacy leakage risk score. It also lists all adjustable tax-sharing parameters, including the adjustment range for each parameter, the cost of privacy impact, and the strength of historical impact on privacy breaches; Three optimization objectives are established, including privacy and security objective one, privacy and security objective two, and economic efficiency objective; Two types of constraints are set: hard constraints and flexible constraints. Using a multi-constraint optimization algorithm, multiple sets of optimization schemes for shared tax parameters are randomly generated. The execution results of the three optimization objectives are evaluated for each scheme, and the execution result scores of the three optimization objectives are obtained. The execution result scores of the three optimization objectives are aggregated and calculated to obtain a comprehensive evaluation value. Select the solution with the highest comprehensive evaluation value from multiple solutions as the final optimized solution; The selected final optimization scheme is converted into actual executable shared instructions; Executable sharing instructions include data encryption instructions, access control instructions, and privacy monitoring instructions; The methods for obtaining the performance scores of the three optimization objectives include: The percentage reduction of the maximum privacy breach value in tax data is used as the performance score for the first privacy and security objective. The dispersion of privacy breach risk scores in different regions is calculated as the execution result score for privacy security objective two. The total resource consumption resulting from adjustments to all tax-shared parameters is used as the performance score for the economic efficiency objective.
2. A data sharing and privacy protection platform in response to the needs of tax digitization as described in claim 1, characterized in that, The method for obtaining multi-source tax information based on a dynamic differential privacy mechanism includes: Multiple data collection modules are deployed in the tax data privacy awareness layer, including a tax declaration interface module, a corporate financial database module, and an economic activity monitoring module, which collect different types of multi-source tax information, including taxpayer identity data, corporate revenue data, and market transaction data. Simulate a privacy protection mechanism, dynamically adjust the data sampling intensity, and define areas involving personal identity information as high-sensitivity areas and areas involving economic statistics as low-sensitivity areas based on a preset privacy risk model. Dynamically define a sensitivity function to divide high-sensitivity and low-sensitivity areas. The privacy feature recognition model analyzes the privacy risk features in multi-source tax information in real time. The input of the privacy feature recognition model is multi-source tax information, and the output is the privacy risk feature region. The boundaries of sensitive areas are dynamically updated based on privacy risk feature regions. If a new privacy risk feature region is identified, it is included in the high-sensitivity region; if no privacy risk features are detected in the high-sensitivity region, it is downgraded to a low-sensitivity region. High-sensitivity areas employ high-intensity encrypted sampling, while low-sensitivity areas use low-intensity conventional sampling. By dynamically adjusting the parameter settings of the data acquisition module through control signals, the final multi-source tax information is obtained.
3. A data sharing and privacy protection platform in response to the needs of tax digitization as described in claim 2, characterized in that, The method for generating an initial privacy risk heatmap by applying a privacy feature extraction model to multi-source tax information includes: The multi-source tax information is standardized and normalized, and the processed multi-source tax information is combined into a multi-dimensional tensor according to the data dimensions. The dimensions of the multi-dimensional tensor include time series, spatial location, data category and number of channels. Construct a privacy feature extraction model structure, which includes a bottom-up feature extraction path, a top-down resolution recovery path, and cross-feature connections. Multi-source tax information is input into the privacy feature extraction model structure. In the bottom-up feature extraction path, the convolutional neural network module is used to extract the feature representation of multi-source tax information, generate feature maps of different granularities, and gradually extract high-level semantic information. In the top-down resolution recovery path, high-level semantic information is gradually restored to spatial resolution through upsampling operations; Establish cross-feature connections between the bottom-up feature extraction path and the top-down resolution recovery path, and fuse feature maps of the same granularity; Point convolution is applied to each output layer of the privacy feature extraction model structure to adjust the feature channels, integrate and fuse feature maps of different granularities, and generate an initial privacy risk heatmap.
4. A data sharing and privacy protection platform for responding to the needs of tax digitization as described in claim 3, characterized in that, The method for segmenting sensitive regions in the initial privacy risk heatmap includes: The privacy risk score for each region in the initial privacy risk heatmap is calculated by weighted aggregation of taxpayer identity data, enterprise revenue data, and market transaction data included in multi-source tax information. Preset privacy risk scoring threshold one and privacy risk scoring threshold two, and compare the privacy risk score with privacy risk scoring threshold one and privacy risk scoring threshold two respectively; If the privacy risk score is lower than the privacy risk score threshold, the area corresponding to the privacy risk score will be marked as a safe zone. If the privacy risk score is higher than the privacy risk score threshold one but lower than the privacy risk score threshold two, then the area corresponding to the privacy risk score will be marked as a warning area. If the privacy risk score is higher than the privacy risk score threshold two, then the area corresponding to the privacy risk score will be marked as a danger zone; Different marking symbols are used to segment sensitive areas in the initial privacy risk heatmap. Different marking symbols represent different sensitive areas. The safe zone is defined as a low-risk sensitive area, the warning zone as a medium-risk sensitive area, and the danger zone as a high-risk sensitive area.
5. A data sharing and privacy protection platform in response to the needs of tax digitization as described in claim 1, characterized in that, The method of integrating an initial privacy risk heatmap and a three-dimensional map of internal privacy distribution in data through a fusion neural network includes: Based on the initial privacy risk heatmap and the three-dimensional map of privacy distribution within the data in the global coordinate system, an undirected network structure is constructed. A multi-layer network architecture based on a multi-head attention mechanism is used to build a fusion neural network to fuse the initial privacy risk heatmap in the global coordinate system and the three-dimensional map of privacy distribution within the data. A fusion neural network includes an input processing layer, a feature projection layer, a network attention layer, a cross-source interaction layer, and an output scoring layer; An undirected network structure is used as the input to the input processing layer of a fusion neural network, and a privacy leakage risk score is generated through the output scoring layer.
6. A data sharing and privacy protection platform in response to the needs of tax digitization as described in claim 5, characterized in that, The method for constructing an undirected network structure includes: Each sensitive region in the initial privacy risk heatmap in the global coordinate system is taken as a data node, and the feature vector of each sensitive region is extracted as the feature of the data node. Each privacy-focused region in the three-dimensional map of the data's internal privacy distribution in the global coordinate system is taken as a distribution node, and the spatial feature vector of each privacy-focused region is extracted as the feature of the distribution node. Collect all data nodes and distributed nodes to form a node set; Traverse all data nodes and calculate the spatial distance between any two data nodes in the global coordinate system; A preset data distance threshold is set. If the spatial distance between any two data nodes in the global coordinate system is less than the data distance threshold, a bidirectional connection is added between the corresponding two data nodes. If the spatial distance between any two data nodes in the global coordinate system is greater than or equal to the data distance threshold, then no connection is added; Traverse all distributed nodes and calculate the spatial distance between any two distributed nodes in the global coordinate system; A preset distribution distance threshold is set. If the spatial distance between any two distribution nodes in the global coordinate system is less than the distribution distance threshold, a bidirectional connection is added between the corresponding two distribution nodes. If the spatial distance between any two distributed nodes in the global coordinate system is greater than or equal to the distribution distance threshold, then no connection is added; By performing a proximity search operation, the nearest distribution node is found for each data node, and a bidirectional connection is added between the data node and the corresponding nearest distribution node. Collect all bidirectional connections to form a connection set; Construct an undirected network structure based on the set of nodes and the set of connections.
7. A data sharing and privacy protection platform in response to the needs of tax digitization as described in claim 1, characterized in that, The method for generating privacy protection strategy evaluation metrics includes: Collect executable shared instruction parameters executed by the tax data processing system, synchronously monitor data change status, calculate the deviation between the data change status and the preset security change threshold, and generate normalized privacy protection strategy evaluation indicators. Preset normalized privacy protection strategy evaluation index threshold value one and normalized privacy protection strategy evaluation index threshold value two; If the normalized privacy protection strategy evaluation index is higher than the critical value of the normalized privacy protection strategy evaluation index, then the privacy protection strategy is judged to be effective. If the normalized privacy protection strategy evaluation index is between the first and second threshold values of the normalized privacy protection strategy evaluation index, then it is determined that the privacy protection strategy needs to be continuously observed. If the normalized privacy protection strategy evaluation index is lower than the normalized privacy protection strategy evaluation index threshold of 1, the privacy protection strategy is invalid, the privacy protection strategy should be optimized, and an early warning notification should be generated immediately.
Citation Information
Patent Citations
Archive management method and system based on big data
CN118551414A
Cross-industry data security sharing method and system based on data desensitization and medium
CN120223391A