Big data visual interactive display method and system
By using sparse projection and dynamic grid technology, combined with reinforcement learning strategies to generate projection results, the problem of achieving a real-time balance between privacy protection and data utility in existing technologies is solved, and adaptive adjustment of the projection direction and dynamic optimization of parameter configuration are achieved.
Patent Information
- Application Number
- CN202510766831.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing big data visualization systems find it difficult to achieve a real-time balance between privacy protection and data utility in dynamic user interaction scenarios, and their parameter configurations lack adaptive adjustment capabilities, resulting in uncontrollable privacy risks and low interaction efficiency.
Through sparse projection processing, low-dimensional projection results are generated, a dynamic grid index is constructed, user interaction behavior is captured and a dynamic intent vector is generated. The local and global projection matrices are combined for hybrid projection, the privacy protection strength is dynamically adjusted, and projection parameter recommendations are generated through reinforcement learning strategies to form a closed-loop optimization mechanism.
It achieves adaptive adjustment of the projection direction, dynamically balances privacy protection and data accuracy, and improves the efficiency and accuracy of the system's parameter configuration in diverse user interaction scenarios.
Smart Images

Figure CN120672560A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data interactive visualization, and in particular to a big data visualization interactive display method and system. Background Art
[0002] With the widespread application of big data visualization analysis in sensitive fields such as financial risk control and medical diagnosis, ensuring data privacy during interactions while maintaining visualization effectiveness has become a pressing technical challenge. Currently, visualization systems based on differential privacy typically employ static noise injection or fixed projection parameters. While these methods provide basic privacy protection, they struggle to adapt to dynamic user interactions and spatial variations in data distribution. These approaches often lead to two extremes: insufficient noise in high-density data areas can lead to group privacy leaks, while excessive noise in sparse areas can mask valid information, severely limiting the depth and reliability of visualization analysis.
[0003] The core flaw of existing technologies lies in the contradiction between their static design logic and the demands of dynamic interaction. On the one hand, the global equal distribution strategy for the privacy budget ignores variations in data spatial density, failing to achieve a dynamic trade-off between privacy protection and data fidelity. On the other hand, the manual configuration of projection parameters relies on empirical experience and lacks the ability to perceive users' real-time intent, resulting in a disconnect between visualization results and interactive behavior. This disconnected optimization mechanism leaves the system facing the dual bottlenecks of uncontrollable privacy risks and low interaction efficiency in complex analysis scenarios, making it difficult to meet users' needs for in-depth exploration of sensitive data.
[0004] In view of the above-mentioned deficiencies in the prior art, the present invention proposes a method and system for interactive visualization of big data. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a method and system for interactive visualization of big data, which solves the problems of difficulty in balancing privacy protection and data utility in real time in dynamic user interaction scenarios and lack of adaptive adjustment capabilities of parameter configuration.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for interactive visualization of big data, comprising the following steps:
[0007] S1. Perform sparse projection processing on the original high-dimensional data to generate low-dimensional projection results and construct a dynamic grid index;
[0008] S2. Capturing user interaction behaviors based on the dynamic grid index, extracting intent features, and generating a dynamic intent vector;
[0009] S3. Generate a local projection matrix based on the dynamic intent vector, combine it with the pre-trained global projection matrix, dynamically adjust the mixing weights through the time decay coefficient, and calculate and generate the mixed projection result;
[0010] S4. Identifying the user's focus area based on the hybrid projection result, performing recursive grid division and incremental statistical aggregation, dynamically adjusting the privacy protection strength, and outputting a dynamic aggregation result, wherein the dynamic aggregation result includes the focus area, grid statistics, and privacy parameters;
[0011] S5. Perform multi-view association rendering based on the dynamic aggregation result, generate a projection parameter recommendation list through a reinforcement learning strategy, and collect user behavior to generate feedback tuples;
[0012] S6. Generate feedback tuples based on user behavior, parse the adoption intention and behavior adjustment signals in the feedback tuples, iteratively update the intention weight matrix and reinforcement learning strategy structure, where the results of the iterative update will act on the next round of recommendation generation to achieve adaptive optimization and closed-loop learning of the system.
[0013] Preferably, the step S1 includes:
[0014] S1-1. Construct a sparse random global projection matrix W based on the original high-dimensional data g , calculate the global skeleton projection Among them, x i is the original high-dimensional data point;
[0015] S1-2, based on global skeleton projection Initialize the dynamic grid tree, divide the coarse-grained grid and store the statistics S j ;
[0016] S1-3, statistics S j Injecting Laplace noise to generate privacy-preserving dynamic grid indexes.
[0017] Preferably, the step S2 includes:
[0018] S2-1. Analyze the center coordinates of the user-selected area (c x , c y ) and area A, generate the regional feature vector φ R =[c x , c y ,A];
[0019] S2-2. Perform Fourier transform on the brush trajectory to extract the The low-frequency components generate the eigenvector φ T , where L is the number of sampling points of the brush trajectory, and p≤10. When L<20, p=L;
[0020] S2.3, the regional feature vector φ R and the trajectory eigenvector φ T Perform weighted splicing to generate the intention feature vector φ of the user interaction behavior at the current moment current =[αφ R , βφ T ], where α and β are dynamic weight coefficients;
[0021] S2-4. Intention feature vector φ based on the current user interaction behavior current , update the intention weight matrix M through the online gradient descent formula, the online gradient descent formula is:
[0022]
[0023] Among them, M t With M t+1 is the value of the intention weight matrix at time steps t and t+1, η is the learning rate, is the gradient operator, φ current is the intention feature vector of the user's interaction behavior at the current moment, ||φ current -M t φ history || 2 is the loss function, φ history is the sliding window mean of the historical intention feature vector, Among them, k is the window length, t is the time variable, is the intention feature vector of the user's interaction behavior at the current moment at time step i.
[0024] Preferably, the step S3 includes:
[0025] S3-1. Generate local projection matrix: According to the intention feature vector φ of the user's interaction behavior at the current moment current Generate the local projection matrix W l (x i );
[0026] S3-2. Get the global projection matrix: Load the pre-trained global projection matrix W from the data preprocessing module g , the pre-trained global projection matrix W g Obtained through offline training of historical data using a sparse autoencoder;
[0027] S3-3. Calculate the mixed projection weight: according to the time attenuation coefficient α(t) = α0e -λt Dynamically adjust the local projection intensity, where α0 is the initial weight of the local projection, λ is the decay rate parameter, and t is the time variable;
[0028] S3-4, mixed projection calculation: According to the local projection matrix and the global projection matrix, a mixed projection result is calculated based on a mixed projection calculation formula. The mixed projection result calculation formula is:
[0029]
[0030] Where α(t) is the time decay coefficient; is the original high-dimensional data point x i The final projection result; x i is the original high-dimensional data point; W g is the global projection matrix; W l (x i ) is the local projection matrix, which depends on the original high-dimensional data point x i Or dynamically generated by user intent.
[0031] Preferably, the step S4 includes:
[0032] S4-1. According to the norm of the current intention vector ||φ current ||2. Adaptively set the grid division granularity g and combine the hybrid projection results to identify the initial grid unit set with significant user attention.
[0033] S4-2. Initial grid cell set Each grid cell in performs recursive subdivision and incrementally updates the local statistics S of the sensitive features of each grid cell i , constructing a grid feature expression structure for differential privacy mechanism
[0034] S4-3. Based on the user interaction behavior of each attention grid unit, calculate the heat index h of each attention grid unit i , and dynamically adjust the privacy protection strength parameters based on the behavior-driven privacy adjustment formula The behavior-driven privacy adjustment formula is as follows:
[0035]
[0036] Where Z is the heat amplification coefficient, ρ base Is the basic privacy protection strength;
[0037] S4-4. Then, using the local statistic S i Based on the adjusted intensity parameters, a privacy protection objective function is constructed, and the optimal privacy budget is solved by the Lagrangian optimization method. Finally, the privacy strength adjustment is calculated Forming a privacy regulation structure As part of the dynamic aggregation result output, the privacy protection objective function is expressed as follows:
[0038]
[0039] in, is the target optimization function, To use the user experience cost or data availability cost brought by higher privacy strength, the coefficient of the loss function that weighs the privacy cost and accuracy, Error(S i ,∈ i ) is the data accuracy loss introduced by privacy perturbation.
[0040] Preferably, the step S5 includes:
[0041] S5-1, Visualization Rendering: Encode the dynamic aggregation results output by step S4 into a heat map, and perform multi-view association rendering with the hybrid projection scatter plot generated by step S3;
[0042] S5-2, Strategy Recommendation: Based on the intention feature vector of the user's interaction behavior at the current moment in step S2 and the time decay coefficient α(t) in step S3, the projection parameter combination (α * ,λ*,β*) recommendation list, and define the adoption flag simultaneously;
[0043] S5-3, Feedback Generation: Collect the adoption identifier δ and the privacy protection strength adjustment behavior of step S4, combined with the norm of the current intent vector ||φ current ||2, generate a matrix containing (δ,ΔW l ,||φ current ||2,∈ * , Δρ) feedback tuple.
[0044] Preferably, the step S6 includes:
[0045] S6-1. Feedback parsing: Parsing the feedback tuple (δ, ΔW l ,||φ current ||2,∈ * ,Δρ), extract the explicit rating value, the number of revocation operations, the privacy budget parameter and the intention strength indicator, and construct the feedback sample set for incremental model update;
[0046] S6-2, Intent Model Update: Based on the feedback samples, the intent weight matrix M of the current user interaction behavior is incrementally updated using the online gradient descent formula. The mapping ability of the intent weight matrix to the intent feature vector is adjusted. The updated results are fed back to step S2 for the next round of intent feature generation.
[0047] S6-3. Optimize the reinforcement learning strategy structure: Based on the adoption flag and privacy adjustment amount in the feedback, update the exploration rate parameter ∈ in the Q-learning strategy, and adjust the strategy score of each projection parameter combination in the recommendation strategy. The updated strategy structure will be used to generate a recommended list of projection parameter combinations in step S5-2.
[0048] The present invention also provides a big data visualization interactive display system, which is applied to the above-mentioned big data visualization interactive display method, including:
[0049] The data preprocessing module is used to perform sparse projection processing on the original high-dimensional data, generate low-dimensional projection results, and construct a dynamic grid index. The grid index is used to capture interactive behaviors and support subsequent privacy protection processing;
[0050] The intent perception module receives dynamic grid indexes, parses the user-selected area and brush trajectory, extracts regional and trajectory features, and weights and concatenates them to generate an intent feature vector. It then updates the intent weight matrix using an online gradient descent formula.
[0051] A dynamic projection module is used to receive the intention feature vector, generate a local projection matrix according to the current intention, and dynamically generate a hybrid projection result based on the time decay coefficient in combination with the pre-trained global projection matrix;
[0052] A dynamic aggregation module is configured to receive the hybrid projection results, identify the user's focus area based on the intent norm, perform recursive grid division and incremental statistical aggregation, calculate the heat index and adjust the privacy protection strength parameter based on the user's interactive behavior, and output the dynamic aggregation result, wherein the dynamic aggregation result includes the focus area, grid statistics, and privacy parameters;
[0053] The visualization rendering module is used to jointly render the dynamic aggregation results and the hybrid projection results into a multi-view interface, and generate a recommended list of projection parameter combinations based on the current intent features and the time decay factor through the Q-learning strategy;
[0054] The feedback iteration module is used to collect user behavioral feedback on the recommended combination, generate feedback tuples, analyze adoption intentions and behavior adjustment signals, iteratively update the intention weight matrix and reinforcement learning strategy structure, and feed back the updated results to the intention perception module and visualization rendering module to achieve system adaptive optimization and closed-loop learning.
[0055] Preferably, the data preprocessing module and the intention perception module transmit a dynamic grid index via a distributed message queue, and the dynamic grid index includes the following information:
[0056] Grid space coordinate range [x min , x max ]×[ymin ,y max ];
[0057] Statistical triplet after noise injection in,, is the noise mean of grid j, is the noise standard deviation of grid j, count j is the number of original data points in grid j.
[0058] Preferably, the intention perception module and the dynamic projection module meet the following real-time constraints:
[0059] The intention feature vector φ of the user's interaction behavior at the current moment current The generation delay does not exceed 50 milliseconds;
[0060] The intention feature vector φ of the user's interaction behavior at the current moment current It is passed to the dynamic projection module through zero-copy memory sharing to ensure the rapid generation and response of hybrid projection results.
[0061] The present invention provides a method and system for interactive visualization of big data, which has the following beneficial effects:
[0062] 1. This invention achieves adaptive adjustment of projection direction based on user interaction intent by extracting intent features in real time and updating the intent weight matrix online, combined with the dynamic generation of global-local hybrid projections. Compared to the view rigidity caused by fixed projection parameters in existing technologies, this method overcomes the inability to capture real-time user intent through local projection matrix adjustment.
[0063] 2. This paper dynamically allocates the privacy budget by constructing a Lagrangian optimization model and differentially adjusts the privacy protection strength parameters for areas with different data densities. Compared with the traditional uniform noise injection method, it effectively balances the privacy leakage risk in high-density areas and the data accuracy loss in low-density areas, solving the imbalance between utility and privacy protection caused by a unified allocation strategy.
[0064] 3. This invention recommends projection parameter combinations based on a Q-learning strategy and expands the state space by integrating user intent feature vectors, enabling cross-scenario adaptive parameter configuration. Compared to manually preset parameters or static rule bases, this overcomes the technical limitations of dynamically adjusting the exploration rate parameter to adapt to diverse user needs and real-time interaction scenarios.
[0065] 4. The present invention utilizes feedback tuple synchronization to optimize the intention weight matrix and the reinforcement learning policy structure, forming a bidirectional feedback closed loop. Compared with the traditional single-model update mechanism, it fills the technical gap of the separate operation between the intention perception module and the recommendation system, enabling the system to continuously improve the intention recognition accuracy and parameter recommendation accuracy during long-term interactions. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is a flowchart of the method of the present invention;
[0067] Figure 2 is a flow framework diagram of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0069] Please refer to the attached Figure 1 , an embodiment of the present invention provides a method for big data visualization interactive display, including the following steps:
[0070] S1. Perform sparse projection processing on the original high-dimensional data to generate a low-dimensional projection result and construct a dynamic grid index;
[0071] In this embodiment, first, sparse projection processing is performed on the original high-dimensional data to achieve data dimension compression and preliminary feature skeleton extraction.
[0072] The original high-dimensional data is represented as a set of data points where is the i-th high-dimensional data point, with dimension d and a total data volume of N.
[0073] Construct a sparse random global projection matrix The matrix satisfies the following conditions:
[0074] Each row vector follows a Laplace or Gaussian sparse distribution;
[0075] The proportion of non-zero elements does not exceed the set sparsity threshold θ, where θ ∈ (0, 1);
[0076] The projection dimension r << d to achieve dimension compression.
[0077] Multiply the original data x i on the left by the sparse global projection matrix W g , to obtain the global skeleton projection result:
[0078] in, Represents the low-dimensional global projection result of the i-th sample point.
[0079] Then, the result is projected onto the global skeleton As input, initialize the dynamic grid index structure. This structure is logically represented by a dynamic grid tree, and physically includes a node record unit, an index pointer unit, and a statistics storage unit. The specific structure includes:
[0080] Grid node unit, used to record the coordinate range of the current projection space;
[0081] Statistics unit, used to store data distribution information S falling within the node area j ;
[0082] The child node pointer unit is used to point to the sub-area after recursive division.
[0083] A coarse-grained grid division is performed on the low-dimensional space, and the entire space is divided into initial square grid areas with a side length of g0, wherein g0 is the minimum division unit set by system initialization.
[0084] Each grid node records the number of samples, mean, variance and other sensitive statistical information that fall within its range, forming an initial statistical set {S j},in,:
[0085] Among them, μ j is the sample mean in the grid cell, is the sample variance, n j is the number of samples.
[0086] In order to meet the requirements of the differential privacy mechanism, Laplace noise needs to be injected into the above statistics while constructing the dynamic grid index.
[0087] Specifically, for each grid cell statistic S i , perform the following noise injection operation:
[0088] Where Lap(·) represents the Laplace distribution sampling function, Δf is the sensitivity (i.e., the maximum change that each statistic may cause due to a single data change), and ∈ is the preset privacy budget parameter.
[0089] Statistics after noise injection It is stored in the grid node statistics unit and serves as the basic data in the subsequent differential privacy protection mechanism.
[0090] The final dynamic grid index structure can be used for user behavior capture, privacy adjustment, and dynamic aggregation result generation. It logically communicates with the intention feature extraction module in step S2 and serves as the spatial basis for intent vector generation.
[0091] Through the above-mentioned sparse projection, skeleton extraction, grid construction and differential privacy noise injection steps, this embodiment can effectively achieve spatial dimensionality reduction representation and index optimization structure construction of high-dimensional data while ensuring user data privacy, providing support for subsequent intent modeling and recommendation strategies.
[0092] S2. Capturing user interaction behaviors based on the dynamic grid index, extracting intent features, and generating a dynamic intent vector;
[0093] In this implementation, the user interaction behavior is captured based on the dynamic grid index, and the dynamic intent vector is generated by multimodal feature fusion. First, the spatial attributes of the user-selected area are analyzed. The selected area is defined by the user in the visual interface through mouse or touch operation, where the center coordinates (c x , c y ) and area A are obtained in real time through the geometric calculation module to generate the regional feature vector:
[0094] φ R =[c x , c y , A], where c x and c y are the coordinates of the center of the selected region in the projection space, and A is the area of the region (in pixels or standardized units).
[0095] At the same time, the time-frequency feature extraction of the user's brush trajectory is performed. The brush trajectory is recorded as a discrete point sequence by the trajectory sampling module. Where L is the number of sampling points. Perform a fast Fourier transform (FFT) on the trajectory coordinate sequence and retain the first p low-frequency components:
[0096] When L≥20;
[0097] p=Lwhen L<20;
[0098] The extracted frequency domain feature vector is:
[0099] φ T =[Re(F1),Im(F1),…,Re(F p ),Im(F p )];
[0100] Among them, F s represents the s-th Fourier coefficient, Re(·) and Im(·) represent the real part and imaginary part, respectively.
[0101] The regional feature vector φ R and the trajectory eigenvector φ T Perform dynamic weighted splicing to generate the current moment intention feature vector:
[0102] φ current =[αφ R , βφ T ];
[0103] Among them, the weight coefficients α and β are dynamically adjusted through user behavior feedback in the sliding window to satisfy α+β=1.
[0104] The intention weight matrix M is updated by the online gradient descent algorithm, and its update formula is:
[0105]
[0106] Among them, M t With M t+1 is the value of the intention weight matrix at time steps t and t+1, η is the learning rate, is the gradient operator, φ current is the intention feature vector of the user's interaction behavior at the current moment, ||φ current -M t φ history || 2 is the loss function;
[0107] φ history is the sliding window mean of the historical intention feature vector, Among them, k is the window length, t is the time variable, is the intention feature vector of the user's interaction behavior at the current moment at time step i.
[0108] The physical implementation of the intention weight matrix M includes a feature mapping unit and a weight storage unit, wherein the feature mapping unit maps historical intention features to the current feature space through matrix multiplication operations, and the weight storage unit stores matrix parameters through a non-volatile memory.
[0109] Through the above-mentioned multimodal feature extraction, dynamic weighted fusion and online learning mechanism, this embodiment can effectively capture the spatiotemporal evolution characteristics of user interaction intentions, and provide an explainable feature-driven basis for subsequent projection parameter recommendations.
[0110] S3. Generate a local projection matrix based on the dynamic intent vector, combine it with the pre-trained global projection matrix, dynamically adjust the mixing weights through the time decay coefficient, and calculate and generate the mixed projection result;
[0111] In this embodiment, the generation of mixed projection results includes four stages: local projection matrix construction, global projection loading, weight adjustment and mixed calculation. First, according to the intention feature vector φ of the user's interaction behavior at the current moment current and the original data point x i ,
[0112] W l (x i )=U·diag(φ current )·V T ;
[0113] in, and is the pre-trained basis matrix obtained through offline training of historical data, m is the implicit dimension, and diag(·) indicates that the vector is converted into a diagonal matrix.
[0114] The original data point x i Multiply it with the local projection matrix to get the local projection component:
[0115] W l (x i )x i =(U·diag(φ current )·V T )x i ;
[0116] Global projection matrix W g Loaded from the preprocessing module, it is optimized by the sparse autoencoder, and the objective function is:
[0117]
[0118] Among them, χ controls the sparsity, ‖·‖1 is the L1 norm, Represents the global projection matrix W g The transposed matrix of .
[0119] The mixing weight is dynamically adjusted by the time decay coefficient, and the calculation formula is:
[0120] α(t)=α0e -λt ;
[0121] in,:
[0122] α0∈[0.5,1.0] is the initial local weight;
[0123] λ>0 is the decay rate (default value λ=0.05);
[0124] t time variable (cumulative value of user interactions).
[0125] The final mixed projection result is calculated by the following formula:
[0126]
[0127] Where α(t) is the time decay coefficient; is the original high-dimensional data point x i The final projection result; x i is the original high-dimensional data point; W g is the global projection matrix; W l (x i ) is the local projection matrix, which depends on the original high-dimensional data point x i Or dynamically generated by user intent.
[0128] In terms of physical implementation, the hybrid projection module includes:
[0129] Basis matrix storage unit: used to store U and V;
[0130] Intent feature interface: receives the φ output from step S2 current ;
[0131] Matrix generation unit: According to W l (x i )=U·diag(φ current )·V T Real-time calculation of W l (x i );
[0132] Hybrid Compute Unit: Execution Linear combination operation.
[0133] In this embodiment, when the user performs five consecutive selection operations (t=5), the time decay coefficient is calculated as α(5)=0.8·e -0.05×5 ≈0.618, at which point the local projection weight decays to 77.3% of the initial value, reflecting the adaptive decay to recent intentions.
[0134] S4. Identifying the user's focus area based on the hybrid projection result, performing recursive grid division and incremental statistical aggregation, dynamically adjusting the privacy protection strength, and outputting a dynamic aggregation result, wherein the dynamic aggregation result includes the focus area, grid statistics, and privacy parameters;
[0135] In this embodiment, the generation of dynamic aggregation results includes four stages: focus area identification, recursive grid division, privacy strength adjustment and optimization solution. First, according to the intention feature vector φ of the user's interaction behavior at the current moment current The L2 norm sets the grid division granularity:
[0136]
[0137] Among them, g base is the reference granularity, μ is the normalization coefficient (the default value is μ=10.0), and g is the grid division granularity.
[0138] Based on the original high-dimensional data point x i The final projection result Identify user attention areas The judgment conditions are:
[0139]
[0140] Among them, n j is the number of samples in unit j, τ is the attention threshold (the default τ=1.5), is the total number of grids.
[0141] right Perform recursive subdivision, halving the edge length each time, until g is satisfied min (Minimum granularity threshold). Statistics S for each subdivision grid i Incremental updates:
[0142]
[0143] Where, ΔS i is the statistic increment of the newly added sample, is the grid statistics before updating, is the updated grid statistics (triplet: mean, variance, count).
[0144] Heat index h i Calculated by user interaction frequency:
[0145]
[0146] Privacy strength parameter The dynamic adjustment formula is:
[0147]
[0148] in,:
[0149] Z is the heat amplification factor (default Z = 0.1);
[0150] ρ base Basic privacy strength (default ρ base =0.5);
[0151] log(·) is the natural logarithm function.
[0152] The privacy protection objective function is defined as:
[0153]
[0154] in, is the target optimization function, To use the user experience cost or data availability cost brought by higher privacy strength, the coefficient of the loss function that weighs the privacy cost and accuracy, Error(S i ,∈ i ) is the data accuracy loss due to privacy perturbation, and the error term Privacy cost item
[0155] Solving the optimal privacy budget through Lagrange multiplier method The final output privacy adjustment structure is:
[0156]
[0157] in, is the privacy strength adjustment amount, The privacy adjustment structure is the final output and is used to guide noise injection.
[0158] In terms of physical implementation, the dynamic aggregation module includes:
[0159] Granularity calculation unit: According to the formula Real-time calculation of grid granularity;
[0160] Recursive subdivision controller: manages the mesh subdivision process;
[0161] Hot statistics unit: accumulates user interaction behavior data;
[0162] Privacy Optimization Solver: Execution Formula Optimization calculation.
[0163] In this embodiment, when a user interacts with a grid unit 10 times (time window 5 minutes), the heat h i =2, the privacy strength is adjusted to Reflects interaction-driven privacy enhancement.
[0164] S5. Perform multi-view association rendering based on the dynamic aggregation result, generate a projection parameter recommendation list through a reinforcement learning strategy, and collect user behavior to generate feedback tuples;
[0165] In this embodiment, multi-view associated rendering and parameter recommendation are implemented through the following process: First, the dynamic aggregation result output from step S4 is encoded into a heat map, and its color mapping function is:
[0166]
[0167] Among them, Cj is the color value of grid cell j, S j is the statistic for each grid cell j, S min and S max The minimum / maximum statistical value in the current view.
[0168] Multi-view association rendering is performed with S3's hybrid projection scatter plot. When the user selects an area in the heat map, the corresponding data point in the scatter plot is highlighted. The association logic is implemented through the spatial coordinate mapping module.
[0169] The projection parameters are recommended to adopt the Q-Learning strategy, whose state space is defined as:
[0170] s t =[φ current ,α(t),||φ current ||2];
[0171] The action space is the parameter combination a=(α * ,λ * ,β * ),in:
[0172] α * ∈[0.1,1.0]: local projection weight adjustment amount;
[0173] λ * ∈[0.01,0.1]: time decay rate adjustment;
[0174] β * ∈[0.1,0.5]: Trajectory feature weight adjustment amount.
[0175] The Q value update formula is:
[0176]
[0177] in,:
[0178] η = 0.1: learning rate;
[0179] γ = 0.9: discount factor;
[0180] r t : Instant reward, calculated as r t =δ·||φ current ||2-Δρ.
[0181] The feedback tuple generation module collects the user adoption flag δ (δ = 1 means adoption of the recommendation, δ = 0 means rejection), combines the privacy adjustment amount Δρ and the current intention strength ||φ current ||2, generate feedback tuple:
[0182] Feedback tuple = (δ, ΔW l ,||φ current ||2,∈ * , Δρ);
[0183] Where ΔW l is the adjustment of the local projection matrix, calculated by matrix difference:
[0184]
[0185] Where ΔW l is the adjustment matrix of the local projection matrix, reflecting the changes in projection parameters caused by user interaction; is the adjusted local projection matrix (dimension r×d); is the local projection matrix before adjustment (dimension r×d).
[0186] In terms of physical implementation, the recommendation system includes:
[0187] Rendering engine: performs coordinate mapping of heat maps and scatter plots;
[0188] Q table storage unit: saves the state-action value matrix;
[0189] Feedback collection interface: capture user interaction events in real time;
[0190] Parameter adjuster: Injects recommended parameters into steps S2-S4.
[0191] In the embodiment, when the user adopts the recommended parameter (α * =0.8, λ*=0.05, β*=0.3), the system updates the Q table and generates a feedback tuple for subsequent strategy optimization.
[0192] S6. Generate feedback tuples based on user behavior, analyze their adoption intentions and behavior adjustment signals, and iteratively update the intention weight matrix and reinforcement learning strategy structure. The results of the iterative update will be used to generate the next round of recommendations, realizing adaptive optimization and closed-loop learning of the system.
[0193] In this embodiment, the adaptive optimization of the system is achieved through three stages: feedback analysis, intention model update and strategy optimization. First, the feedback tuple (δ, ΔW l ,||φ current ||2,∈ * , Δρ), extract the following features to construct training samples:
[0194] Sample = (δ, ΔW l ,||φ current ||2,∈ * , Δρ,t);
[0195] Where t is the time variable, which is used to identify the temporal relationship of feedback; δ is the adoption flag (1 = adoption, 0 = rejection); ΔW l is the local projection matrix adjustment (from step S5); ||φ current ||2 is the L2 norm of the current intention vector;∈ * is the optimized privacy budget (from step S4); Δρ is the privacy strength adjustment (from step S4).
[0196] The online update formula of the intention weight matrix M is:
[0197]
[0198] Among them, η = 0.01 is the learning rate; φ current is the current intention feature vector (from step S2); ΔW l is the local projection matrix adjustment amount (from step S5); M new is the updated intention weight matrix; M old is the intention weight matrix before updating; η is the learning rate (default 0.01), which controls the update step size; φ current is the intention feature vector of the user's interaction behavior at the current moment; is the transpose of the intent vector.
[0199] The dynamic adjustment formula of the exploration rate parameter ∈ of the reinforcement learning strategy is:
[0200]
[0201] Where K = 100 is the decay coefficient. When the user adopts the recommendation (δ = 1), the exploration rate decreases to utilize the existing knowledge; ∈ new is the adjusted exploration rate; ∈ old is the exploration rate before adjustment (initial value 0.2); δ is the adoption flag; ||φ current ||2 is the L2 norm of the intention vector; exp(·) is the natural exponential function.
[0202] The Q table update rule is expanded to:
[0203]
[0204] Where Q(s,a) is the strategy score of action a in state s; η Q is the learning rate of Q learning; γ is the discount factor, which measures the importance of future rewards; is the maximum expected reward for the next state s′; r t For immediate rewards.
[0205] New reward function:
[0206] rt =δ·||φ current ||2-0.5·Δρ;
[0207] In terms of physical implementation, the adaptive optimization module includes:
[0208] Feedback parser: extracts key features from tuples;
[0209] Matrix update unit: performs the gradient descent calculation of formula (2);
[0210] Policy Optimizer: manages updates to the Q-table and exploration rate parameters.
[0211] In this embodiment, when the user adopts the recommendation (δ=1) and ||φ current ||2=8.5, the exploration rate is updated to:∈ new =0.2·exp(-8.5 / 100)≈0.184;
[0212] At the same time, the scores of the corresponding state-action pairs in the Q-table are increased to strengthen the effective recommendation strategy.
[0213] The big data visualization interactive display system described below and the big data visualization interactive display method described above can refer to each other.
[0214] Please see the attached Figure 2 A big data visualization interactive display system is applied to the above-mentioned big data visualization interactive display method, comprising:
[0215] The data preprocessing module is used to perform sparse projection processing on the original high-dimensional data, generate low-dimensional projection results, and construct a dynamic grid index. The grid index is used to capture interactive behaviors and support subsequent privacy protection processing;
[0216] The intent perception module receives dynamic grid indexes, parses the user-selected area and brush trajectory, extracts regional and trajectory features, and weights and concatenates them to generate an intent feature vector. It then updates the intent weight matrix using an online gradient descent formula.
[0217] A dynamic projection module is used to receive the intention feature vector, generate a local projection matrix according to the current intention, and dynamically generate a hybrid projection result based on the time decay coefficient in combination with the pre-trained global projection matrix;
[0218] A dynamic aggregation module is configured to receive the hybrid projection results, identify the user's focus area based on the intent norm, perform recursive grid division and incremental statistical aggregation, calculate the heat index and adjust the privacy protection strength parameter based on the user's interactive behavior, and output the dynamic aggregation result, wherein the dynamic aggregation result includes the focus area, grid statistics, and privacy parameters;
[0219] The visualization rendering module is used to jointly render the dynamic aggregation results and the hybrid projection results into a multi-view interface, and generate a recommended list of projection parameter combinations based on the current intent features and the time decay factor through the Q-learning strategy;
[0220] The feedback iteration module is used to collect user behavioral feedback on the recommended combination, generate feedback tuples, analyze adoption intentions and behavior adjustment signals, iteratively update the intention weight matrix and reinforcement learning strategy structure, and feed back the updated results to the intention perception module and visualization rendering module to achieve system adaptive optimization and closed-loop learning.
[0221] The data preprocessing module and the intent perception module transmit the dynamic grid index through a distributed message queue. The dynamic grid index contains the following information:
[0222] Grid space coordinate range [x min ,x max ]×[y min ,y max ];
[0223] Statistical triplet after noise injection in,, is the noise mean of grid j, is the noise standard deviation of grid j, count j is the number of original data points in grid j.
[0224] The intent perception module and the dynamic projection module meet the following real-time constraints:
[0225] The intention feature vector φ of the user's interaction behavior at the current moment current The generation delay does not exceed 50 milliseconds;
[0226] The intention feature vector φ of the user's interaction behavior at the current moment current It is passed to the dynamic projection module through zero-copy memory sharing to ensure the rapid generation and response of hybrid projection results.
[0227] The system of this embodiment can be used to execute the above method embodiments, and its principles and technical effects are similar, so they will not be repeated here.
[0228] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for interactive visualization of big data, characterized in that: The following steps are involved: S1. Perform sparse projection processing on the original high-dimensional data to generate low-dimensional projection results and construct a dynamic grid index; S2. Capturing user interaction behaviors based on the dynamic grid index, extracting intent features, and generating a dynamic intent vector; S3. Generate a local projection matrix based on the dynamic intent vector, combine it with the pre-trained global projection matrix, dynamically adjust the mixing weights through the time decay coefficient, and calculate and generate the mixed projection result; S4. Identifying the user's focus area based on the hybrid projection result, performing recursive grid division and incremental statistical aggregation, dynamically adjusting the privacy protection strength, and outputting a dynamic aggregation result, wherein the dynamic aggregation result includes the focus area, grid statistics, and privacy parameters; S5. Perform multi-view association rendering based on the dynamic aggregation result, generate a projection parameter recommendation list through a reinforcement learning strategy, and collect user behavior to generate feedback tuples; S6. Generate feedback tuples based on user behavior, parse the adoption intention and behavior adjustment signals in the feedback tuples, iteratively update the intention weight matrix and reinforcement learning strategy structure, where the results of the iterative update will act on the next round of recommendation generation to achieve adaptive optimization and closed-loop learning of the system.
2. A big data visualization interactive display method according to claim 1, characterized in that: The steps of S1 include: S1-1. Construct a sparse random global projection matrix W based on the original high-dimensional data g , calculate the global skeleton projection Among them, x i is the original high-dimensional data point; S1-2, based on global skeleton projection Initialize the dynamic grid tree, divide the coarse-grained grid and store the statistics S j ; S1-3, statistics S j Injecting Laplace noise to generate privacy-preserving dynamic grid indexes.
3. A big data visualization interactive display method according to claim 1, characterized in that: The steps of S2 include: S2-1. Analyze the center coordinates of the user-selected area (c x , c y ) and area A, generate the regional feature vector φ R =[c x ,c y ,A]; S2-2. Perform Fourier transform on the brush trajectory to extract the The low-frequency components generate the eigenvector φ T , where L is the number of sampling points of the brush trajectory, and p≤10. When L<20, p=L; S2.3, the regional feature vector φ R and the trajectory eigenvector φ T Perform weighted splicing to generate the intention feature vector φ of the user interaction behavior at the current moment current =[αφ R , βφ T ], where α and β are dynamic weight coefficients; S2-4. Intention feature vector φ based on the current user interaction behavior current , update the intention weight matrix M through the online gradient descent formula, the online gradient descent formula is: Among them, M t With M t+1 is the value of the intention weight matrix at time steps t and t+1, η is the learning rate, is the gradient operator, φ current is the intention feature vector of the user's interaction behavior at the current moment, ||φ current -M t φ history || 2 is the loss function, φ history is the sliding window mean of the historical intention feature vector, Among them, k is the window length, t is the time variable, is the intention feature vector of the user's interaction behavior at the current moment at time step i.
4. A method for interactive visualization of big data according to claim 1, characterized in that: The steps of S3 include: S3-1. Generate local projection matrix: According to the intention feature vector φ of the user's interaction behavior at the current moment current Generate the local projection matrix W l (x i ); S3-2. Get the global projection matrix: Load the pre-trained global projection matrix W from the data preprocessing module g , the pre-trained global projection matrix W g Obtained through offline training of historical data using a sparse autoencoder; S3-3. Calculate the mixed projection weight: according to the time attenuation coefficient α(t) = α0e -λt Dynamically adjust the local projection intensity, where α0 is the initial weight of the local projection, λ is the decay rate parameter, and t is the time variable; S3-4, mixed projection calculation: According to the local projection matrix and the global projection matrix, a mixed projection result is calculated based on a mixed projection calculation formula. The mixed projection result calculation formula is: Where α(t) is the time decay coefficient; is the original high-dimensional data point x i The final projection result; x i is the original high-dimensional data point; W g is the global projection matrix; W l (x i ) is the local projection matrix, which depends on the original high-dimensional data point x i Or dynamically generated by user intent.
5. The method for interactive visualization of big data according to claim 1, characterized in that: The steps of S4 include: S4-1. According to the norm of the current intention vector ||φ current ||2. Adaptively set the grid division granularity g and combine the hybrid projection results to identify the initial grid unit set with significant user attention. S4-2. Initial grid cell set Each grid cell in performs recursive subdivision and incrementally updates the local statistics S of the sensitive features of each grid cell i , constructing a grid feature expression structure for differential privacy mechanism S4-3. Based on the user interaction behavior of each attention grid unit, calculate the heat index h of each attention grid unit i , and dynamically adjust the privacy protection strength parameters based on the behavior-driven privacy adjustment formula The behavior-driven privacy adjustment formula is as follows: Where Z is the heat amplification coefficient, ρ base Is the basic privacy protection strength; S4-4. Then, using the local statistic S i Based on the adjusted intensity parameters, a privacy protection objective function is constructed, and the optimal privacy budget is solved by the Lagrangian optimization method. Finally, the privacy strength adjustment is calculated Forming a privacy regulation structure As part of the dynamic aggregation result output, the privacy protection objective function is expressed as follows: in, is the target optimization function, To use the user experience cost or data availability cost brought by higher privacy strength, the coefficient of the loss function that weighs the privacy cost and accuracy, Error(S i ,∈ i ) is the data accuracy loss introduced by privacy perturbation.
6. A big data visualization interactive display method according to claim 1, characterized in that: The steps of S5 include: S5-1, Visualization Rendering: Encode the dynamic aggregation results output by step S4 into a heat map, and perform multi-view association rendering with the hybrid projection scatter plot generated by step S3; S5-2, Strategy Recommendation: Based on the intention feature vector of the user's interaction behavior at the current moment in step S2 and the time decay coefficient α(t) in step S3, the projection parameter combination (α * ,λ*,β*) recommendation list, and define the adoption flag simultaneously; S5-3, Feedback Generation: Collect the adoption identifier δ and the privacy protection strength adjustment behavior of step S4, combined with the norm of the current intent vector ||φ current ||2, generate (δ, ΔW l ,||φ current ||2,∈ * , Δρ) feedback tuple.
7. A big data visualization interactive display method according to claim 1, characterized in that: The steps of S6 include: S6-1. Feedback parsing: Parsing the feedback tuple (δ, ΔW l ,||φ current ||2,∈ * ,Δρ), extract the explicit rating value, the number of revocation operations, the privacy budget parameter and the intention strength index, and construct the feedback sample set for incremental model update; S6-2, Intent Model Update: Based on the feedback samples, the intent weight matrix M of the current user interaction behavior is incrementally updated using the online gradient descent formula. The mapping ability of the intent weight matrix to the intent feature vector is adjusted. The updated results are fed back to step S2 for the next round of intent feature generation. S6-3. Optimize the reinforcement learning strategy structure: Based on the adoption flag and privacy adjustment amount in the feedback, update the exploration rate parameter ∈ in the Q-learning strategy, and adjust the strategy score of each projection parameter combination in the recommendation strategy. The updated strategy structure will be used to generate a recommended list of projection parameter combinations in step S5-2.
8. A big data visualization interactive display system, characterized in that: A method for interactive visualization of big data as described in any one of claims 1 to 7 above, comprising: The data preprocessing module is used to perform sparse projection processing on the original high-dimensional data, generate low-dimensional projection results, and construct a dynamic grid index. The grid index is used to capture interactive behaviors and support subsequent privacy protection processing; The intent perception module receives dynamic grid indexes, parses the user-selected area and brush trajectory, extracts regional and trajectory features, and weights and concatenates them to generate an intent feature vector. It then updates the intent weight matrix using an online gradient descent formula. A dynamic projection module is used to receive the intention feature vector, generate a local projection matrix according to the current intention, and dynamically generate a hybrid projection result based on the time decay coefficient in combination with the pre-trained global projection matrix; A dynamic aggregation module is configured to receive the hybrid projection results, identify the user's focus area based on the intent norm, perform recursive grid division and incremental statistical aggregation, calculate the heat index and adjust the privacy protection strength parameter based on the user's interactive behavior, and output the dynamic aggregation result, wherein the dynamic aggregation result includes the focus area, grid statistics, and privacy parameters; The visualization rendering module is used to jointly render the dynamic aggregation results and the hybrid projection results into a multi-view interface, and generate a recommended list of projection parameter combinations based on the current intent features and the time decay factor through the Q-learning strategy; The feedback iteration module is used to collect user behavioral feedback on the recommended combination, generate feedback tuples, analyze adoption intentions and behavior adjustment signals, iteratively update the intention weight matrix and reinforcement learning strategy structure, and feed back the updated results to the intention perception module and visualization rendering module to achieve system adaptive optimization and closed-loop learning.
9. A big data visualization interactive display system according to claim 8, characterized in that: The data preprocessing module and the intention perception module transmit the dynamic grid index through a distributed message queue. The dynamic grid index contains the following information: Grid space coordinate range [x min , x max ]×[y min ,y max ]; Statistical triplet after noise injection in,, is the noise mean of grid j, is the noise standard deviation of grid j, count j is the number of original data points in grid j.
10. A big data visualization interactive display system according to claim 8, characterized in that: The intention perception module and the dynamic projection module meet the following real-time constraints: The intention feature vector φ of the user's interaction behavior at the current moment current The generation delay does not exceed 50 milliseconds; The intention feature vector φ of the user's interaction behavior at the current moment current It is passed to the dynamic projection module through zero-copy memory sharing to ensure the rapid generation and response of hybrid projection results.
Citation Information
Patent Citations
User-centered personalized recommendation privacy protection method and user-centered personalized recommendation privacy protection system
CN112035755A
Personalized multi-view federal recommendation system
CN114564641A
Next-generation point-of-interest security recommendation strategy based on asynchronous advantage reinforcement learning
CN118484596A
Cloud side-end collaborative sparse traffic flow prediction privacy protection method and device
CN119066712A
Dynamic distribution method for privacy budget
CN119128974A