Data security sharing method under AI platform
By constructing a confusion center and edge nodes in the AI platform, generating dynamic confusion bubbles, and combining game boundary models and vector confusion enhancement models, the problem of mismatch between protection strategies and needs in data sharing under the AI platform is solved. Dynamic encryption and multi-dimensional authorization of data are realized, reducing the risk of data leakage and improving the security and efficiency of the system.
Patent Information
- Application Number
- CN202511303280.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-11-28
AI Technical Summary
In the data sharing process under the existing AI platform, the obfuscation protection strategy does not match the actual security needs, resulting in a high risk of data leakage. Furthermore, traditional methods cannot be dynamically adjusted to adapt to complex and ever-changing data access patterns and real-time threat environments.
In the AI platform, obfuscation center nodes and edge nodes are established. Dynamic obfuscation bubbles are generated through knowledge vector extraction and vectorization. Combined with game boundary model and vector obfuscation enhancement model, dynamic encryption and multi-dimensional authorization of data are achieved. Multi-level authentication technology is used to ensure data security.
It achieves a precise match between obfuscation protection strategies and actual security needs, reduces the risk of data leakage, improves the security and efficiency of data sharing systems, and can monitor and respond to abnormal access behavior in real time, maintaining data availability and semantic integrity.
Smart Images

Figure CN121037091A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of AI platforms, and in particular, relates to a data security sharing method under an AI platform. BACKGROUND
[0002] In the field of artificial intelligence platform data sharing, traditional data protection technologies mainly use static encryption, access control, and data desensitization methods to ensure data security. These traditional technologies are widely used in enterprise-level data sharing platforms, medical data exchange systems, financial data service platforms, and other scenarios in cloud computing environments. By establishing fixed security boundaries and permission management mechanisms, data access is controlled. However, traditional data protection methods have significant drawbacks. The static protection strategy they use cannot be dynamically adjusted according to changes in data usage scenarios and threat environments, resulting in a mismatch between protection strength and actual needs. In current AI platform data sharing applications, due to complex and variable data access patterns, dynamic changes in user permission requirements, and real-time updates of threat intelligence, traditional technologies cannot accurately match the confusion protection strategy with actual security needs, which can easily lead to over-protection resulting in low system efficiency or insufficient protection leading to data leakage risks. That is, there is a technical problem in the prior art that the confusion protection strategy does not match the actual security needs in the data sharing process under the AI platform, leading to data leakage risks. SUMMARY
[0003] Therefore, the present application provides a data security sharing method under an AI platform, which can solve the technical problem of data leakage risks caused by the mismatch between the confusion protection strategy and the actual security needs in the data sharing process under the AI platform in the prior art.
[0004] The application is implemented in the following manner: the application provides a data security sharing method under an AI platform, which includes establishing a confusion center node and a plurality of confusion edge nodes in an AI platform management server, deploying a data monitoring module to collect data access records and operation behavior information in real time; performing knowledge vector extraction on data to be shared, converting the data into a vector representation in a high-dimensional knowledge vector space through vectorization processing, automatically extracting tool feature labels, and calculating the sensitivity weight value of each knowledge vector; generating a plurality of confusion bubbles in the vector space based on the distribution characteristics of the knowledge vectors, and constructing a multi-dimensional authorization strategy; separating the cross-overlapping areas of the confusion bubbles using a game boundary model, determining the attribution boundary of the cross-overlapping areas through game solving; establishing a data training and reasoning mapping relationship according to the confusion bubble separation results, calculating the confusion strength parameters through a confusion strength optimization function using a multi-factor authentication technology; starting a vector confusion enhancement model to perform differential privacy processing on the data vectors in the confusion bubbles, and realizing dynamic encryption of the data; establishing a multi-level authentication mechanism, performing distributed verification on the user identity through the confusion edge nodes, automatically classifying and labeling the data using AI technology, and automatically distributing the data to the corresponding isolated areas according to the preset strategy; and performing real-time confusion protection during the data transmission process, and monitoring the data flow in real time to prevent data leakage and illegal sharing.
[0005] The confusion center node is responsible for generating confusion bubble identifiers and confusion vector ranges, dynamically updating the position and size of the confusion bubbles by analyzing historical data access patterns and security threat intelligence, and ensuring that the confusion protection strategy matches the actual security requirements. The confusion edge nodes are distributed at various access portals of the platform and are responsible for performing local confusion operations.
[0006] The knowledge vector extraction process uses semantic encoding technology to convert raw data into a vector representation containing semantic information, while preserving the core features of the data for subsequent similarity matching and retrieval operations. AI technology is used to analyze historical usage data and user evaluations of tools to automatically extract tool feature labels.
[0007] The confusion bubbles are distributed in a high-dimensional vector space in the shape of an ellipsoid. The vector range of each confusion bubble covers the corresponding knowledge vector area, and the adjacent confusion bubbles form cross-overlapping areas. The center coordinates of each confusion bubble correspond to a knowledge cluster center, and the lengths of the major and minor axes of the ellipsoid are determined according to the distribution density of the data in the knowledge cluster center.
[0008] The multi-dimensional authorization strategy is constructed based on role permissions, attribute permissions, data sensitivity permissions, and operation behavior permissions. The area of the cross-overlapping region accounts for 15% to 25% of the total area of each confusion bubble. The game boundary model determines the attribution confusion bubble of the data vector in the cross-overlapping region by solving the Nash equilibrium.
[0009] The game boundary model includes an upper game model aiming to maximize data protection intensity and a lower game model aiming to minimize computing resource consumption, and the data training and inference mapping relationship is established based on the life cycle stage of the AI model, and the data vector in each confusion bubble is associated with the corresponding AI model training stage and inference stage.
[0010] The data vector of the training stage is assigned to a high-security-level confusion bubble, the data vector of the inference stage is assigned to a standard-security-level confusion bubble, the data vector of the verification stage is assigned to a medium-security-level confusion bubble, and the multi-factor authentication technology combines username and password authentication, fingerprint recognition authentication, face recognition authentication, and SMS verification code authentication.
[0011] The objective function of the upper game model is to maximize the data protection intensity function, the input of the data protection intensity function includes sensitivity weight value, access frequency statistical value, user permission level value, and threat detection score, and the output is a data protection intensity coefficient, and the sensitivity weight value is derived from the calculation result of knowledge vector extraction.
[0012] The objective function of the lower game model is to minimize the computing resource consumption function, the input of the computing resource consumption function includes processor occupancy, memory usage, network bandwidth demand, and storage space occupancy, and the output is a resource consumption weight coefficient, and the processor occupancy and memory usage are derived from the system performance monitoring data of the data monitoring module.
[0013] The objective functions of the upper game model and the lower game model are associated through a security efficiency balance coupling term, the security efficiency balance coupling term establishes a constraint condition based on the product relationship between the data protection intensity coefficient and the resource consumption weight coefficient, and the data protection intensity coefficient and the resource consumption weight coefficient are used for game equilibrium solution value calculation of the confusion intensity optimization function.
[0014] The confusion intensity optimization function is used to calculate the confusion intensity parameter suitable for the current data sharing scenario according to the solution result of the game boundary model and the data training and inference mapping relationship, the input includes knowledge vector sensitivity weight value, user permission level value, data access frequency statistical value, sharing range identifier code, and game equilibrium solution value, and the output is a confusion intensity parameter value between 0 and 1.
[0015] As Figure 2As shown, the structure of the vector confusion enhancement model is a deep generative model based on a variational autoencoder architecture, including three main components: an encoder, a latent space transformation layer, and a decoder. The encoder maps the input knowledge vector to the latent space. The latent space transformation layer adjusts the randomness scale parameter of noise addition according to the confusion strength parameter. The decoder reconstructs the output vector from the confused latent vector. According to the confusion strength parameter and the data training inference mapping relationship, the temperature parameter of the noise addition mechanism is dynamically adjusted.
[0016] The training data set of the vector confusion enhancement model includes collecting a large number of multi-domain knowledge vector samples as original training data, labeling the sensitivity level and knowledge category information of each sample, constructing training sample pairs containing original vectors, target confusion vectors, and semantic consistency constraints, and expanding the training set size through data augmentation techniques to improve the generalization ability and robustness of the model.
[0017] The vector confusion enhancement model training step specifically includes training the generator and discriminator networks simultaneously using an adversarial training strategy. The generator learns to effectively confuse the knowledge vector while preserving semantic information. The discriminator learns to distinguish between the original vector and the confused vector. The model parameters are optimized by minimizing the weighted combination of reconstruction loss and semantic preservation loss. During training, a curriculum learning strategy is used to gradually increase the confusion difficulty to improve model performance.
[0018] The multi-level authentication mechanism combines user biometric recognition, behavior pattern analysis, and dynamic permission verification. The data protection strength function uses a nonlinear combination method to process input parameters. The sensitivity weight value and threat detection score are exponentially operated, and the access frequency statistics value and user permission level value are logarithmically operated. Then, the two operation results are combined through weighted average to obtain the final data protection strength coefficient.
[0019] The dynamic encryption strategy uses different encryption algorithms and key management strategies according to data sensitivity and user permissions. The computational resource consumption function establishes a nonlinear relationship between input parameters through polynomial fitting. The square term of processor occupancy rate is multiplied by the cubic root term of memory usage. Network bandwidth demand and storage space occupation are coupled through a trigonometric function relationship. Dynamic encryption strategy and vector perturbation technology are continuously applied during data transmission from the confusion bubble to the target node to ensure data security.
[0020] The application realizes accurate matching of the confusion protection strategy and actual security requirements by establishing a distributed architecture of confusion center nodes and edge nodes and dynamically adjusting the confusion bubble distribution combined with a game boundary model. The application dynamically updates the confusion strength parameters according to real-time data access patterns and threat intelligence by using a vector confusion enhancement model, overcomes the defects that the traditional static protection strategy cannot adapt to complex changing environments, and ensures that the protection strength is neither excessive nor insufficient through multi-dimensional authorization strategies and game equilibrium solving. The application solves the fundamental problem of mismatching between the confusion protection strategy and actual security requirements from the technical principle by vectorizing the data and constructing dynamic confusion bubbles in a high-dimensional space, and balances security and efficiency combined with a double-layer game model, thereby effectively reducing the data leakage risk. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 is a flowchart of the method of the application.
[0022] Figure 2 is a schematic diagram of the vector confusion enhancement model involved in the application.
[0023] Figure 3 is a schematic diagram of the confusion bubble distribution in a high-dimensional vector space.
[0024] Figure 4 is a performance monitoring data graph during the operation of the system in Example 2. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical scheme and advantages of the embodiments of the application clearer, the technical scheme in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application.
[0026] As Figure 1 shown, is a flowchart of a data security sharing method under an AI platform provided by the application, and the method comprises the following steps:
[0027] S01, confusion center nodes and a plurality of confusion edge nodes are established in an AI platform management server, the confusion center nodes are responsible for generating confusion bubble identifiers and confusion vector ranges, the confusion edge nodes are distributed at various access portals of the platform and are responsible for performing local confusion operations, and a data monitoring module is deployed to collect data access records and operation behavior information in real time;
[0028] S02, knowledge vectors are extracted from data to be shared, the data is converted into a vector representation in a high-dimensional knowledge vector space through vectorization processing, tool feature labels are automatically extracted by using historical use data and user evaluation of an AI technology analysis tool, and the sensitivity weight values of each knowledge vector are calculated;
[0029] S03, generate multiple confusion bubbles in the vector space based on the distribution characteristics of the knowledge vectors, the vector range of each confusion bubble covers the corresponding knowledge vector area, and an overlapping area is formed between adjacent confusion bubbles, and a multi-dimensional authorization strategy based on role permission, attribute permission, data sensitivity permission and operation behavior permission is constructed;
[0030] S04, separate the overlapping area of the confusion bubbles by using a game boundary model, the game boundary model includes an upper game model aiming to maximize data protection strength and a lower game model aiming to minimize calculation resource consumption, and the attribution boundary of the overlapping area is determined by game solving;
[0031] S05, establish a data training and inference mapping relationship according to the confusion bubble separation result, associate the data vectors in each confusion bubble with the corresponding AI model training stage and inference stage, use multi-factor authentication technology combined with username and password authentication, fingerprint recognition authentication, face recognition authentication, and SMS verification code authentication, and calculate the confusion strength parameter by using a confusion strength optimization function;
[0032] S06, start a vector confusion enhancement model to perform differential privacy processing on the data vectors in the confusion bubbles, and use different encryption algorithms and key management strategies to realize dynamic encryption of data according to data sensitivity and user permissions, the vector confusion enhancement model dynamically adjusts the temperature parameter of its noise adding mechanism according to the confusion strength parameter and the data training and inference mapping relationship;
[0033] S07, establish a multi-level authentication mechanism, combine user biometric recognition, behavior pattern analysis and dynamic permission verification, perform distributed verification on user identity through confusion edge nodes, and use AI technology to automatically classify and label data and automatically distribute data to corresponding isolated areas according to the preset strategy;
[0034] S08, perform real-time confusion protection during data transmission, continuously apply dynamic encryption strategy and vector perturbation technology to ensure data security during data transmission from confusion bubbles to target nodes, and monitor data flow in real time to prevent data leakage and illegal sharing.
[0035] Among them, the confusion center node dynamically updates the position and size of the confusion bubble by analyzing historical data access patterns and security threat intelligence, ensures that the confusion protection strategy matches the actual security demand, and the historical data access pattern is derived from the collection result of the data monitoring module in step S01.
[0036] The knowledge vector extraction process converts the original data into a vector representation containing semantic information using semantic encoding techniques, while preserving the core features of the data for subsequent similarity matching and retrieval operations. The tool feature label is used for the generation and distribution determination of the confusion bubble in step S03.
[0037] The confusion bubble adopts an ellipsoidal shape in the high-dimensional vector space, with each confusion bubble center coordinate corresponding to a knowledge cluster center. The lengths of the major and minor axes of the ellipsoid are determined according to the distribution density of the data within the knowledge cluster center. The multi-dimensional authorization strategy is used as input for the calculation of the confusion strength parameter in step S05.
[0038] The area of the intersection overlap region accounts for 15% to 25% of the total area of each confusion bubble. The game boundary model determines the attribution confusion bubble of the data vector in the intersection overlap region by solving the Nash equilibrium. The attribution boundary result is used to establish the data training inference mapping relationship in step S05.
[0039] The data training inference mapping relationship is established based on the life cycle stage of the AI model. The data vector in the training stage is assigned to a high-security level confusion bubble, the data vector in the inference stage is assigned to a standard security level confusion bubble, and the data vector in the verification stage is assigned to a medium security level confusion bubble. The authentication result of the multi-factor authentication technology is used to determine the user permission level value.
[0040] The objective function of the upper game model is to maximize the data protection strength function. The input of the data protection strength function includes sensitivity weight value, access frequency statistical value, user permission level value, threat detection score, and the output is the data protection strength coefficient. The sensitivity weight value is derived from the calculation result of the knowledge vector extraction in step S02.
[0041] The objective function of the lower game model is to minimize the computational resource consumption function. The input of the computational resource consumption function includes processor occupancy, memory usage, network bandwidth demand, and storage space occupancy, and the output is the resource consumption weight coefficient. The processor occupancy and memory usage are derived from the system performance monitoring data of the data monitoring module in step S01.
[0042] The objective functions of the upper game model and the lower game model are coupled through a security efficiency balance coupling term. The security efficiency balance coupling term establishes a constraint condition based on the product relationship between the data protection strength coefficient and the resource consumption weight coefficient. The data protection strength coefficient and the resource consumption weight coefficient are used for the game equilibrium solution value calculation of the confusion strength optimization function in step S05.
[0043] The confusion strength optimization function is used to calculate the confusion strength parameter suitable for the current data sharing scenario according to the solution result of the game boundary model and the data training inference mapping relationship, the input includes the knowledge vector sensitivity weight value, the user permission level value, the data access frequency statistical value, the sharing range identifier code and the game equilibrium solution value, and the output is the confusion strength parameter value between 0 and 1, and the data access frequency statistical value is derived from the access record statistics of the data monitoring module in step S01.
[0044] As shown in Figure 2 The structure of the vector confusion enhancement model is a deep generative model based on a variational autoencoder architecture, including an encoder, a latent space transformation layer and a decoder, the encoder maps the input knowledge vector to the latent space, the latent space transformation layer adjusts the randomness scale parameter of noise addition according to the confusion strength parameter, and the decoder reconstructs the latent vector after confusion processing into an output vector, and the confusion strength parameter is derived from the output result of the confusion strength optimization function in step S05.
[0045] The training data set of the vector confusion enhancement model is established, which specifically includes collecting a large number of multi-domain knowledge vector samples as original training data, labeling the sensitivity level and knowledge category information of each sample, constructing a training sample pair containing the original vector, the target confusion vector and the semantic consistency constraint, and expanding the training set scale through data enhancement technology to improve the generalization ability and robustness of the model, and the knowledge category information is consistent with the classification standard of the tool feature label in step S02.
[0046] The vector confusion enhancement model training step specifically includes simultaneously training the generator and the discriminator network by using the adversarial training strategy, the generator learns to effectively confuse the knowledge vector while maintaining the semantic information, the discriminator learns to distinguish the original vector and the confused vector, and the model parameters are optimized by minimizing the weighted combination of reconstruction loss and semantic preservation loss, and the course learning strategy is adopted in the training process to gradually increase the confusion difficulty to improve the model performance.
[0047] The data protection strength function adopts a nonlinear combination method to process the input parameters, performs exponential operation on the sensitivity weight value and the threat detection score, performs logarithmic operation on the access frequency statistical value and the user permission level value, and then combines the two operation results by weighted average method to obtain the final data protection strength coefficient, and the threat detection score is derived from the security threat intelligence analysis result in step S01.
[0048] The computing resource consumption function establishes a non-linear relationship between input parameters through a polynomial fitting method, the square term of processor occupancy rate is multiplied by the cubic root term of memory usage, network bandwidth demand and storage space occupation are coupled through a trigonometric function relationship, and finally the resource consumption weight coefficient is output. The network bandwidth demand and storage space occupation are derived from the system resource statistics of the real-time monitoring data in step S08.
[0049] The game boundary model solving process adopts a Nash equilibrium algorithm to find the equilibrium solution of the upper game model and the lower game model, and the equilibrium solution is used as the basis for dividing the boundary of the confusion bubble cross-overlapping area. The equilibrium solution is used as the input parameter of the game equilibrium solution value of the confusion strength optimization function in step S05.
[0050] The dynamic encryption strategy ensures that the access rights of users to data are strictly limited according to data sensitivity classification and operation behavior analysis. Strong encryption algorithms and long keys are used for high-sensitivity data. Different keys are generated for different sharing range data. The keys are distributed to legitimate users through a secure key distribution mechanism. The sharing range identifier code is derived from the permission allocation result of the multi-dimensional authorization strategy in step S03, and is used as an input parameter of the confusion strength optimization function in step S05.
[0051] The automatic classification and labeling process automatically assigns data to the corresponding isolated area according to a preset strategy and implements strict access control using AI technology. The preset strategy is determined based on the security level allocation rules of the data training inference mapping relationship in step S05. The access control result of the isolated area is used for permission verification in the data flow monitoring in step S08.
[0052] The game boundary model is a double-layer optimization model for accurately separating the confusion bubble cross-overlapping area. It balances the data security protection demand and system resource limitation to realize dynamic adjustment of the confusion strategy and clear division of the region attribution. The confusion bubble distribution in a high-dimensional vector space is as shown in Figure 3 .
[0053] The tool feature label is identification information describing tool function characteristics, performance indicators, and applicable data types, which is automatically extracted by AI technology analyzing tool historical usage data and user evaluations.
[0054] The multi-dimensional authorization strategy is a comprehensive permission management mechanism based on role permissions, attribute permissions, data sensitivity permissions, and operation behavior permissions, which is used to ensure that the access rights of users to data accurately match the actual needs.
[0055] The data training and inference mapping relationship is an allocation mechanism for associating data vectors in the confusion bubble with AI model life cycle stages, and is used to implement a differentiated security protection strategy based on the differences between the model training stage, the inference stage, and the verification stage.
[0056] The game equilibrium solution value is a numerical representation of the optimal strategy combination obtained by solving the upper game model and the lower game model through the Nash equilibrium algorithm, and is used to guide the boundary division of the confusion bubble cross-overlapping area and the calculation of the confusion strength parameter.
[0057] The security efficiency balance coupling term is a constraint mechanism that connects the objective functions of the upper game model and the lower game model, and realizes dynamic balance optimization of security and efficiency through the product relationship of the data protection strength coefficient and the resource consumption weight coefficient.
[0058] The specific implementation of the above steps is described in detail below.
[0059] The specific implementation of step S01 is to deploy a confusion center node as a core control unit in the distributed architecture of the AI platform management server. The center node adopts a master-slave architecture design, and the master node is responsible for the generation of confusion bubble identifiers and the calculation of the global confusion vector range, and the slave node is responsible for backup and load balancing. The confusion bubble identifier is generated using a 64-bit hash algorithm to ensure the uniqueness of the identifier in a high-concurrency environment. Confusion edge nodes are deployed at each access portal of the platform, and each edge node is deployed using a lightweight containerization method, and the nodes exchange data through an encrypted communication protocol. The data monitoring module uses an event-driven architecture to collect user access records in real time, and the monitoring frequency is set to 100 times per second. When abnormal access behavior is detected, an alarm mechanism is triggered. The operation behavior information is processed in real time by a log analysis engine, and the access pattern data is stored in a time series database to provide a data foundation for subsequent threat intelligence analysis.
[0060] The specific implementation of step S02 is to convert the original data into a 512-dimensional high-dimensional vector representation using a knowledge vector extraction algorithm based on a variational autoencoder. The semantic encoding process uses an attention mechanism to extract key features of the data, and captures complex relationships between data through a multi-head attention network. The vectorization process uses a distributed computing framework to support parallel processing of TB-level data. The tool feature label extraction uses a machine learning classification algorithm to automatically generate a set of labels describing the tool's functionality by analyzing the operation frequency, execution time, error rate, and other indicators in the historical usage data. The user evaluation data is processed by a sentiment analysis algorithm to extract user satisfaction scores as the basis for tool quality assessment. The sensitivity weight value calculation uses the entropy method to determine the importance of each knowledge vector in the data set by calculating its information entropy, and the weight value range is set to 0.1 to 1.0, where 0.1 represents low sensitivity and 1.0 represents high sensitivity.
[0061] The specific implementation of step S03 is to generate confusion bubbles in a high-dimensional vector space using an ellipsoid clustering algorithm, and the geometric shape of each confusion bubble is represented by a multi-dimensional ellipsoid. The center coordinates of the ellipsoid are determined by the K-means clustering algorithm, and the number of clusters is adaptively adjusted according to the data distribution density, generally set to the square root of the total amount of data. The length of the long and short axes of the ellipsoid is determined by the eigenvalue decomposition of the covariance matrix, and the ratio of the long and short axes is controlled between 1.5 and 3.0 to ensure the confusion effect. The overlapping area of adjacent confusion bubbles is set to 20%, and the precise boundary of the overlapping area is calculated by a spatial geometry algorithm. The multi-dimensional authorization strategy uses an attribute-based access control model, role permissions are managed through a role inheritance tree structure, attribute permissions use a multi-value attribute matching algorithm, data sensitivity permissions are classified according to the weight values of step S02, and operation behavior permissions are dynamically adjusted through a behavior pattern recognition algorithm.
[0062] The specific implementation of step S04 is to build a double-layer game boundary model to handle the overlapping area between confusion bubbles. The upper game model uses non-cooperative game theory with the objective function of maximizing data protection strength, and the game participants are the confusion bubbles, and the strategy space is the ownership selection of the overlapping area. The lower game model aims to minimize the consumption of computing resources, and solves the optimal solution of resource allocation through linear programming method. The game solution uses the Nash equilibrium algorithm to find a strategy combination that makes all participants have no improvement motivation through iterative calculation. The convergence judgment standard of the equilibrium solution is set to the strategy change amplitude of 10 consecutive iterations being less than 0.01. The ownership boundary of the overlapping area is determined by the weighted Voronoi diagram algorithm, and the weight coefficient is provided by the game equilibrium solution to ensure the optimality and stability of the boundary division.
[0063] The specific implementation of step S05 is to establish a data training inference mapping relationship according to the life cycle stage of the AI model. Training stage data is allocated to a high-security obfuscation bubble with a security level of 5, inference stage data is allocated to a standard security obfuscation bubble with a security level of 3, and verification stage data is allocated to a medium security obfuscation bubble with a security level of 4. The multi-factor authentication adopts a parallel verification mechanism, the username and password authentication is verified by a salted hash algorithm, the fingerprint recognition adopts a minutia matching algorithm, the matching threshold is set to 85%, the face recognition adopts a deep neural network model, the recognition accuracy is required to be more than 95%, the SMS verification code adopts a timestamp verification mechanism, and the effective period is set to 300s. The obfuscation strength optimization function adopts a multi-objective optimization algorithm, the input parameters include the knowledge vector sensitivity weight value, the user permission level value, the data access frequency statistical value, the shared range identifier code and the game equilibrium solution value, the obfuscation strength parameter is calculated by weighted summation, and the parameter value range is limited between 0 and 1.
[0064] The specific implementation of step S06 is to deploy a vector obfuscation enhancement model based on a variational autoencoder architecture to perform differential privacy processing. The encoder adopts a multi-layer perceptron structure, containing 3 hidden layers, the number of neurons in each layer is 256, 128 and 64 respectively, and the activation function adopts a ReLU function. The latent space transformation layer dynamically adjusts the variance of Gaussian noise according to the obfuscation strength parameter, and the noise intensity is positively correlated with the obfuscation parameter. The decoder structure is symmetrical with the encoder, and the stability of gradient propagation is maintained through residual connection. The privacy budget ε value of differential privacy is dynamically set according to the data sensitivity, the ε value of high sensitivity data is set to 0.1, the ε value of medium sensitivity data is set to 0.5, and the ε value of low sensitivity data is set to 1.0. The dynamic encryption strategy adopts a hybrid encryption system, AES-256 algorithm is adopted for high sensitivity data, the key length is 256 bits, AES-128 algorithm is adopted for standard sensitivity data, the key management adopts a hierarchical key structure, and the master key is protected by a hardware security module.
[0065] The specific implementation of step S07 is to establish a multi-level authentication mechanism based on biometric features and behavior patterns. Biometric recognition uses multi-modal fusion technology, combining fingerprint, face, and voiceprint biometrics, to improve recognition accuracy through a weighted fusion algorithm. Behavior pattern analysis uses machine learning anomaly detection algorithms to establish user profiles by analyzing user login times, operation frequencies, access paths, and other behavior characteristics. Dynamic permission verification uses a time and geographic location-based access control strategy to dynamically adjust permission levels based on the user's current location and access time. Distributed verification is achieved through obfuscated edge nodes, and Byzantine fault-tolerant algorithms are used between nodes to ensure consistency of verification results, with a fault tolerance threshold set to one-third of the total number of nodes. Data automatic classification uses a deep learning classification model to extract data features through a convolutional neural network, with a classification accuracy requirement of over 90%. Isolation areas are divided into five levels according to security levels, with different access control strategies and encryption strengths used for different levels.
[0066] The specific implementation of step S08 is to implement a real-time obfuscation protection mechanism during data transmission. Data transmission uses an end-to-end encryption protocol, with dynamic encryption strategies applied continuously during transmission, and encryption keys updated every 1MB of data transmitted. Vector perturbation technology uses the Laplace mechanism to add noise, with noise scale parameters dynamically adjusted according to the security level of the transmission path. Real-time monitoring uses a streaming processing framework to analyze data streams in real time, with monitoring indicators including transmission rate, data integrity, access permission compliance, and others. Data leakage detection uses machine learning-based anomaly detection algorithms to identify potential leakage behavior by analyzing statistical characteristics of data streams, with a detection accuracy requirement of over 95% and a false positive rate controlled below 5%. Illegal sharing protection is achieved through digital watermarking technology, embedding invisible identification information in the data, with watermark strength adjusted according to data importance, with a watermark redundancy of 15% for important data and 8% for general data. When violations are detected, the system automatically triggers an emergency response mechanism, including access blocking, log recording, administrator alerts, and other measures.
[0067] The key technical ideas of the present application mainly reflect in the following aspects. First, the synergy mechanism of confusion bubble and game boundary model, by constructing an ellipsoidal confusion bubble in a high-dimensional vector space, combining a double-layer game model to optimize the boundary division of the overlapping area, compared with the traditional static partition method, this mechanism can dynamically adjust the protection boundary according to the data distribution characteristics and security requirements, maximize the resource utilization efficiency on the premise of ensuring data security. Secondly, the vector confusion enhancement technology based on variational autoencoder, through the deep generative model to maintain the semantic information of the data while realizing effective confusion, compared with the traditional noise adding method, this technology can provide stronger privacy protection without destroying the data usability, especially in processing high-dimensional knowledge vectors, it shows better confusion effect and semantic preservation ability. Third, the fusion security mechanism of multi-dimensional authorization and multi-level authentication, by combining role permission, attribute permission, data sensitivity permission and operation behavior permission to build a comprehensive permission management system, cooperating with multi-level authentication of biometric recognition and behavior pattern analysis, compared with single access control method, this mechanism can provide more fine-grained and more dynamic permission management, effectively prevent internal threats and abuse of permissions. The synergistic effect of these key technical ideas forms a complete data security sharing solution, through the dynamic protection of confusion bubble, the privacy processing of vector enhancement and the access control of multi-dimensional authentication, a full life cycle security protection system from data storage, processing to transmission is constructed, compared with the existing single-point protection or static protection method, this synergistic mechanism can provide more comprehensive, more intelligent and more efficient data security protection in the complex AI platform environment, especially in processing large-scale knowledge vector data, it shows significant security and practicality advantages.
[0068] It should be noted that the present application also solves the technical problems that the present application solves the technical problems of lack of real-time monitoring and dynamic response capability of data access behavior in traditional data sharing system. In the existing data sharing platform, most systems can only provide post-audit and static monitoring function, and cannot detect and respond to abnormal access behavior in real time, which leads to the fact that data security threats cannot be discovered and handled in time. The present application collects data access records and operation behavior information in real time through the deployment of a data monitoring module, analyzes user behavior patterns in combination with AI technology, can identify abnormal access patterns in time and trigger corresponding security response mechanism, and significantly improves the real-time security protection of the data sharing system. The present application also solves the technical problems of semantic information loss and difficulty in balancing privacy protection strength in traditional vectorization data protection method. The existing data vectorization processing technology often adopts simple noise addition or data perturbation method, which can provide a certain degree of privacy protection, but often leads to serious loss of semantic information of data, affecting the subsequent data utilization effect. The present application adopts an adversarial training strategy through the vector confusion enhancement model, the generator effectively confuses the knowledge vector on the premise of maintaining semantic information, the discriminator is responsible for distinguishing the original vector and the confused vector, and the model parameters are optimized by minimizing the weighted combination of reconstruction loss and semantic preservation loss, so that the availability and semantic integrity of data are maximized while providing strong privacy protection.
[0069] Specifically, the principle of the present application is that the core principle of the technical solution of the present application to solve the problem of mismatch between confusion protection strategy and actual security demand is to build a dynamic adaptive confusion protection mechanism based on game theory. First, by converting the data to be shared into a vector representation in a high-dimensional knowledge vector space, the semantic features and sensitivity weights of the data can be quantitatively described, providing a mathematical basis for subsequent precise protection. Secondly, the ellipsoidal confusion bubble generated in the vector space is not statically distributed, but is dynamically adjusted according to the historical data access patterns and security threat intelligence collected by the confusion center node, ensuring that the confusion region covers the actual data distribution characteristics. Thirdly, the game boundary model maximizes the data protection strength by the upper game model and minimizes the computational resource consumption by the lower game model, uses the Nash equilibrium algorithm to solve the optimal strategy combination considering safety and efficiency, and avoids the problem of over-protection or insufficient protection in traditional methods. Finally, the vector confusion enhancement model is based on the variational autoencoder architecture, dynamically adjusts the temperature parameter of the noise addition mechanism according to the game solution result, realizes the real-time matching of confusion strength and current security demand. The whole technical scheme forms a closed loop control through the real-time feedback of the data monitoring module, so that the confusion protection strategy can continuously adapt to environmental changes, thereby completely solving the technical problem of mismatch between strategy and demand.
[0070] A specific embodiment 1 of the present application is provided below, and the specific implementation of each step in embodiment 1 is described in detail as follows.
[0071] In this embodiment, the specific implementation of step S01 is the same as described above, and will not be described in detail here.
[0072] The specific implementation of step S02 is to convert data into a vector representation in a high-dimensional knowledge vector space using vectorization processing. The mathematical expression of the knowledge vector extraction process is as follows:
[0073] V i = f enc (D i , θ enc );
[0074] In the formula, V i is the knowledge vector representation of the i th data sample; D i is the i th original data sample; f enc is the encoding function; and θ enc is the encoder parameter. The sensitivity weight value is calculated using the information entropy method, and the specific expression is as follows:
[0075]
[0076] In the formula, w is the sensitivity weight value of the i th knowledge vector; p ij is the probability distribution of the i th vector in the j th feature dimension; n is the number of vector dimensions; and α is the context weight coefficient, which is in the range of 0.1 to 0.3. is the context sensitivity evaluation value of the i th vector. The parameter acquisition method is as follows: p ij is obtained through normalization processing, and the calculation process is as follows: where V ij is the component value of the i th knowledge vector in the j th dimension, and V ik is the component value of the i th knowledge vector in the k th dimension. The range is 0.05 to 0.95, which is determined by analyzing the importance of data in the business scenario.
[0077] The specific implementation of step S03 is to generate an ellipsoidal confusion bubble based on the distribution characteristics of the knowledge vector. The mathematical expression for calculating the ellipsoidal parameters is as follows:
[0078]
[0079] In the formula, E k is the ellipsoidal region of the k th confusion bubble; μ k is the coordinate vector of the k th cluster center; and Σ kcovariance matrix of the kth cluster; r k radius parameter of the kth ellipsoid; n is the dimension of the vector space. The calculation expression of the cross-overlapping area is:
[0080]
[0081] wherein, is the overlapping area of the kth and the lth confusion bubbles; E k ∩E l denotes the intersection area of two ellipsoids. The parameter acquisition method is: μ k is calculated by the K-means clustering algorithm; ∑ k is obtained by sample covariance matrix estimation, and the calculation formula is wherein N k is the number of samples in the kth cluster; r k is determined according to the sample distribution density within the cluster, and the range is 1.5 to 3.0.
[0082] The specific implementation of step S04 is to process the cross-overlapping area by using a double-layer game boundary model. The objective function expression of the upper game model is:
[0083]
[0084] wherein, P strength is a data protection strength function; s k is the strategy selection of the kth confusion bubble; is a sensitivity weight value; is an access frequency statistical value; is a user permission level value; is a threat detection score; β1, β2, β3, β4 are undetermined coefficients; ∈1 is an error term, and the range is 0.01 to 0.05. The objective function expression of the lower game model is:
[0085]
[0086] wherein, C resource is a computing resource consumption function; a k is the resource allocation strategy of the kth confusion bubble; CPU rate is a processor occupancy rate; MEM usage is a memory usage amount; NET demand is a network bandwidth demand; STOR occupy is a storage space occupancy amount; γ1, γ2 are coupling coefficients; ∈2 is an error term, and the range is 0.02 to 0.08. The Nash equilibrium solving expression of the game equilibrium solution is:
[0087]
[0088] In the formula, s -k represents the strategy combination of other bubbles except the kth bubble; a -k represents the resource allocation strategy of other bubbles except the kth bubble; φ k is the trade-off coefficient of the kth bubble, with a value range of 0.3 to 0.7. The expression of the data training inference mapping relationship is:
[0089]
[0090] In the formula, is the security level mapping value of the ith data vector; Stage i is the AI model life cycle stage identifier corresponding to the ith data vector. The parameter acquisition method is: β1, β2, β3, β4 are determined by historical data regression fitting, with a value range of 0.2 to 0.8, 0.1 to 0.5, 0.3 to 0.7, and 0.15 to 0.45, respectively; γ1, γ2 are determined by system performance test, with a value range of 0.5 to 2.0; φ k According to the importance of each confusion bubble.
[0091] The specific implementation of step S05 is to calculate the confusion strength parameter, and the mathematical expression of the confusion strength optimization function is:
[0092]
[0093] In the formula, I confuse is the confusion strength parameter value, with a value range of 0 to 1; ω1, ω2, ω3, ω4, ω5 are weight coefficients; R share is the shared range identifier encoding; G nash is the game equilibrium solution value; ζ is the adjustment term, ranging from -0.1 to 0.1. The normalization constraint condition of the weight coefficient is:
[0094]
[0095] The parameter acquisition method is: ω1, ω2, ω3, ω4, ω5 are determined by a multi-objective optimization algorithm, with initial values of 0.3, 0.25, 0.2, 0.15, and 0.1, respectively; G nash The double-layer game model is solved by Nash equilibrium algorithm.
[0096] The specific implementation of step S06 is to deploy the vector confusion enhancement model to perform differential privacy processing, and the temperature parameter adjustment expression of the noise addition mechanism is:
[0097] T noise =T base ·(1+δ·I confuse )λ ;
[0098] wherein T noise is the adjusted noise temperature parameter; T base is the base temperature parameter, taking a value of 1.0; δ is the confusion strength influence factor, taking a value range of 0.5 to 2.0; λ is the exponential adjustment parameter, taking a value range of 0.8 to 1.2. The noise addition expression of differential privacy is:
[0099]
[0100] wherein is the knowledge vector after adding noise; N(0, σ 2 ·T noise ) is Gaussian noise with a mean of 0 and a variance of σ 2 ·T noise ; σ is the noise scale parameter, determined according to the privacy budget ε dp , and the calculation formula is wherein ε dp is the privacy budget parameter of differential privacy, and δ dp is the differential privacy parameter, taking a value of 10 -5 .
[0101] The specific implementation of steps S07-S08 is the same as the foregoing, and will not be described in detail here.
[0102] It should be explained that the mathematical expression of the security-efficiency balance coupling term is:
[0103] C balance = P strength ·C resource + η·|P strength -P target | + ξ·|C resource -C limit |;
[0104] wherein C balance is the security-efficiency balance coupling term; P target is the target protection strength; C limit is the resource consumption limit; η and ξ are balance adjustment coefficients, taking value ranges of 0.1 to 0.5 and 0.2 to 0.8, respectively.
[0105] The sensitivity weight value calculation formula is based on the information entropy theory, and realizes sensitivity evaluation by quantifying the information content of data in a high-dimensional vector space. The formula introduces a context sensitivity evaluation term The formula can consider the internal characteristics of data and external environmental factors, and compared with the traditional single weight allocation method, the formula can more accurately identify key sensitive data and improve the pertinence and effectiveness of data protection. The ellipsoid confusion bubble generation formula adopts a multi-dimensional ellipsoid geometric model, and through the covariance matrix ∑ k The formula describes the data distribution characteristics, and can dynamically adjust the protection boundary according to the actual distribution of data. Compared with the fixed geometric shape of the protection area, the ellipsoid model can better fit the natural clustering characteristics of data, minimize the protection resource consumption while ensuring the coverage rate. The objective function of the double-layer game model adopts a nonlinear combination form, and the upper model realizes the nonlinear optimization of security strength through the combination of exponential function and logarithmic function ;
[0106]
[0107] In the formula, P upper is the objective function value of the upper game model, representing the comprehensive evaluation result of data protection strength. The lower model realizes accurate modeling of resource consumption through the coupling of polynomial term and trigonometric function term sin(γ1NET demand )cos(γ2STOR occupy );
[0108]
[0109] In the formula, C lower is the objective function value of the lower game model, representing the quantitative evaluation result of computing resource consumption.
[0110] The double-layer structure can find the optimal balance point in the complex multi-objective optimization environment, and compared with the single-layer optimization model, the double-layer game can better handle the conflict between security and efficiency. The Nash equilibrium solving formula finds the stable strategy combination by maximizing the difference P strength (s k , s -k )-φ k ·C resource (a k , a -k ) of the utility functions of each participant, which ensures that any participant cannot obtain greater benefits by changing the strategy unilaterally. Compared with the simple optimization algorithm, the Nash equilibrium can guarantee the stability and fairness of the solution. The data training and inference mapping relationship formula adopts a piecewise function form According to the AI model life cycle stage, different security levels are allocated, the mapping mechanism can realize differentiated security protection strategy, compared with uniform security level, the stage mapping can guarantee the safety of the key stage while improving the overall efficiency of the system. The confusion strength optimization function adopts a multi-parameter linear weighted model, by introducing a logarithmic transformation log(R share +1) processing shared range coding, it can effectively handle the balance problem between parameters of different orders of magnitude. The normalization constraint condition of the function ensures the rationality and stability of the parameters. Compared with simple linear combination, the optimization function can more accurately reflect the contribution of each influencing factor to the confusion strength.
[0111] The noise temperature parameter adjustment formula adopts an exponential growth model T noise =T base ·(1+δ·I confuse ) λ , by dynamically adjusting the randomness of noise addition through the confusion strength parameter, the formula can adaptively adjust the privacy protection strength according to the security requirements of data. Compared with the differential privacy method with fixed noise parameter, the dynamic adjustment mechanism can maximize the availability of data while protecting privacy. The security efficiency balance coupling term is realized through the combination of product relationship P strength ·C resource and deviation penalty term η·|P strength -P target |+ξ·|C resource -C limit |, which realizes the dynamic balance optimization between security strength and resource consumption. The coupling mechanism can prevent the over-optimization of a single target, ensuring that the system maintains reasonable resource efficiency while pursuing high security. Compared with independent security and efficiency optimization, coupled optimization can obtain more stable and practical system performance.
[0112] In this embodiment, the vector confusion enhancement model adopts a deep generative model design based on variational autoencoder (VAE) architecture, which includes three core components. The encoder part adopts a multi-layer perceptron structure, containing 3 hidden layers with neuron numbers of 256, 128 and 64 respectively, using ReLU activation function, responsible for mapping the input 512-dimensional knowledge vector to a low-dimensional latent space. The latent space transformation layer is the key innovative part of the model, which dynamically adjusts the variance of Gaussian noise according to the confusion strength parameter. The noise intensity is positively related to the confusion parameter, realizing differential privacy protection. The decoder structure is symmetrical with the encoder, maintaining the stability of gradient propagation through residual connection, reconstructing the latent vector after confusion processing into the output vector, ensuring the semantic consistency of data while adding noise.
[0113] The establishment of the training data set follows the multi-domain knowledge vector collection principle. First, a large number of multi-domain knowledge vector samples are collected as original training data, covering different data types such as text, image, audio, etc. Each sample is accurately labeled, including sensitivity level classification (high, medium, and low three levels) and knowledge category information, which is consistent with the tool feature label classification standard. When constructing the training sample pair, each sample contains three elements: original vector, target confusion vector, and semantic consistency constraint, to ensure that the confused vector can protect privacy and not lose the original semantic information. Through data augmentation techniques, more samples are generated by methods such as rotation, scaling, and noise addition to improve the generalization ability and robustness of the model. Finally, the training set size reaches the million level.
[0114] The model training adopts an adversarial training strategy, training the generator and discriminator networks simultaneously. The generator learns to effectively confuse the knowledge vector while maintaining semantic information, and the discriminator learns to distinguish between the original vector and the confused vector, forming a dynamic game process. The training objective function optimizes the model parameters by minimizing the weighted combination of reconstruction loss and semantic preservation loss, where the reconstruction loss ensures the structural similarity between the output vector and the input vector, and the semantic preservation loss ensures the availability of the confused data. The training process adopts a curriculum learning strategy, starting with simple low-dimensional vector confusion and gradually increasing to high-dimensional complex vectors, while gradually increasing the confusion difficulty and noise intensity. After each training epoch, the model performance on the validation set is evaluated, including confusion effect, semantic preservation degree, and computational efficiency. When the validation loss does not decrease for 5 consecutive epochs, the training is stopped to ensure that the model reaches the optimal performance state.
[0115] To better understand and implement the present application, the following provides an embodiment 2 of a specific application scenario of the present application: A technical team is commissioned to build an AI-based data security sharing system for a large e-commerce platform. The platform processes about 5 million commodity data, about 20 million user behavior data, and about 1 million transaction data per day. The platform needs to realize data sharing of AI models such as commodity recommendation, user portrait analysis, and fraud detection under the premise of ensuring data security. The technical team adopts the data security sharing method of the present application to build a confusion protection system covering the entire platform.
[0116] The technical team first deployed the confusion center node in the distributed service architecture of the e-commerce platform. The node adopts a master-slave architecture design, with the master node deployed in the core data center, responsible for generating confusion bubble identifiers and global confusion vector range calculation. The confusion bubble identifier is generated by the SHA-256 hashing algorithm to generate a 64-bit unique identifier, ensuring the uniqueness of the identifier in a 100,000 concurrent access environment per second. Confusion edge nodes are deployed at the entrances of 12 main business modules such as product management, user management, order management, and payment management on the platform. Each edge node is deployed using the Docker containerization method, and nodes exchange data through the TLS1.3 encryption communication protocol. The data monitoring module uses the Apache Kafka event-driven architecture to collect user access records in real time, with a monitoring frequency of 100 times per second. When detecting abnormal access behavior such as accessing more than 1000 product detail pages within 5 minutes for a single user, the alarm mechanism is triggered. Operation behavior information is processed in real time by the ELK log analysis engine, and access pattern data is stored in the InfluxDB time series database to provide a data foundation for subsequent threat intelligence analysis.
[0117] The technical team extracted knowledge vectors from the e-commerce platform's product data, user behavior data, and transaction data. A vector extraction model based on the Transformer architecture was used to convert raw data such as product descriptions, user reviews, and transaction records into 512-dimensional high-dimensional vector representations. The semantic encoding process uses a multi-head attention mechanism to extract key features from the data, capturing complex relationships between product attributes, user preferences, and transaction patterns through 8-head attention networks. Vectorization processing uses the Apache Spark distributed computing framework, supporting TB-level parallel processing with a single processing capacity of 5 million product data. Tool feature label extraction uses a random forest classification algorithm to automatically generate a label set describing tool functions by analyzing historical usage data from AI tools such as recommendation systems, search engines, and risk control systems, including call frequency, response time, accuracy, and other indicators. 15 categories of labels such as "product recommendation", "user portrait", and "fraud detection". User review data is processed by the BERT sentiment analysis model to extract user satisfaction scores as a basis for tool quality evaluation, with a score range of 1 to 5. Sensitivity weight value calculation uses the information entropy method to determine the importance of each knowledge vector in the data set, with a product price information weight value of 0.85, a user personal information weight value of 0.92, and a transaction record weight value of 0.78.
[0118] The technical team adopts an ellipsoid clustering algorithm to generate confusion bubbles in a 512-dimensional vector space, each confusion bubble adopting a multi-dimensional ellipsoid geometric shape. The center coordinates of the ellipsoid are determined by improving the K-means clustering algorithm, and the number of clusters is adaptively adjusted according to the data distribution density, finally determined as 2236 confusion bubbles. The lengths of the major and minor axes of the ellipsoid are determined by eigenvalue decomposition of the covariance matrix, and the length ratio of the major and minor axes of the commodity data confusion bubble is 2.3, the length ratio of the major and minor axes of the user behavior data confusion bubble is 1.8, and the length ratio of the major and minor axes of the transaction data confusion bubble is 2.1. The cross-overlapping area of adjacent confusion bubbles accounts for 22%, and the precise boundary of the overlapping area is calculated by a spatial geometry algorithm. The multi-dimensional authorization strategy adopts an attribute-based access control model, and the role permissions include data analyst, algorithm engineer, business operation, system administrator, etc. 4 main roles, and the permission relationship is managed through a role inheritance tree structure. The attribute permission adopts a multi-value attribute matching algorithm, including 12 attribute dimensions such as department attribute, project attribute, and data type attribute. The data sensitivity permission is classified into 5 levels according to the aforementioned weight value, and the operation behavior permission is dynamically adjusted through a behavior pattern recognition algorithm, identifying 8 types of operations such as reading, writing, deleting, and sharing.
[0119] The technical team constructs a double-layer game boundary model to handle the cross-overlapping area between the 2236 confusion bubbles. The upper game model adopts a non-cooperative game theory, with the maximum data protection strength as the objective function, and the game participants are the confusion bubbles, and the strategy space is the ownership selection of the overlapping area. The lower game model minimizes the consumption of computing resources as the target, and solves the optimal solution of resource allocation through linear programming method. The game solution adopts Nash equilibrium algorithm, and finds the equilibrium solution through iterative calculation, sets the maximum iteration number to 1000 times, and the convergence judgment standard is that the strategy change amplitude of continuous 10 times iteration is less than 0.01. After 673 iterations, the convergence is reached, and the ownership boundary of the cross-overlapping area is determined by the weighted Voronoi diagram algorithm, and the weight coefficient is provided by the game equilibrium solution, ensuring the optimality and stability of the boundary division.
[0120] The technical team establishes a data training inference mapping relationship according to the life cycle stage of the AI model. The training stage data of the commodity recommendation model is allocated to a high security obfuscation bubble with a security level of 5, containing 256 bubbles, the inference stage data is allocated to a standard security obfuscation bubble with a security level of 3, containing 892 bubbles, and the verification stage data is allocated to a medium security obfuscation bubble with a security level of 4, containing 445 bubbles. The user portrait analysis model is allocated to the corresponding security level obfuscation bubble according to the same rule. The multi-factor authentication adopts a parallel verification mechanism, the username and password authentication is verified by the bcrypt salted hash algorithm, the fingerprint recognition adopts the minutia matching algorithm, the matching threshold is set to 85%, the face recognition adopts the FaceNet deep neural network model, the recognition accuracy reaches 97.3%, the SMS verification code adopts the TOTP timestamp verification mechanism, and the effective period is set to 300s. The confusion strength optimization function adopts a genetic algorithm for multi-objective optimization, the input parameters include knowledge vector sensitivity weight value, user permission level value, data access frequency statistical value, shared range identifier code and game equilibrium solution value, the confusion strength parameter is calculated by weighted summation. The confusion strength parameter of commodity data is 0.73, the user behavior data is 0.86, and the transaction data is 0.91.
[0121] As shown in Figure 2 The technical team deploys a vector obfuscation enhancement model based on a variational autoencoder architecture to perform differential privacy processing. The encoder adopts a multi-layer perceptron structure, containing 3 hidden layers with neuron counts of 256, 128 and 64 respectively, and the activation function adopts the ReLU function. The latent space transformation layer dynamically adjusts the standard deviation of Gaussian noise according to the confusion strength parameter, the noise standard deviation of commodity data is 0.15, the user behavior data is 0.22, and the transaction data is 0.28. The decoder structure is symmetrical with the encoder, and the residual connection is used to maintain the stability of gradient propagation. The privacy budget ε value of differential privacy is dynamically set according to the data sensitivity, the transaction data with high sensitivity sets ε value to 0.1, the user behavior data with medium sensitivity sets to 0.5, and the commodity data with low sensitivity sets to 1.0. The dynamic encryption strategy adopts a hybrid encryption system, AES-256 algorithm is used for transaction data with a key length of 256 bits, AES-192 algorithm is used for user behavior data, and AES-128 algorithm is used for commodity data. The key management adopts a three-layer key structure, and the master key is protected by the Huawei Kunpeng 920 hardware security module.
[0122] The technology team establishes a multi-level authentication mechanism based on biometric and behavioral patterns. Biometric recognition uses multi-modal fusion technology, combining fingerprint, face, and voiceprint biometrics, and improves recognition accuracy to 99.1% through a weighted fusion algorithm. Behavioral pattern analysis uses the Isolation Forest anomaly detection algorithm to analyze 17 behavioral characteristics such as login time, operation frequency, and access path to establish a user profile, with an anomaly detection accuracy of 94.7%. Dynamic permission verification uses a time and geographic location-based access control strategy, with permission levels increasing by one level during working hours (9:00-18:00) and decreasing by one level during non-working hours, and additional identity verification required for out-of-town access. Distributed verification is achieved through 12 confusion edge nodes, with PBFT Byzantine fault tolerance algorithm used between nodes to ensure consistency of verification results, with a fault tolerance threshold of 4 nodes. Data automatic classification uses a ResNet-50 deep learning classification model to extract data features through a convolutional neural network, with a classification accuracy of 93.8%. Isolation areas are divided into five levels according to security levels, with different access control strategies and encryption strengths, as shown in Table 1.
[0123] Table 1 Data Security Level Classification Table
[0124] Security Level Data Type Encryption Algorithm Access Rights Logging Frequency 5 Transaction Record AES-256 Admin + Auditor Real-time 4 User Personal Information AES-192 Admin + Analyst Every Minute 3 User Behavior Data AES-128 All Roles Every 5 Minutes 2 Product Attributes AES-128 Operations + Analyst Every 10 Minutes 1 Product Description DES All Roles Every Hour
[0125] The technology team implements a real-time confusion protection mechanism during data transmission. Data transmission uses the TLS1.3 end-to-end encryption protocol, with dynamic encryption strategies applied continuously during transmission, with encryption keys updated every 1MB of data transmitted, with 150,000 key updates per day. Vector perturbation technology uses the Laplace mechanism to add noise, with noise scale parameters dynamically adjusted according to the security level of the transmission path, with a noise scale of 0.05 for internal network transmission, 0.12 for external network transmission, and 0.18 for cross-regional transmission. Real-time monitoring uses the Apache Storm streaming processing framework to analyze data streams in real time, with monitoring indicators including transmission rate, data integrity, access permission compliance, and other indicators. Data leakage detection uses an anomaly detection algorithm based on the LSTM neural network to identify potential leakage behavior by analyzing the statistical characteristics of data streams, with a detection accuracy of 96.2% and a false positive rate of 3.8%. Illegal sharing protection is achieved through digital watermarking technology, which embeds invisible identification information in the data, with watermark strength adjusted according to data importance, with a watermark redundancy of 15% for transaction data, 10% for user behavior data, and 8% for product data. When illegal behavior is detected, the system automatically triggers an emergency response mechanism, including access blocking, log recording, administrator alerts, and other measures, with a response time of less than 5 seconds.
[0126] As Figure 4As shown, the performance monitoring data during system operation shows that the CPU usage of the confusion center node is maintained at about 65%, the memory usage is 78%, and the peak network bandwidth occupancy is 2.3 Gbps. The average response time of the confusion edge node is 12 ms, and the single-node processing capacity reaches 80,000 requests per second. The training time of the vector confusion enhancement model is 72 hours, the model size is 256 MB, and the inference delay is 8.5 ms. The impact of differential privacy processing on data accuracy is controlled within 5%, meeting the business requirements. The average verification time of the multi-factor authentication is 3.2 s, and the user experience is good. The encryption and decryption overhead of data transmission accounts for 8% of the total transmission time, and the overall performance loss of the system is controlled within 15%.
[0127] This embodiment successfully solves the problem of data security sharing in the AI model training and inference process of e-commerce platforms. After 6 months of stable operation, the system has processed more than 1 billion data records for security sharing, and no data leakage events have occurred. The training effect of the commodity recommendation model is compared with the traditional method, and the recommendation accuracy only decreases by 1.2%, but the data security is significantly improved. The user portrait analysis model still maintains good analysis accuracy under the protection of confusion, and the user clustering accuracy reaches 91.5%. The fraud detection model is trained through confused transaction data, and the detection accuracy reaches 94.3%, with a false positive rate of 5.7%.
[0128] The present application brings significant technical progress compared to traditional data security sharing methods. Traditional methods usually use simple data desensitization or access control mechanisms, lack protection of data semantic information, and are vulnerable to inference of original data through correlation analysis and other means. The present application constructs a multi-layer protection system in a high-dimensional vector space through the confusion bubble technology, not only protecting the direct information of the data, but also protecting the association relationship and distribution characteristics between the data. The introduction of the game boundary model realizes the dynamic balance between security and efficiency, avoiding the problem of excessive security measures leading to system performance degradation in traditional methods. The vector confusion enhancement model based on deep learning technology can achieve accurate privacy protection while maintaining data usability, which is not achievable by traditional static confusion methods. The multi-dimensional authorization strategy combined with behavior pattern analysis realizes fine-grained permission control, which is more flexible and secure than traditional role-based access control. The real-time confusion protection mechanism ensures the security of data throughout its life cycle, while traditional methods often only focus on protecting data during the storage stage, ignoring the security risks during transmission and use.
[0129] It should be noted that the variables involved in the present application are explained in detail as shown in Table 2.
[0130] Table 2 Variable Explanation Table
[0131]
[0132] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. A method for secure data sharing under an AI platform, characterized in that, This includes establishing a central obfuscation node and multiple obfuscation edge nodes in the AI platform management server, deploying a data monitoring module to collect data access records and operation behavior information in real time; extracting knowledge vectors from the data to be shared, converting the data into a vector representation in a high-dimensional knowledge vector space through vectorization processing, automatically extracting tool feature labels, and calculating the sensitivity weight value of each knowledge vector; generating multiple obfuscation bubbles in the vector space based on the distribution characteristics of knowledge vectors, and constructing a multi-dimensional authorization strategy. A game-theoretic boundary model is used to separate the overlapping regions of confused bubbles, and the ownership boundary of the overlapping regions is determined by solving the game. A data training inference mapping relationship is established based on the separation results of confused bubbles, and a multi-factor authentication technique is used to calculate the confusion intensity parameter through a confusion intensity optimization function. The system initiates a vector obfuscation enhancement model to perform differential privacy processing on the data vectors within the obfuscation bubble, achieving dynamic data encryption; it establishes a multi-level authentication mechanism to perform distributed verification of user identities through obfuscated edge nodes, and uses AI technology to automatically classify and label data and automatically allocate data to corresponding isolated areas according to preset strategies; it performs real-time obfuscation protection during data transmission and monitors data flow in real time to prevent data leakage and illegal sharing.
2. The data security sharing method under the AI platform according to claim 1, characterized in that, The obfuscation center node is responsible for generating obfuscation bubble identifiers and obfuscation vector ranges. It dynamically updates the position and size of obfuscation bubbles by analyzing historical data access patterns and security threat intelligence to ensure that the obfuscation protection strategy matches the actual security needs. The obfuscation edge nodes are distributed at various access points of the platform and are responsible for performing local obfuscation operations.
3. The data security sharing method under the AI platform according to claim 2, characterized in that, The knowledge vector extraction process uses semantic encoding technology to convert the original data into a vector representation containing semantic information, while retaining the core features of the data for subsequent similarity matching and retrieval operations. AI technology is used to analyze the historical usage data and user reviews of the tool to automatically extract tool feature tags.
4. The data security sharing method under the AI platform according to claim 3, characterized in that, The confusion bubbles are distributed in a high-dimensional vector space in an ellipsoidal shape. The vector range of each confusion bubble covers the corresponding knowledge vector region. Adjacent confusion bubbles form overlapping regions. The center coordinates of each confusion bubble correspond to a knowledge cluster center. The lengths of the major and minor axes of the ellipsoid are determined according to the distribution density of the data within the knowledge cluster center.
5. The data security sharing method under the AI platform according to claim 4, characterized in that, The multi-dimensional authorization strategy is constructed based on role permissions, attribute permissions, data sensitivity permissions, and operation behavior permissions. The area of the overlapping region accounts for 15% to 25% of the total area of each confusion bubble. The game boundary model determines the belonging of the data vector in the overlapping region to the confusion bubble by solving the Nash equilibrium.
6. The data security sharing method under the AI platform according to claim 5, characterized in that, The game boundary model includes an upper-level game model that aims to maximize data protection strength and a lower-level game model that aims to minimize computational resource consumption. The data training and inference mapping relationship is established based on the life cycle stages of the AI model, and the data vector in each confusion bubble is associated with the corresponding AI model training and inference stages.
7. The data security sharing method under the AI platform according to claim 6, characterized in that, During the training phase, data vectors are assigned to high-security-level obfuscation bubbles; during the inference phase, data vectors are assigned to standard-security-level obfuscation bubbles; and during the verification phase, data vectors are assigned to medium-security-level obfuscation bubbles. Multi-factor authentication technology combines username / password authentication, fingerprint authentication, facial recognition authentication, and SMS verification code authentication.
8. The data security sharing method under the AI platform according to claim 7, characterized in that, The objective function of the upper-level game model is to maximize the data protection strength function. The inputs of the data protection strength function include sensitivity weight values, access frequency statistics, user permission level values, and threat detection scores. The output is the data protection strength coefficient. The sensitivity weight values are derived from the calculation results of knowledge vector extraction.
9. The data security sharing method under the AI platform according to claim 8, characterized in that, The objective function of the lower-level game model is to minimize the computational resource consumption function. The inputs to the computational resource consumption function include processor utilization, memory usage, network bandwidth requirements, and storage space usage. The output is the resource consumption weight coefficient. The processor utilization and memory usage are derived from the system performance monitoring data of the data monitoring module.
10. The data security sharing method under the AI platform according to claim 9, characterized in that, The objective functions of the upper-level game model and the lower-level game model are linked through a security-efficiency balance coupling term. The security-efficiency balance coupling term establishes constraints based on the product relationship between the data protection strength coefficient and the resource consumption weight coefficient. The data protection strength coefficient and the resource consumption weight coefficient are used for numerical calculation of the game equilibrium solution of the confusion strength optimization function.
Citation Information
Cited By
Enterprise-level image-text document management method, medium and system based on dynamic authority control
CN121434113A