Data secrecy method and system for client

By constructing a pseudo-dataset and a global model, identifying heterogeneous features, determining the distribution of local features, and establishing a personalized model and a dynamic security system, the problem of low accuracy of personalized models in multiple client interactions is solved, and high-accuracy, multi-layered security of data is achieved.

CN121966911APending Publication Date: 2026-05-01ZHEJIANG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG NORMAL UNIV
Filing Date
2025-11-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, the impact of heterogeneous features is ignored when multiple clients interact, resulting in low accuracy of personalized models and an inability to effectively protect the multi-layered confidentiality system of data.

Method used

By collecting target layer features and models from various clients, constructing a pseudo dataset, identifying heterogeneous features, establishing a global model, determining the distribution of local features, constructing a personalized model, and building a dynamic security system based on the personalized model and client data, multiple layers of data security are achieved.

Benefits of technology

It improves the accuracy of personalized models, enables holistic consideration of data isolation and confidentiality events, and enhances the accuracy of the multi-layered data confidentiality system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121966911A_ABST
    Figure CN121966911A_ABST
Patent Text Reader

Abstract

The invention discloses a data secrecy method and system for clients, and relates to the technical field of data secrecy, and the method comprises the steps: building a corresponding global model according to heterogeneous features and corresponding clients; determining neurons of each target layer according to the recognition of the global model, determining local feature distribution of each client based on the neurons of each target layer and the corresponding transmission plan, and constructing a corresponding personalized model according to the local feature distribution of each client and the personalized factors corresponding to the client, and the accuracy of the personalized model is improved. Therefore, a plurality of dynamic secrecy items are determined based on the identification of the dynamic secrecy system of the client, a corresponding data isolation event is constructed according to the plurality of dynamic secrecy items and the data distribution system corresponding to the client, and a multi-secrecy system of data is determined based on the data isolation event, the data secrecy event and the client. And the accuracy of the multiple secrecy system of the data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of data confidentiality, and more particularly to a method and system for client-side data confidentiality. Background Technology

[0002] With the development of technology, multiple clients interact and transmit data. Data is sent across multiple clients. Current technologies collect target-layer features from each client, control these features, and label the corresponding global model. However, this ignores the impact of heterogeneous features and fails to control the distribution of local features across different clients, affecting the accuracy of personalized models and resulting in lower accuracy of multi-layered data security systems. Summary of the Invention

[0003] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides a method and system for client-side data confidentiality.

[0004] This invention provides a client-side method for protecting data confidentiality, comprising: Collect target layer features and corresponding models from various clients, and determine the corresponding pseudo dataset based on the integration of target layer features and corresponding models; Based on the identification of pseudo datasets, the corresponding heterogeneous features are determined, and a corresponding global model is constructed based on the heterogeneous features and the corresponding client. The neurons of each target layer are determined based on the recognition of the global model. The local feature distribution of each client is determined based on the neurons of each target layer and the corresponding transmission plan. The corresponding personalized model is constructed based on the local feature distribution of each client and the personalized factors corresponding to the client. A personalized model is embedded into the corresponding client. Based on the personalized model and the client's data, a corresponding data confidentiality event is constructed. The dynamic confidentiality system of the client is determined according to the data confidentiality event, the importance level of the client's data, and the client's previous confidentiality events. Based on the identification of the client's dynamic security system, multiple dynamic security items are determined. According to the multiple dynamic security items and the data distribution system corresponding to the client, corresponding data isolation events are constructed. Based on the data isolation events, data security events, and the client, a multi-layered security system for the data is determined.

[0005] This invention provides a client-side data confidentiality system, which is applied to the aforementioned client-side data confidentiality method. The client-side data confidentiality system includes: The pseudo dataset module is used to collect target layer features and corresponding models from various clients, and to determine the corresponding pseudo dataset based on the integration of target layer features and corresponding models. The global model module is used to determine the corresponding heterogeneous features based on the identification of pseudo datasets, and to build the corresponding global model based on the heterogeneous features and the corresponding client. The personalized model module is used to determine the neurons of each target layer based on the recognition of the global model, determine the local feature distribution of each client based on the neurons of each target layer and the corresponding transmission plan, and construct the corresponding personalized model based on the local feature distribution of each client and the personalized factors corresponding to the client. The dynamic security system module is used to embed personalized models into corresponding clients, construct corresponding data security events based on the personalized models and the client's data, and determine the client's dynamic security system based on the data security events, the importance level of the client's data, and the client's previous security events. The multi-layered security system module is used to identify multiple dynamic security items based on the client's dynamic security system. It constructs corresponding data isolation events based on the multiple dynamic security items and the data distribution system corresponding to the client. Based on the data isolation events, data security events, and the client, it determines the multi-layered security system of the data.

[0006] Compared with the prior art, the beneficial effects of the present invention are: In this embodiment of the invention, the method collects target layer features and corresponding models from various clients, determines corresponding pseudo datasets based on the integration of target layer features and corresponding models, identifies corresponding heterogeneous features based on the identification of pseudo datasets, constructs corresponding global models based on these heterogeneous features and corresponding clients, identifies neurons in each target layer based on the identification of the global model, determines the local feature distribution of each client based on the neurons in each target layer and the corresponding transmission plan, and constructs corresponding personalized models based on the local feature distribution of each client and the personalized factors corresponding to that client. The introduction of a global model controls the local feature distribution of each client, taking into account both the local feature distribution of each client and the overall personalized factors corresponding to that client, thus improving the accuracy of the personalized models.

[0007] Therefore, a personalized model is embedded into the corresponding client. Based on the personalized model and the client's data, a corresponding data confidentiality event is constructed. The dynamic confidentiality system of the client is determined according to the data confidentiality event, the importance level of the client's data, and the client's previous confidentiality events. Based on the identification of the client's dynamic confidentiality system, multiple dynamic confidentiality items are identified. Based on the multiple dynamic confidentiality items and the data distribution system corresponding to the client, a corresponding data isolation event is constructed. Based on the data isolation event, the data confidentiality event, and the client, a multi-layered data confidentiality system is determined. The dynamic confidentiality system of the client is introduced, further controlling the data isolation event. This achieves a holistic consideration of the data isolation event, the data confidentiality event, and the client, improving the accuracy of the multi-layered data confidentiality system. Attached Figure Description

[0008] Figure 1 This is a flowchart illustrating the client-side data confidentiality method in an embodiment of the present invention; Figure 2 This is a flowchart illustrating step S11 of the client-side data confidentiality method in an embodiment of the present invention. Figure 3 This is a flowchart illustrating step S12 of the client-side data confidentiality method in an embodiment of the present invention. Figure 4 This is a flowchart illustrating step S13 of the client-side data confidentiality method in an embodiment of the present invention. Figure 5 This is a flowchart illustrating step S14 of the client-side data confidentiality method in an embodiment of the present invention. Figure 6 This is a flowchart illustrating step S15 of the client-side data confidentiality method in an embodiment of the present invention. Figure 7 This is a schematic diagram of the structural composition of the client-side data confidentiality system in an embodiment of the present invention. Detailed Implementation

[0009] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0010] Please see Figures 1 to 7 A client-side method for protecting data confidentiality, applied in data confidentiality scenarios; the client-side method for protecting data confidentiality includes: Step S11: Collect target layer features and corresponding models from each client, and determine the corresponding pseudo dataset based on the integration of target layer features and corresponding models; Step S12: Based on the identification of the pseudo dataset, determine the corresponding heterogeneous features, and construct the corresponding global model based on the heterogeneous features and the corresponding client; Step S13: Based on the recognition of the global model, determine the neurons of each target layer, determine the local feature distribution of each client based on the neurons of each target layer and the corresponding transmission plan, and construct the corresponding personalized model based on the local feature distribution of each client and the personalized factors corresponding to the client. Step S14: Embed the personalized model into the corresponding client, construct the corresponding data confidentiality event based on the personalized model and the client's data, and determine the client's dynamic confidentiality system based on the data confidentiality event, the importance level of the client's data, and the client's previous confidentiality events; Step S15: Based on the identification of the client's dynamic security system, determine multiple dynamic security items, construct corresponding data isolation events based on the multiple dynamic security items and the data distribution system corresponding to the client, and determine the multi-layered security system of the data based on the data isolation events, data security events, and the client.

[0011] refer to Figure 2 In step S11, the specific steps are as follows: S111: Label each client, determine the corresponding feature set based on the detection of each client, determine the target layer features and the corresponding model based on the traversal of the feature set, so as to collect the target layer features and the corresponding model; S112: Determine the corresponding matching coefficient based on the matching between the target layer features and the corresponding model. Determine the corresponding integration method based on the mapping relationship between the matching coefficient, the target layer features, and the integration method. Determine the corresponding pseudo dataset based on the integration method, the target layer features, and the corresponding model.

[0012] In the embodiments of this application, each client is marked, and a corresponding feature set is determined based on the detection of each client. The target layer features and the corresponding model are determined by traversing the feature set, so as to collect the target layer features and the corresponding model. This approach takes into account the overall consideration of feature set traversal and ensures the accuracy of the target layer features and the corresponding model.

[0013] At this point, in the federated learning network architecture, the server must have the ability to uniquely identify each participating client. This process is achieved by assigning globally unique identifiers (such as client ID k ∈ {1, 2, …, K}). These identifiers not only bear the key responsibility of data routing, but also play a core role in indexing and access control in the subsequent aggregation algorithm and personalized model distribution stages. From an information security perspective, this marking mechanism establishes the technical starting point for the system's trust boundary, ensuring that knowledge flow is strictly limited to authorized entities, and laying the foundation for identity authentication for the entire confidentiality system.

[0014] The system prioritizes deep hidden layers (such as fully connected layers and BatchNorm layers) that are close to the model output because the high-level abstract features extracted by these layers best reflect the unique distribution characteristics of the client's data, while shallow features (such as edges and textures) have limited contribution to personalization due to their strong generality.

[0015] After determining the target layer, the client performs a forward propagation operation using its local private dataset. For each local data sample x_i, the model's output at the target layer is the feature vector h_i. By traversing all samples or representative batches in the local dataset, a set of feature vectors {h_i} is generated. This set essentially constitutes the digital fingerprint of the client's local data distribution.

[0016] The system iterates through the feature set {h_i} generated in the previous step and records the corresponding original model output logits z_i for each feature vector h_i. These logits contain the model's confidence scores for the sample belonging to each category without softmax normalization, so that the knowledge of each sample is represented as a structured data pair (h_i, z_i).

[0017] The information collected and uploaded to the server by the client contains three key components: the target layer feature set {h_i} representing the local data distribution, the corresponding model output {z_i} representing the local model's interpretation of the features, and the local model parameters w_k used by the server for preliminary global model averaging. These three parts together constitute the complete knowledge package uploaded by the client. By packaging features and corresponding logits, the client provides contextual knowledge, enabling the server to perform more effective knowledge distillation. At the same time, since the original data is not uploaded, and the server subsequently performs secondary training using a pseudo-dataset, the direct connection between the client's model parameters and the final global model is further isolated, forming a dual technical protection mechanism.

[0018] Specifically, each client corresponds to a hospital, and the hospital is assigned a unique client ID, Client_A. All data packets sent from the hospital carry this identification tag, ensuring the traceability of data sources and the accuracy of access control. In the process of determining the corresponding feature set based on the hospital's detection, the federated learning framework pre-specifies that all clients use the standard ResNet-18 model and designates the fully connected layer after the global average pooling layer (GAP) as the target layer. The features output by this layer best represent the abstract patterns of different types of pneumonia lesions.

[0019] The hospital used a private dataset containing thousands of chest X-ray images on its internal server to perform forward propagation on a local ResNet-18 model. For each X-ray image x_i, the model outputs a 512-dimensional feature vector h_i at the target layer. After traversing the entire dataset, a feature set {h_i} containing thousands of 512-dimensional vectors is generated. This set abstractly represents the distribution of radiographic features of pneumonia cases treated by the hospital.

[0020] The hospital system then traverses the dataset again, recording not only the 512-dimensional feature vector h_i of the target layer for each X-ray image, but also the logits z_i (e.g., z_i = [2.1, -0.5]) of the last layer output of the model, which includes the two categories of pneumonia / normal. This packs the knowledge of each image into (h_i, z_i) pairs. The hospital then encrypts three key data structures and uploads them to the central server: a large matrix containing the 512-dimensional feature vectors of all local X-ray images, a matrix containing all the corresponding logits, and the weights w_A of the ResNet-18 model trained locally by the hospital.

[0021] Furthermore, the matching coefficient is determined based on the matching between the target layer features and the corresponding model. The corresponding integration method is determined based on the mapping relationship between the matching coefficient, the target layer features, and the integration method. The corresponding pseudo dataset is determined based on the integration method, the target layer features, and the corresponding model. This approach takes into account the overall considerations of the integration method, the target layer features, and the corresponding model, ensuring the accuracy of the corresponding pseudo dataset.

[0022] At this point, the server precisely pairs the target layer feature set {h_i^k} uploaded by each client with its corresponding model output {z_i^k} to form a complete (h_i^k, z_i^k) knowledge unit, ensuring a tight binding between features and interpretation. The matching coefficient c_k serves as a quantitative indicator, and its determination strategy is highly flexible: it can be based on data volume weight, with the coefficient proportional to the number of local data samples N_k on the client, and clients with larger data volumes receiving higher weights due to their more representative knowledge distribution; it can also adopt a data quality weight strategy, evaluating data quality based on indicators such as the local validation accuracy of the client model and the diversity of feature distribution, with high-quality clients receiving corresponding weights; in some scenarios that protect clients with small data volumes, an equal weight strategy can also be used.

[0023] Determining the corresponding integration method based on the mapping relationship between the matching coefficient, target layer features, and integration method is a multi-parameter input intelligent decision-making process. It determines the optimal fusion strategy for all client knowledge. This decision-making process comprehensively considers three key input parameters: the matching coefficient c_k determines the volume of each client knowledge during fusion; the target layer features {h_i^k} represent the distribution of core knowledge provided by all clients; and the mapping relationship of integration methods is a preset rule base or strategy function that defines the specific integration algorithm to be used under different conditions.

[0024] The system selects the optimal integration strategy based on these input parameters. In the FLKI framework, the core integration method is to construct a pseudo-dataset for knowledge distillation, which can better handle data heterogeneity among clients compared to simple parameter averaging. The mapping relationship can be simply defined as: IF (high data heterogeneity) THEN (integrate using knowledge distillation) ELSE (integrate using parameter averaging). Optionally, the FLKI framework is a federated learning framework under knowledge isolation.

[0025] The server begins constructing a pseudo dataset D_pseudo according to the selected knowledge distillation and integration method. It gathers all knowledge units (h_i^k, z_i^k) from all clients k into a large set, forming the pseudo dataset D_pseudo = {(h_i, z_i)}, where h_i and z_i come from any client. The pseudo nature of this dataset is that it does not contain any real original data, but only abstract feature vectors and model outputs. However, in a statistical sense, it simulates a huge virtual dataset containing the knowledge of all clients.

[0026] Specifically, the server received data packets from 10 hospitals. Hospital B uploaded 5,000 (h_i, z_i) pairs, Hospital B uploaded 3,000, and the number of pairs uploaded by the other hospitals varied. Based on a preset data volume weighting strategy, the server calculated the matching coefficient c_A for each hospital as 5,000 divided by the total number of samples from all hospitals, and for Hospital B as 3,000 divided by the total number of samples from all hospitals, and so on. This weighting ensures that hospitals, as major data contributors, have greater say in the global fusion process, reflecting a positive correlation between contribution and weight.

[0027] The server's data analysis revealed that the hospital's X-ray images mainly came from the elderly population, while Hospital C's samples mainly came from the children population, showing a significant difference in data distribution (heterogeneity) between the two. Based on the pre-defined high heterogeneity, a knowledge distillation mapping relationship was adopted. The server determined that this aggregation would use feature-level knowledge distillation based on pseudo-datasets as the integration method, rather than simple parameter averaging, thereby avoiding the dilution of the specific characteristics of the children or elderly populations during the aggregation process and ensuring the complete preservation of the characteristics of each group.

[0028] The server began performing the specific integration operation, aggregating all 5,000 (h_i, z_i) pairs from Hospital B, 3,000 from Hospital C, and 4,000 from Hospital C, etc.; ultimately constructing a pseudo dataset D_pseudo containing tens of thousands of (h_i, z_i) pairs. This dataset is like a database containing tens of thousands of virtual X-ray images, but none of them are real patient images; each virtual X-ray image h_i is accompanied by a diagnostic opinion z_i given by a hospital model, forming a complete virtual knowledge base.

[0029] refer to Figure 3 In step S12, the specific steps are as follows: S121: Collect a pseudo dataset, determine multiple sub-heterogeneous data based on the detection of the pseudo dataset, determine the corresponding heterogeneous features based on the multiple sub-heterogeneous data, and mark the feature position and corresponding feature shape of the heterogeneous features. S122: Determine the first level of global content based on the feature location of the heterogeneous feature and the corresponding client, determine the second level of global content based on the feature shape of the heterogeneous feature and the corresponding client, and construct the corresponding global model based on the first level of global content, the second level of global content and the overall shape of the client.

[0030] In the embodiments of this application, a pseudo dataset is collected, and multiple sub-heterogeneous data are determined based on the detection of the pseudo dataset. The corresponding heterogeneous features are determined according to the multiple sub-heterogeneous data, and the feature position and corresponding feature shape of the heterogeneous features are marked. This approach takes into account the overall consideration of multiple sub-heterogeneous data and ensures the accuracy of the corresponding heterogeneous features.

[0031] At this point, the server directly obtains the pseudo dataset D_pseudo constructed in step S112. This dataset appears to be an unlabeled mixed feature pool, but in reality, it contains a collection of knowledge from all clients. By detecting this metadata analysis process, the server checks the source label of each feature vector h_i, that is, identifies the client k that uploaded it. Due to the natural heterogeneity of the data distribution of different clients, the features from different clients will naturally form different clusters statistically.

[0032] Based on these source labels, the server logically divides the entire pseudo dataset D_pseudo into K subsets. Each subset D_pseudo^k mainly contains features h_i and corresponding output z_i from client k. These subsets are the sub-heterogeneous data. While they together constitute the complete pseudo dataset, they each maintain the distribution characteristics of the source clients.

[0033] The server explicitly defines the feature vector h_i in each sub-heterogeneous data D_pseudo^k as a heterogeneous feature. This naming emphasizes that they do not come from the same distribution, but each represents a unique data pattern from its source client. These heterogeneous features become the atomic units for all subsequent analysis and operations, providing the system with clear technical operation objects.

[0034] The labeling of feature locations defines the specific level and dimension of the feature in the neural network architecture. All clients follow the protocol to upload features from the same target layer, so the feature locations of all heterogeneous features are unified. For example, the 512-dimensional feature vector after the global average pooling layer of the ResNet-18 model. This unified labeling ensures that all heterogeneous features are in the same coordinate system, enabling them to be aligned, compared and mathematically operated on, which is a prerequisite for knowledge distillation.

[0035] The feature morphology label refers to the statistical distribution characteristics of each sub-heterogeneous dataset. The server calculates a statistical description for the feature set of each client k, including the mean vector μ_k (representing the central position of the client's feature in the feature space) and the covariance matrix Σ_k (describing the dispersion of the feature and the correlation between different dimensions). This morphology quantifies the statistical differences between the features of different clients, providing a precise mathematical basis for subsequent morphology alignment and personalized adaptation.

[0036] Specifically, the server collected a pseudo dataset D_pseudo containing tens of thousands of (h_i, z_i) pairs. By detecting the source label of each h_i, the server logically divided D_pseudo into 10 subsets, where D_pseudo^A contains 5,000 knowledge pairs from hospitals, D_pseudo^B contains 3,000 from hospital B, and so on. These 10 subsets are the sub-heterogeneous data, and each subset maintains the data distribution characteristics of the corresponding hospital.

[0037] The server defines the 5000 512-dimensional feature vectors in D_pseudo^A as heterogeneous features of hospitals, the 3000 feature vectors in D_pseudo^B as heterogeneous features of hospital B, and so on; each hospital's feature vector becomes an independent heterogeneous feature entity, laying the foundation for subsequent personalized processing.

[0038] The server assigns a uniform location label to the heterogeneous features of all 10 hospitals: a 512-dimensional feature layer after the GAP of ResNet-18, ensuring that a feature vector of a hospital can be used to calculate distance with the feature vector of Hospital B in the same 512-dimensional space.

[0039] The server calculates the feature morphology of the heterogeneous features of the hospital: it calculates the mean vector μ_A of 5000 feature vectors of the hospital and finds that they form a dense cluster in a certain region of the feature space; it calculates the covariance matrix Σ_A and finds that its variance is small, indicating that the case features of the hospital are relatively concentrated and typical; the server also calculates the unique μ_k and Σ_k of the other 9 hospitals, and finally obtains digital archives containing 10 different morphologies, which fully prepares for the next step of precise knowledge fusion.

[0040] Furthermore, the first layer of global content is determined based on the feature location of the heterogeneous feature and the corresponding client, and the second layer of global content is determined based on the feature shape of the heterogeneous feature and the corresponding client. A corresponding global model is constructed based on the first layer of global content, the second layer of global content and the overall shape of the client, which takes into account the overall consideration of the first layer of global content, the second layer of global content and the overall shape of the client, and ensures the accuracy of the corresponding global model.

[0041] At this point, the first layer of global content is determined based on the feature location of the heterogeneous feature and the corresponding client. The feature location refers to the feature uploaded by all clients in accordance with the protocol and from the same target layer. The server uses this unified location information and the identity k of each client to perform a weighted average operation of the model parameters.

[0042] The server collects the local model parameters w_k uploaded by all clients and performs a weighted average based on the amount of data N_k from each client to obtain a preliminary global model ω_{t+1}. This model is the average of the model structures of all clients and incorporates the model architecture contributions of all participants.

[0043] The second layer of global content is determined based on the feature shape of the heterogeneous feature and the corresponding client. The feature shape is the statistical distribution (such as mean μ_k and covariance Σ_k) calculated for each client in S121. The server uses this morphological information to perform knowledge distillation on the pseudo dataset D_pseudo constructed in S11.

[0044] The server uses the pseudo dataset D_pseudo as training data to refine the initial global model ω_{t+1} for the first layer of content generation. During the distillation process, for a feature h_i from client k, the global model is trained to mimic the output (soft label z_i) of the client k model for that feature, thereby enabling the global model to learn different forms of knowledge and absorb the behavioral patterns of all clients.

[0045] A global model was constructed based on the first-level global content, the second-level global content, and the overall form of the client, and the final synthesis of the global model was completed. Based on the first-level global content, the second-level global content, and the overall form of the client, the overall form of the client can be understood as the comprehensive knowledge distribution reflected by the entire pseudo dataset. The server starts with the first-level content ω_{t+1} and optimizes it through the distillation training process of the second-level content.

[0046] After multiple iterations, the initial model was optimized into the final global model ω_final. This model possesses the average structure of all client model parameters (first layer) and has learned expert-level judgment ability on various heterogeneous data distributions through distillation (second layer), becoming a powerful virtual expert model with outstanding generalization ability.

[0047] Specifically, the server takes the ResNet-18 model parameters uploaded by 10 hospitals and calculates a weighted average based on their respective data volumes (5000 for Hospital B, 3000 for Hospital B, etc.) to generate a preliminary global model ω_{t+1}. This model is the average skeleton of the model structure of the 10 hospitals, integrating the model contributions of all participants, and laying the structural foundation for subsequent knowledge fusion.

[0048] The server begins knowledge distillation; it inputs the features h_i^A from the hospitals in the pseudo-dataset into the initial model and forces the model's output p_g(h_i^A) to approximate the output z_i^A (soft label) given by the original hospital model. This process teaches the global model to interpret the features of elderly cases like hospital experts. Similarly, the server is also trained with features and outputs from hospitals B, C, etc., teaching the global model how to interpret different forms of knowledge such as pediatric cases and rare cases, thus achieving a comprehensive improvement in the model's capabilities.

[0049] After multiple rounds of distillation training on a pseudo dataset containing tens of thousands of virtual samples, the initial model ω_{t+1} was refined into the final global model ω_final. The body of this final model is derived from the average structure of the models from 10 hospitals, while its core integrates all the knowledge from hospitals such as their expertise in geriatric pneumonia and Hospital B's expertise in pediatric pneumonia, becoming an unprecedentedly powerful virtual expert in pneumonia detection. Throughout the entire process, the original patient data of all participating parties, including hospitals, was kept strictly confidential, achieving a perfect balance between data value and privacy protection.

[0050] refer to Figure 4 In step S13, the specific steps are as follows: S131: Label the global model, determine the corresponding global distribution system based on the recognition of the global model, determine multiple target layers based on the detection of the global distribution system, and label the neurons of each target layer; S132: Based on the detection of the client, determine the previous transmission events, determine the corresponding transmission plan based on the identification of the previous transmission events, and determine the local feature distribution of each client based on the neurons of each target layer and the corresponding transmission plan. S133: Collect the personalized factors corresponding to the client, and determine the construction of the corresponding personalized model based on the local feature distribution of each client and the corresponding personalized factors.

[0051] In the embodiments of this application, the global model is labeled, the corresponding global distribution system is determined based on the recognition of the global model, multiple target layers are determined based on the detection of the global distribution system, and the neurons of each target layer are labeled, which is compatible with the overall consideration of the detection of the global distribution system and ensures the accuracy of multiple target layers.

[0052] At this point, the server explicitly marks and versions the final global model ω_final obtained after knowledge distillation and optimization in S12. This model, as the crystallization of the wisdom of all clients, becomes the sole and authoritative blueprint for all subsequent personalized operations.

[0053] The server performs structured analysis on ω_final to form a globally distributed system. This is not just a simple list of model files, but a hierarchical knowledge architecture diagram. The server identifies different functional modules in the model, such as the input layer, multiple convolutional blocks, residual blocks, pooling layers, and the final fully connected layer. This system describes in detail how knowledge is abstracted and flows from the original input layer by layer, and finally forms a complete path for decision-making.

[0054] The server scans the knowledge architecture graph established in the previous step and, based on preset rules or heuristic algorithms, finds the most suitable layer for personalized modification. These target layers are usually deep hidden layers close to the output, such as fully connected layers or BatchNorm layers. The server will determine one or more such target layers.

[0055] The technical basis for choosing these layers is that they extract the most discriminative high-level abstract features, which directly determine the model's final decision and best reflect the distribution differences between different client data. In contrast, choosing shallow layers is not effective because shallow layers learn low-level features with strong generality, such as edges and textures.

[0056] After determining the target layer (e.g., the last fully connected layer), the server delves into the interior of that layer, which is essentially a huge weight matrix W_G. The server identifies and labels each row of this matrix as an independent neuron, and each neuron W_G[i] is a weight vector, corresponding to a specific judgment criterion for the model from high-level features to the final output category. By labeling all neurons, the server breaks down the model's decision-making ability into thousands of tiny, actionable knowledge units.

[0057] Specifically, the server labels this global ResNet-18 model as v1.0-global-pneumonia-detector and analyzes its structure to establish a global distribution system map, clearly showing the complete data flow from the input 64x64 pixel image, through multiple convolutional layers and residual blocks for feature extraction, and finally to the global average pooling layer and fully connected classifier.

[0058] The server scans this knowledge architecture graph and selects layers that are close to the output and contain high-level abstract features according to the rules. Finally, it determines the fully connected layer (fc layer) after the global average pooling layer as the target layer for this personalization. The server did not choose shallower convolutional layers because those layers learn general edges and textures and are not sensitive to hospital-specific data.

[0059] The server delves into this fully connected layer, assuming it is a layer that maps 512-dimensional feature vectors to 200 fine-grained pneumonia categories, with a weight matrix W_G of size 200×512. The server labels each row of W_G as an independent neuron. The first row of neurons W_G[1] is responsible for identifying bacterial pneumonia-type A, the second row of neurons W_G[2] is responsible for identifying viral pneumonia-type B, and so on, with a total of 200 such judgment criteria neurons labeled.

[0060] Furthermore, based on the detection of the client, previous transmission events are determined, and corresponding transmission plans are determined based on the identification of these previous transmission events. The local feature distribution of each client is determined based on the neurons of each target layer and the corresponding transmission plans, which takes into account the overall consideration of the neurons of each target layer and the corresponding transmission plans, and ensures the accuracy of the local feature distribution of each client.

[0061] At this point, the server retrieves all historical information related to client A from the database. In the context of the FLKI framework, past transmission events do not refer to communications that have occurred in the past, but rather to a sophisticated technical metaphor. It refers to a static snapshot that represents the data distribution of the client—that is, the target layer feature set Y^A uploaded by the hospital in step S11. This set serves as the digital fingerprint or knowledge DNA of the hospital's local data distribution, becoming the fundamental basis for personalized modifications.

[0062] The server uses the identified hospital feature set Y^A to construct an optimal transmission problem. The source distribution is the weight vector W_G[i] of the n neurons in the target layer of the global model labeled in S131, where each neuron is considered as a component with equal importance. The target distribution is the m feature vectors y_j of the hospital, where each feature vector is considered as a part that needs to be filled.

[0063] The server calculates an n×m cost matrix M, where M_ij represents the effort (such as the Euclidean distance of the feature vector) required to move the i-th global neuron to match the j-th local feature. The efficient Sinkhorn algorithm is used to solve the optimal transport problem, and finally an n×m transport plan matrix Ω_A is obtained.

[0064] The server uses the neuron weight matrix W_G of the target layer of the global model marked in S131, and the transmission plan matrix Ω_A calculated for the hospital in the previous step, to perform the key matrix operation W̃_A = Ω_A ⋅ W_G; it performs a weighted linear combination on the original global neuron weights W_G to generate a brand new, personalized weight matrix W̃_A specific to the hospital. This new matrix defines the local feature distribution of the hospital. The server completes the personalized transformation of the global model at a purely mathematical level. The final W̃_A is a unique feature distribution that belongs entirely to the hospital. It can be directly used to build a personalized model, and its internal knowledge representation is different from that of any other hospital, fundamentally achieving knowledge isolation.

[0065] Specifically, the server retrieves the hospital from the database and accesses the 512-dimensional feature vector set Y^A that it uploaded in S11, representing the distribution of its data. This set contains the hospital's past transmission events or knowledge background, including the hospital's unique case feature distribution information.

[0066] The server begins solving the optimal transport problem for the hospital; it treats the 200 neurons of the global model as 200 source distributions and the 512 feature vectors of the hospital as 512 target distributions; the server calculates a 200×512 cost matrix M, where M_ij measures the degree of mismatch between the i-th global neuron and the j-th feature of the hospital, runs the Sinkhorn algorithm, and finally obtains a 200×512 transport plan matrix Ω_A, which precisely describes how to reorganize the knowledge of the 200 global neurons to adapt to the 512 local features of the hospital.

[0067] The server uses the original weight matrix W_G (200×512) of the global model target layer and the transmission plan Ω_A calculated for the hospital to perform the matrix operation W̃_A = Ω_A ⋅ W_G, generating a brand new 200×512 personalized weight matrix W̃_A. This W̃_A matrix defines the local feature distribution of the hospital. It is a knowledge hybrid in which each new personalized neuron is a linear combination of global neurons according to the recipe of Ω_A.

[0068] Therefore, by collecting the personalized factors corresponding to each client, and determining the construction of the corresponding personalized model based on the local feature distribution of each client and the corresponding personalized factors, the overall consideration of the local feature distribution of each client and the corresponding personalized factors is taken into account, ensuring the accuracy of the corresponding personalized model. At the same time, a global model is introduced to control the local feature distribution of each client, taking into account the local feature distribution of each client and the overall consideration of the personalized factors corresponding to that client, thereby improving the accuracy of the personalized model.

[0069] At this point, the server retrieves all the personalized instructions and parameters prepared for a specific client k (such as a hospital) from its computing module or database. In the FLKI framework, the most core personalization factor is the transmission plan matrix Ω_k calculated through optimal transmission in step S132. This matrix realizes a precise mapping from global knowledge to local knowledge. In addition, personalization factors also include other auxiliary parameters, such as learning rate adjustment and regularization strength of specific layers. However, in the core implementation of FLKI, Ω_k is the most critical factor.

[0070] The server begins the model building process, where the local feature distribution refers to the new personalized weight matrix W̃_k generated for client k in S132. This matrix represents the unique knowledge representation of client k. The server uses the loaded personalized factors, mainly the transport plan matrix Ω_k, to guide how to generate the personalized weight matrix W̃_k from the global model ω_final.

[0071] The server completes the model construction through a replacement operation: it copies the complete structure of the global model ω_final and replaces the target layer weights with the personalized weight matrix W̃_k calculated for client k in S132. After this replacement operation, a brand new personalized model ω̃_k that belongs entirely to client k is built. The personalized model is directly deployed to the client. Due to the uniqueness of its internal knowledge representation, it cannot be used to infer information from any other client, achieving perfect knowledge isolation.

[0072] Specifically, the server retrieves the core personalized factor customized for the hospital (Client_A) from its computing cache—the transmission plan matrix Ω_A (a 200×512 matrix). This matrix is ​​the part that translates global knowledge into local hospital knowledge and contains all the mapping information of the hospital's unique case characteristics.

[0073] The server begins building the final personalized model ω̃_A for the hospital; the server copies the complete structure of the global model ω_final (a complete ResNet-18) and performs a crucial replacement operation: the original weight matrix W_G of the fully connected layer (fc layer) after the global average pooling layer in the copied model is replaced with the personalized weight matrix W̃_A calculated for the hospital in S132; after this replacement operation, a brand new, unique, and completely personalized pneumonia detection model ω̃_A belonging entirely to the hospital is built.

[0074] refer to Figure 5 In step S14, the specific steps are as follows: S141: Collect personalized models and input them into the corresponding client to achieve client embedding. At the same time, determine the client's data based on the client's detection and construct the corresponding data confidentiality event based on the personalized model and the client's data. S142: Collect client data and mark the importance level of the client data. Determine the first level of dynamic confidentiality content based on the data confidentiality event and the importance level of the client data. S143: Determine the second layer of dynamic confidentiality content based on the data confidentiality event and the client's previous confidentiality events, and determine the client's dynamic confidentiality system based on the first layer of dynamic confidentiality content and the second layer of dynamic confidentiality content.

[0075] In the embodiments of this application, a personalized model is collected and input into the corresponding client to achieve client embedding. At the same time, the client's data is determined based on the client's detection. A corresponding data confidentiality event is constructed based on the personalized model and the client's data, which takes into account the overall consideration of the personalized model and the client's data and ensures the accuracy of the corresponding data confidentiality event.

[0076] At this point, the server retrieves a personalized model ω̃_k tailored for a specific client k (such as a hospital) from its model generation module. This model is a product of S13, incorporating global intelligence and adapting to local data distribution. The server sends the model file ω̃_k to the corresponding client k through a secure, certified communication channel (such as an encrypted API channel using the TLS 1.3 protocol). The transmission process also includes model integrity verification (such as using hash value verification) to ensure that the model has not been tampered with during transmission.

[0077] After receiving the model, the client k deploys it in a local computing environment (such as a hospital's internal server or edge computing device). This process is called embedding, which means that the personalized model is now part of the client's local technology stack and can be invoked at any time to process local data.

[0078] The client's local system identifies and monitors its own data assets. When the personalized model ω̃_k is invoked, the system will detect the data object currently being processed. The system will specify the metadata of the data, including the data source (such as PACS image archiving system, EMR electronic medical record system), data type (such as chest X-ray, CT scan image or structured laboratory test results), and data identifier (such as anonymous patient ID, examination sequence number, etc., identifiers that can uniquely identify the data item).

[0079] Based on the personalized model and the client's data, a corresponding data confidentiality event is constructed. This data confidentiality event is a structured log entry containing rich metadata, including at least an event ID (a globally unique identifier used for tracking), a timestamp (the precise time the event occurred), a client identifier (e.g., Client_A), a model identifier (e.g., ω̃_A_v1.2), a data object identifier (a unique ID of the processed data), an operation type (e.g., inference, training), and an operation result summary (e.g., output_label: pneumonia). Furthermore, client data is collected and its importance level is marked. The first layer of dynamic confidentiality content is determined based on the importance level of the data confidentiality event and the client's data, ensuring the accuracy of the first layer of dynamic confidentiality content while considering the overall importance levels of both the data confidentiality event and the client's data.

[0080] At this point, the system identifies and extracts metadata from the data that the client is about to process or is currently processing. The system scans information such as the source, format, and related business processes of the data. Based on preset rules or strategies, the system assigns a multi-dimensional importance level to each type or each data item, including sensitivity level (such as public, internal, confidential, top secret, which determines the potential harm of data leakage), business impact level (such as low, medium, high, critical, which determines the impact of data loss or damage on business operations), and compliance requirement level (such as whether it is subject to strict constraints of laws and regulations such as GDPR, HIPAA, and cybersecurity laws, which determines the legal risks of data processing).

[0081] The system determines the first level of dynamic confidentiality content based on the importance level of the data confidentiality event and the client's data. The system combines the data confidentiality event generated in S141 (representing a specific model-data interaction) with the data importance level determined in the previous step to form a complete context. Based on this complete context, the system will query or calculate a set of confidentiality policies, namely the first level of dynamic confidentiality content. This is a dynamically generated set of policies whose content is directly linked to the risk level of the data.

[0082] For high-level data (such as top secret), the first level of measures includes forcibly enabling the highest level of access control, enabling end-to-end encryption, recording the most detailed operation audit logs, prohibiting any form of model result export, and triggering real-time security monitoring alarms; for low-level data (such as public), the first level of measures only includes regular access log recording; for medium-level data (such as confidential), the first level of measures is somewhere in between.

[0083] Therefore, the second layer of dynamic confidentiality content is determined based on the data confidentiality event and the client's previous confidentiality events. The dynamic confidentiality system of the client is determined based on the first and second layers of dynamic confidentiality content, which takes into account the overall consideration of the first and second layers of dynamic confidentiality content and ensures the accuracy of the client's dynamic confidentiality system.

[0084] At this point, the system not only focuses on the current data confidentiality events, but also delves into and analyzes all the previous confidentiality event records of client k within the historical time window (such as the past 30 days or 90 days). Based on the multi-dimensional analysis of historical events, the system generates a risk profile or trust score about the client's behavior. This profile is the second layer of dynamic confidentiality content.

[0085] The analysis dimensions include compliance history (whether the client frequently triggers access control policies, and whether there are records of multiple failed privilege escalation attempts), security event history (whether the client environment has experienced security events such as malware infection or abnormal network connection), abnormal operation patterns (whether the client's data access volume, time pattern, and operation frequency suddenly deviate from its baseline behavior), and abnormal model interaction (whether the performance of the personalized model, such as accuracy and confidence distribution, shows an unexplained decline or fluctuation). The system quantifies these analysis results into a trust score (e.g., a maximum score of 100) or a behavioral risk label (e.g., low risk, medium risk, high risk).

[0086] The system uses a policy fusion engine to weight and combine the content of the first layer of dynamic confidentiality (from S142, representing the risk of the current task, i.e. how important the data being processed is) and the second layer of dynamic confidentiality (from S143 sub-step 1, representing the risk of the client, i.e. whether the operator is trustworthy) to generate a complete, real-time dynamic confidentiality system for the client k.

[0087] A dynamic confidentiality system is a specific and actionable set of strategies. Its strength and scope are determined by two risk dimensions: When dealing with high-value data and high-trust clients, the system adopts standard confidentiality strategies; when dealing with high-value data and low-trust clients, the system will activate the strictest zero-trust mode, such as implementing continuous multi-factor authentication, real-time auditing of every operation, restricting model functionality, and even running the model in an isolated sandbox environment; when dealing with low-value data and high-trust clients, the system will relax some non-core confidentiality measures to improve operational efficiency; when dealing with low-value data and low-trust clients, the system will still maintain a moderate level of monitoring, because low-trust behavior itself is a risk signal.

[0088] refer to Figure 6 In step S15, the specific steps are as follows: S151: Real-time monitoring of the client's dynamic security system, determining multiple dynamic surface markers based on the detection of the client's dynamic security system, determining the corresponding dynamic security items based on the tracing of each dynamic surface marker, and collecting multiple dynamic security items. S152: Determine the data distribution system corresponding to the client based on the client's system identification, and construct corresponding data isolation events based on the data distribution system corresponding to the client and multiple dynamic confidentiality projects; S153: Based on the data isolation event and the client, determine the first layer of data security system; based on the data isolation event and the data security event, determine the second layer of data security system; and construct a multi-layered data security system based on the first and second layers of security systems.

[0089] In the embodiments of this application, the dynamic security system of the client is monitored in real time. Multiple dynamic surface markers are determined based on the detection of the client's dynamic security system. The corresponding dynamic security items are determined based on the tracing of each dynamic surface marker. Multiple dynamic security items are collected, which takes into account the overall consideration of tracing each dynamic surface marker and ensures the accuracy of the corresponding dynamic security items.

[0090] At this point, the monitoring targets all components of the dynamic confidentiality system built for the client in S14, including policy enforcement monitoring (monitoring whether access control policies are executed correctly and whether encryption policies are effective), behavior monitoring (monitoring user and system behavior patterns, such as login time, data access volume, API call frequency, etc.), performance monitoring (monitoring the running status of the personalized model ω̃_k, such as inference latency, CPU / GPU utilization, output confidence distribution, etc.), network traffic monitoring (monitoring all data packets entering and leaving the client's network boundary and detecting abnormal connections or data transmissions), and log and audit monitoring (collecting and analyzing logs from the operating system, applications, databases, and security devices in real time).

[0091] The system uses various detection engines to analyze the monitoring data collected in the previous step in real time. These engines include rule-based engines (matching predefined attack patterns or violations, such as multiple failed login attempts), statistics-based engines (detecting abnormal fluctuations in statistical indicators, such as a sudden surge in network traffic), and machine learning-based engines (learning a baseline of normal behavior to identify abnormal behaviors that deviate from the baseline, such as users logging in from abnormal locations outside of working hours).

[0092] When any detection engine detects an anomaly, it immediately generates a dynamic surface marker. This marker is a lightweight, standardized risk signal. It is not a complete security event in itself, but a clue that requires further investigation.

[0093] Once the system generates a dynamic surface tag, it immediately initiates an automated tracing process. This process includes correlation analysis (associating the tag with other relevant monitoring data, logs, and other tags to find the temporal and logical relationships between them), context enrichment (retrieving contextual information related to the tag from asset management systems, user directories, CMDBs, etc., such as which server the involved IP address belongs to, who the involved user is, and what the function of the server is), and threat intelligence matching (matching the IOCs in the tag with external threat intelligence databases to see if they are known attack activities).

[0094] After tracing and enriching the context, if the system confirms that this is a real security risk or event, it will establish it as a dynamic confidential project. This project is a structured security event record with rich context, which usually includes event ID, severity level, detailed description, affected assets, timeline, preliminary root cause hypothesis, and recommended remedial measures. All established dynamic confidential projects will be collected into a unified security event management platform or work order system for unified tracking, assignment, investigation, and closed-loop management.

[0095] Furthermore, based on the client's system identification, the data distribution system corresponding to the client is determined. Based on the data distribution system corresponding to the client and multiple dynamic security projects, corresponding data isolation events are constructed. This approach takes into account the overall considerations of the client's data distribution system and multiple dynamic security projects, ensuring the accuracy of the corresponding data isolation events.

[0096] At this point, the system performs a comprehensive and automated system identification of the entire IT environment of client k. This is not just a simple scan, but a deep discovery and correlation through various technical means. Through system identification, the system constructs a dynamic and multi-dimensional data distribution system. This system is a structured knowledge graph that details the data storage layer (where the data is physically or logically stored, such as specific database instances, file server paths, cloud storage buckets), the data processing layer (which applications, services, or models are processing this data, such as the inference server running the personalized model ω̃_k, EMR electronic medical record applications), the data flow layer (how the data flows between different systems and network areas, such as the data flow from the PACS system to the AI ​​inference server), and the data ownership layer (which business department, project, and confidentiality level the data belongs to, such as all patient image data belonging to the radiology department and marked as confidential).

[0097] The system intelligently correlates and overlays the dynamic confidentiality items (i.e. confirmed security risk events) collected in S151 with the data distribution system (i.e. data map) constructed in the previous step. When a dynamic confidentiality item is precisely located to one or more specific assets in the data distribution system, the system will immediately trigger and construct a data isolation event. This event is a specific set of action instructions that can be executed automatically, with the aim of severing the connection between the risk and the core data in the shortest possible time.

[0098] Data isolation incidents include network layer isolation (immediately blocking the network connection between the attack source IP and the network segment where the affected data asset is located via firewall or SDN controller), host layer isolation (disconnecting the infected host from the network or placing it in an isolated VLAN via endpoint detection and response system), application layer isolation (immediately disabling API access permissions for affected application accounts via API gateway or application firewall), and data layer isolation (immediately setting the affected database tablespace to read-only mode or revoking access permissions for specific users via database management system).

[0099] Therefore, based on the data isolation event and the client, a first layer of data confidentiality is determined, and based on the data isolation event and the data confidentiality event, a second layer of data confidentiality is determined. Based on the first and second layers of confidentiality, a multi-layered data confidentiality system is constructed, which is compatible with the overall consideration of data isolation events and data confidentiality events, ensuring the accuracy of the second layer of data confidentiality. At the same time, the dynamic confidentiality system of the client is introduced to further control the data isolation event, realize the overall consideration of data isolation events, data confidentiality events and the client, and improve the accuracy of the multi-layered data confidentiality system. At this point, the system will analyze the data isolation events triggered in S152. These real-world risks expose weaknesses in the existing defense system. The system will then feed back these practical experiences to strengthen and define the first layer of security. This layer is a relatively static set of technical and management strategies built around the data assets themselves, with the goal of ensuring that the data is protected at all times.

[0100] Its core content includes data encryption strategy (defining encryption standards and algorithms for data at rest, in transit, and in use), data isolation and segmentation strategy (defining network segmentation rules, virtual LAN division, and physical or logical isolation requirements for data of different sensitivity levels), access control strategy (defining role-based access control model and clarifying the access permissions of different roles to different data assets), and data leakage prevention strategy (defining rules for deploying DLP at network boundaries, terminals, and the cloud to identify and prevent unauthorized transmission of sensitive data).

[0101] The system simultaneously analyzes data isolation events (occurring risks) in S152 and data confidentiality events in S141 (all normal model-data interactions). These two types of events together constitute a complete behavior log. Based on this, the system determines a second layer of confidentiality, which is a dynamic set of monitoring and governance strategies built around the entire data lifecycle behavior. Its goal is to ensure that all operations on the data are visible, controllable, and traceable.

[0102] Its core content includes a comprehensive audit and logging strategy (defining the types, formats, storage periods, and audit frequencies of logs that need to be recorded), a behavior monitoring and anomaly detection strategy (defining baseline models and anomaly detection algorithms for user and entity behavior analysis to discover potential insider threats or compromised accounts), a compliance reporting strategy (defining the compliance report templates and generation processes that must be generated regularly to meet the requirements of regulations such as GDPR and HIPAA), and an incident response and event management strategy (defining a complete incident response process, roles, responsibilities, and communication mechanisms from incident discovery, analysis, and handling to post-incident review).

[0103] The system organically and deeply integrates the first layer of security (static defense) and the second layer of security (dynamic governance) defined in the previous two steps to construct an interconnected, mutually reinforcing, and three-dimensional multi-layered security system. This multi-layered security system has the characteristics of defense in depth (attackers need to break through the encryption, isolation, and access control of the first layer of the system, while also evading the monitoring and auditing of the second layer of the system, making the attack difficulty increase exponentially), linkage response characteristics (there is a linkage mechanism between the two layers of the system; for example, after the anomaly detection engine of the second layer of the system detects abnormal behavior, it can immediately trigger the network isolation policy of the first layer of the system, achieving a second-level response from detection to blocking), and self-evolution characteristics (the system will learn from each data isolation event and automatically adjust and optimize the policy parameters in the two layers of the system).

[0104] Please see Figure 7 , Figure 7 This is a schematic diagram of the structural composition of the client-side data confidentiality system in an embodiment of the present invention; the client-side data confidentiality system includes: The pseudo dataset module 21 is used to collect target layer features and corresponding models from various clients, and to determine the corresponding pseudo dataset based on the integration of target layer features and corresponding models. The global model module 22 is used to determine the corresponding heterogeneous features based on the identification of pseudo datasets, and to build the corresponding global model based on the heterogeneous features and the corresponding client. The personalized model module 23 is used to determine the neurons of each target layer based on the recognition of the global model, determine the local feature distribution of each client based on the neurons of each target layer and the corresponding transmission plan, and construct the corresponding personalized model based on the local feature distribution of each client and the personalized factors corresponding to the client. The dynamic security system module 24 is used to embed the personalized model into the corresponding client, construct the corresponding data security event based on the personalized model and the client's data, and determine the client's dynamic security system based on the data security event, the importance level of the client's data, and the client's previous security events. The multi-layered security system module 25 is used to identify multiple dynamic security items based on the identification of the client's dynamic security system, construct corresponding data isolation events based on the multiple dynamic security items and the data distribution system corresponding to the client, and determine the multi-layered security system of the data based on the data isolation events, data security events, and the client.

[0105] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A method for client-side data confidentiality, characterized in that, include: Collect target layer features and corresponding models from various clients, and determine the corresponding pseudo dataset based on the integration of target layer features and corresponding models; Based on the identification of pseudo datasets, the corresponding heterogeneous features are determined, and a corresponding global model is constructed based on the heterogeneous features and the corresponding client. The neurons of each target layer are determined based on the recognition of the global model. The local feature distribution of each client is determined based on the neurons of each target layer and the corresponding transmission plan. The corresponding personalized model is constructed based on the local feature distribution of each client and the personalized factors corresponding to the client. The personalized model is directly deployed to the client. Due to the uniqueness of its internal knowledge representation, it cannot be used to infer information from any other client. A personalized model is embedded into the corresponding client. Based on the personalized model and the client's data, a corresponding data confidentiality event is constructed. The dynamic confidentiality system of the client is determined according to the data confidentiality event, the importance level of the client's data, and the client's previous confidentiality events. Based on the identification of the client's dynamic confidentiality system, multiple dynamic confidentiality items are determined. Based on the multiple dynamic confidentiality items and the data distribution system corresponding to the client, a corresponding data isolation event is constructed. Based on the data isolation event, the data confidentiality event, and the client, a multi-layered confidentiality system for the data is determined. The multi-layered security system features defense-in-depth, coordinated response, and self-evolution.

2. The client-side data confidentiality method according to claim 1, characterized in that, The process of collecting target layer features and corresponding models from various clients, and determining the corresponding pseudo dataset based on the integration of target layer features and corresponding models, includes: Each client is labeled, and the corresponding feature set is determined based on the detection of each client. The target layer features and the corresponding model are determined by traversing the feature set, so as to collect the target layer features and the corresponding model. The matching coefficient is determined based on the matching between the target layer features and the corresponding model. The corresponding integration method is determined based on the mapping relationship between the matching coefficient, the target layer features, and the integration method. The corresponding pseudo dataset is determined based on the integration method, the target layer features, and the corresponding model.

3. The client-side data confidentiality method according to claim 1, characterized in that, The process of identifying corresponding heterogeneous features based on pseudo-datasets, and constructing a corresponding global model based on these heterogeneous features and the corresponding client, includes: Collect a pseudo dataset, identify multiple sub-heterogeneous data based on the detection of the pseudo dataset, determine the corresponding heterogeneous features based on the multiple sub-heterogeneous data, and mark the feature location and corresponding feature shape of the heterogeneous features. The first layer of global content is determined based on the feature location of the heterogeneous feature and the corresponding client. The second layer of global content is determined based on the feature shape of the heterogeneous feature and the corresponding client. A corresponding global model is constructed based on the first layer of global content, the second layer of global content, and the overall shape of the client.

4. The client-side data confidentiality method according to claim 1, characterized in that, The process of determining neurons in each target layer based on the recognition of the global model, determining the local feature distribution of each client based on the neurons in each target layer and the corresponding transmission plan, and constructing a corresponding personalized model based on the local feature distribution of each client and the personalized factors corresponding to that client includes: The global model is labeled, and the corresponding global distribution system is determined based on the recognition of the global model. Multiple target layers are determined based on the detection of the global distribution system, and the neurons of each target layer are labeled.

5. The client-side data confidentiality method according to claim 4, characterized in that, The process of determining neurons in each target layer based on the recognition of the global model, determining the local feature distribution of each client based on the neurons in each target layer and the corresponding transmission plan, and constructing a corresponding personalized model based on the local feature distribution of each client and the personalized factors corresponding to that client, further includes: Based on the detection of the client, previous transmission events are determined, and corresponding transmission plans are determined based on the identification of these previous transmission events. The local feature distribution of each client is determined based on the neurons of each target layer and the corresponding transmission plans. Collect the personalized factors corresponding to the client, and determine the construction of the corresponding personalized model based on the local feature distribution of each client and the corresponding personalized factors.

6. The client-side data confidentiality method according to claim 1, characterized in that, The personalized model is embedded into the corresponding client. Based on the personalized model and the client's data, a corresponding data confidentiality event is constructed. A dynamic confidentiality system for the client is determined based on this data confidentiality event, the importance level of the client's data, and the client's past confidentiality events, including: A personalized model is collected and input into the corresponding client to achieve client embedding. At the same time, the data of the client is determined based on the detection of the client, and a corresponding data confidentiality event is constructed based on the personalized model and the data of the client.

7. The client-side data confidentiality method according to claim 6, characterized in that, The personalized model is embedded into the corresponding client. Based on the personalized model and the client's data, a corresponding data confidentiality event is constructed. The client's dynamic confidentiality system is determined based on this data confidentiality event, the importance level of the client's data, and the client's past confidentiality events. The system also includes: Collect client data and mark the importance level of the client data. Determine the first level of dynamic confidentiality content based on the confidentiality event and the importance level of the client data. The second layer of dynamic security content is determined based on the data security incident and the client's previous security incidents. The dynamic security system of the client is then determined based on the first and second layers of dynamic security content.

8. The client-side data confidentiality method according to claim 1, characterized in that, The identification of the dynamic security system based on the client determines multiple dynamic security items. Based on these dynamic security items and the data distribution system corresponding to the client, a corresponding data isolation event is constructed. Based on this data isolation event, the data security event, and the client, a multi-layered security system for the data is determined, including: The client's dynamic security system is monitored in real time. Multiple dynamic surface markers are determined based on the detection of the client's dynamic security system. The corresponding dynamic security items are determined based on the tracing of each dynamic surface marker, so as to collect multiple dynamic security items. Based on the client's system identification, the data distribution system corresponding to the client is determined, and based on the data distribution system corresponding to the client and multiple dynamic confidentiality projects, corresponding data isolation events are constructed.

9. The client-side data confidentiality method according to claim 8, characterized in that, The identification of multiple dynamic security items based on the client's dynamic security system, the construction of corresponding data isolation events based on the multiple dynamic security items and the data distribution system corresponding to the client, and the multi-layered security system based on the data isolation events, data security events, and the client to determine the data further include: Based on the data isolation event and the client, a first layer of data security is determined. Based on the data isolation event and the data security event, a second layer of data security is determined. Based on the first layer of security and the second layer of security, a multi-layered data security system is constructed.

10. A client-side data confidentiality system, characterized in that, The client-side data confidentiality system is applied to the client-side data confidentiality method as described in any one of claims 1-9, and the client-side data confidentiality system includes: The pseudo dataset module is used to collect target layer features and corresponding models from various clients, and to determine the corresponding pseudo dataset based on the integration of target layer features and corresponding models. The global model module is used to determine the corresponding heterogeneous features based on the identification of pseudo datasets, and to build the corresponding global model based on the heterogeneous features and the corresponding client. The personalized model module is used to determine the neurons of each target layer based on the recognition of the global model, determine the local feature distribution of each client based on the neurons of each target layer and the corresponding transmission plan, and construct the corresponding personalized model based on the local feature distribution of each client and the personalized factors corresponding to the client. The dynamic security system module is used to embed personalized models into corresponding clients, construct corresponding data security events based on the personalized models and the client's data, and determine the client's dynamic security system based on the data security events, the importance level of the client's data, and the client's previous security events. The multi-layered security system module is used to identify multiple dynamic security items based on the client's dynamic security system. It constructs corresponding data isolation events based on the multiple dynamic security items and the data distribution system corresponding to the client. Based on the data isolation events, data security events, and the client, it determines the multi-layered security system of the data.