Data space-oriented data processing method and device and electronic equipment

By introducing an intermediate node for hierarchical fusion in federated clustering and dynamically adjusting the privacy budget by combining hierarchical weights and node sensitivity, the problems of ignoring hierarchical semantic relationships and fixed budgets in existing technologies are solved, resulting in more efficient clustering results with higher stability and accuracy.

CN121743914APending Publication Date: 2026-03-27GRG BANKING IT
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing federated clustering methods ignore hierarchical semantic relationships in data organization, resulting in clustering results lacking structural constraints and semantic associations. Furthermore, fixed-difference privacy budgets struggle to balance clustering accuracy with the strength of privacy protection.

Method used

A multi-layer federated clustering model is constructed by introducing an intermediate end. The intermediate end receives clustering results from the client and the global end and performs hierarchical fusion. The privacy budget is dynamically adjusted by combining hierarchical weights and node data sensitivity. A credibility coefficient is used for weighted fusion to optimize the stability and accuracy of the clustering results.

Benefits of technology

It achieves hierarchical fusion of cross-domain data, ensures the preservation of correlation information between data layers, balances privacy protection and clustering accuracy, improves the stability and accuracy of clustering results, and adapts to multi-layer data structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121743914A_ABST
    Figure CN121743914A_ABST
Patent Text Reader

Abstract

The invention discloses a data space-oriented data processing method and apparatus, and an electronic device, and belongs to the field of federal clustering. The data space-oriented data processing method comprises the following steps: receiving first information corresponding to a target round sent by each client; the first information at least comprises a first clustering result corresponding to the target round; a middle end clustering model corresponding to the target round is adopted to train the first clustering result, and a second clustering result corresponding to the target round is obtained; sending second information corresponding to the target round to the global end; and receiving third information corresponding to the target round sent by the global end. According to the data space-oriented data processing method provided by the invention, through hierarchical aggregation, the stability and precision of a clustering result are improved, and the method can better adapt to a multi-layer data structure, so that the effect of clustering analysis is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of federated clustering, and in particular relates to a data processing method, apparatus and electronic device for data space. Background Technology

[0002] With the continuous development of the digital economy, data has become a crucial element supporting intelligent decision-making. To achieve cross-institutional data sharing and joint analysis, federated learning has gradually become a mainstream privacy-preserving computing framework. It enables joint modeling of multi-source data by retaining data locally at each participating node and only uploading processed model parameters. However, in related technologies, federated clustering methods often employ a single-layer aggregation structure, ignoring the hierarchical semantic relationships in real-world data organization, resulting in clustering results lacking structural constraints and semantic connections. Furthermore, existing methods typically employ a fixed differential privacy budget, applying a uniform level of noise to all nodes, making it difficult to balance clustering accuracy and privacy protection strength, thus affecting the overall model's clustering accuracy. Summary of the Invention

[0003] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a data processing method, apparatus, and electronic device oriented towards data space, which improves the stability and accuracy of clustering results, can better adapt to multi-layer data structures, and thus optimizes the effect of clustering analysis.

[0004] In a first aspect, this application provides a data processing method oriented towards a data space, applied to an intermediate terminal, wherein the intermediate terminal corresponds to at least one client and is connected to a global terminal; the method includes: The system receives first information corresponding to a target wheel sent by each of the clients; the first information includes at least a first clustering result corresponding to the target wheel; the first clustering result is obtained by the client using the client clustering model corresponding to the target wheel and clustering based on local data. The intermediate clustering model corresponding to the target wheel is used to train the first clustering result to obtain the second clustering result corresponding to the target wheel; The second information corresponding to the target wheel is sent to the global terminal; the second information includes at least the second clustering result corresponding to the target wheel; the second clustering result is used by the global terminal to train the global terminal clustering model corresponding to the target wheel to obtain the third clustering result corresponding to the target wheel; Receive the third information corresponding to the target wheel sent by the global terminal.

[0005] According to the data processing method for data space in this application, by introducing an intermediate end, a multi-layer federated clustering model is formed with the client and the global end, realizing the hierarchical fusion of cross-domain data. The intermediate end receives the clustering results uploaded by the client and performs district-level hierarchical clustering fusion, effectively combining the data characteristics of each layer, ensuring that the correlation information between data layers is preserved, and achieving a balance between privacy protection and clustering accuracy. At the same time, through hierarchical aggregation, the stability and accuracy of clustering results are improved, which can better adapt to multi-layer data structures, thereby optimizing the effect of clustering analysis.

[0006] According to one embodiment of this application, the step of using the intermediate clustering model corresponding to the target wheel to train the first clustering result and obtain the second clustering result corresponding to the target wheel includes: Based on the hierarchical weights corresponding to the intermediate nodes and the data sensitivity of the nodes in the first clustering result, the privacy budget of each node is determined. The credibility coefficient of each node is determined based on the clustering quality of the nodes, the noise perturbation intensity, and the privacy budget. Based on the credibility coefficient, each node is weighted and fused to obtain the first clustering result after weighted fusion; Based on the first clustering result after weighted fusion, clustering is performed to obtain the second clustering result corresponding to the target round.

[0007] According to one embodiment of this application, after receiving the third information corresponding to the target wheel sent by the global terminal, the method further includes: Based on the third information, the privacy budget of the intermediate node corresponding to the next round of the target round is determined.

[0008] Secondly, this application provides a data processing method oriented towards a data space, applied to a client, wherein the client corresponds to an intermediate terminal, and the client is connected to the intermediate terminal. The method includes: The client clustering model corresponding to the target round is used and trained based on its own node data to obtain the first clustering result corresponding to the target round; The first information corresponding to the target wheel is sent to the intermediate terminal; the first information includes at least the first clustering result corresponding to the target wheel. Receive the second information corresponding to the target wheel sent by the intermediate terminal and the third information corresponding to the target wheel sent by the global terminal.

[0009] According to the data processing method for data space in this application, by introducing an intermediate end, a multi-layer federated clustering model is formed with the client and the global end, realizing the hierarchical fusion of cross-domain data. The intermediate end receives the clustering results uploaded by the client and performs district-level hierarchical clustering fusion, effectively combining the data characteristics of each layer, ensuring that the correlation information between data layers is preserved, and achieving a balance between privacy protection and clustering accuracy. At the same time, through hierarchical aggregation, the stability and accuracy of clustering results are improved, which can better adapt to multi-layer data structures, thereby optimizing the effect of clustering analysis.

[0010] According to one embodiment of this application, the step of using a client clustering model corresponding to the target round, training it based on its own node data, and obtaining the first clustering result corresponding to the target round includes: Based on the hierarchical weight corresponding to the client and the data sensitivity of its own node data, the privacy budget of each of its own nodes is determined; Based on the data sensitivity and the privacy budget, each of its own nodes is clustered to obtain the first clustering result corresponding to the target round.

[0011] According to one embodiment of this application, after receiving the second information corresponding to the target wheel sent by the intermediate terminal and the third information corresponding to the target wheel sent by the global terminal, the method further includes: Based on the second information and the third information, the privacy budget of each node in the client corresponding to the next round of the target round is determined.

[0012] Thirdly, this application provides a data processing method oriented towards a data space, applied to a global endpoint, wherein the global endpoint corresponds to at least one intermediate endpoint, and the global endpoint is connected to the intermediate endpoint. The method includes: The intermediate terminal receives second information corresponding to the target wheel sent by each of the intermediate terminals; the second information includes at least the second clustering result corresponding to the target wheel; the second clustering result is obtained by the intermediate terminal using the intermediate terminal clustering model corresponding to the target wheel and training based on the first clustering result; The second clustering result is trained using the global end clustering model corresponding to the target wheel to obtain the third clustering result corresponding to the target wheel; The third information corresponding to the target round is sent to the client and the intermediate end; the third information is used to determine the privacy budget of each node in the intermediate end and the client corresponding to the next round of the target round.

[0013] According to the data processing method for data space in this application, by introducing an intermediate end, a multi-layer federated clustering model is formed with the client and the global end, realizing the hierarchical fusion of cross-domain data. The intermediate end receives the clustering results uploaded by the client and performs district-level hierarchical clustering fusion, effectively combining the data characteristics of each layer, ensuring that the correlation information between data layers is preserved, and achieving a balance between privacy protection and clustering accuracy. At the same time, through hierarchical aggregation, the stability and accuracy of clustering results are improved, which can better adapt to multi-layer data structures, thereby optimizing the effect of clustering analysis.

[0014] According to one embodiment of this application, after training the second clustering result using the global end-clustering model corresponding to the target wheel and obtaining the third clustering result corresponding to the target wheel, the method further includes: Based on the clustering quality of the target wheel and the previous wheel of the target corresponding to the global end clustering model, the clustering accuracy gain corresponding to the target wheel is determined. Based on the privacy consumption rate corresponding to the target round and the clustering accuracy gain, the privacy budget of the node corresponding to the next round of the target round is determined.

[0015] Fourthly, this application provides a data processing apparatus oriented towards a data space, applied to an intermediate terminal, wherein the intermediate terminal corresponds to at least one client and is connected to a global terminal, the apparatus comprising: A first processing module is configured to receive first information corresponding to a target wheel sent by each of the clients; the first information includes at least a first clustering result corresponding to the target wheel; the first clustering result is obtained by the client using the client clustering model corresponding to the target wheel and clustering based on local data; The second processing module is used to train the first clustering result using the intermediate clustering model corresponding to the target wheel, and obtain the second clustering result corresponding to the target wheel. The third processing module is used to send the second information corresponding to the target wheel to the global terminal; the second information includes at least the second clustering result corresponding to the target wheel; the second clustering result is used by the global terminal to train the global terminal clustering model corresponding to the target wheel to obtain the third clustering result corresponding to the target wheel; The fourth processing module is used to receive the third information corresponding to the target wheel sent by the global terminal.

[0016] According to the data processing device for data space of this application, by introducing an intermediate end, a multi-layer federated clustering model is formed with the client and the global end, realizing the hierarchical fusion of cross-domain data. The intermediate end receives the clustering results uploaded by the client and performs district-level hierarchical clustering fusion, effectively combining the data characteristics of each layer, ensuring that the correlation information between data layers is preserved, and achieving a balance between privacy protection and clustering accuracy. At the same time, through hierarchical aggregation, the stability and accuracy of clustering results are improved, which can better adapt to multi-layer data structures, thereby optimizing the effect of clustering analysis.

[0017] Fifthly, this application provides a data processing apparatus for a data space, applied to a client, wherein the client corresponds to an intermediate terminal, and the client is connected to the intermediate terminal. The apparatus includes: The fifth processing module is used to train the client clustering model corresponding to the target round based on its own node data to obtain the first clustering result corresponding to the target round; The sixth processing module is used to send the first information corresponding to the target wheel to the intermediate terminal; the first information includes at least the first clustering result corresponding to the target wheel; The seventh processing module is used to receive the second information corresponding to the target wheel sent by the intermediate terminal and the third information corresponding to the target wheel sent by the global terminal.

[0018] Sixthly, this application provides a data processing apparatus oriented towards a data space, applied to a global endpoint, wherein the global endpoint corresponds to at least one intermediate endpoint, and the global endpoint is correspondingly connected to the intermediate endpoint; the apparatus includes: The eighth processing module is used to receive second information corresponding to the target wheel sent by each of the intermediate terminals; the second information includes at least the second clustering result corresponding to the target wheel; the second clustering result is obtained by the intermediate terminal using the intermediate terminal clustering model corresponding to the target wheel and training based on the first clustering result; The ninth processing module is used to train the second clustering result using the global end clustering model corresponding to the target wheel, and obtain the third clustering result corresponding to the target wheel; The tenth processing module is used to send the third information corresponding to the target round to the client and the intermediate end; the third information is used to determine the privacy budget of each node in the intermediate end and the client corresponding to the next round of the target round.

[0019] In a seventh aspect, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the data processing method oriented towards data space as described in the first, second, or third aspects above.

[0020] Fourthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data processing method oriented towards data space as described in the first, second, or third aspects above.

[0021] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the data processing method oriented towards data space as described in the first, second, or third aspects above.

[0022] The above-described one or more technical solutions in the embodiments of this application have at least one of the following technical effects: By introducing an intermediate layer to form a multi-layer federated clustering model with the client and global layers, hierarchical fusion of cross-domain data is achieved. The intermediate layer receives the clustering results uploaded by the client and performs district-level hierarchical clustering fusion, effectively combining the data characteristics of each layer and ensuring that the correlation information between data layers is preserved. A balance is achieved between privacy protection and clustering accuracy. At the same time, through hierarchical aggregation, the stability and accuracy of clustering results are improved, which can better adapt to multi-layer data structures, thereby optimizing the effect of clustering analysis.

[0023] Furthermore, by combining hierarchical weights and node data sensitivity, the privacy budget of each node is determined. Using the clustering quality, noise perturbation intensity, and privacy budget, a credibility coefficient is calculated and applied to perform weighted fusion of the clustering results. This weighted fusion optimizes the contributions of nodes at different levels, ensuring a balance between privacy protection and clustering accuracy, thus improving the robustness and accuracy of the clustering results. This credibility assessment mechanism comprehensively considers the stability of clustering quality, noise perturbation, and privacy budget, improving the stability and accuracy of the clustering results. It can better adapt to multi-layered data structures, thereby optimizing the effect of clustering analysis.

[0024] Furthermore, according to the data processing method for data space provided in the embodiments of this application, by introducing clustering accuracy gain as the basis for dynamic allocation of privacy budget, monitoring the change in clustering accuracy in each iteration, calculating the clustering accuracy gain, and adjusting the privacy budget of subsequent iterations based on the gain and privacy consumption rate, the privacy budget is reasonably allocated under the premise of achieving privacy protection, thereby improving the stability and accuracy of clustering results, better adapting to multi-layer data structures, and thus optimizing the effect of clustering analysis.

[0025] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0026] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is one of the flowcharts illustrating a data processing method for data space provided in an embodiment of this application; Figure 2 This is a second schematic flowchart of a data processing method for data space provided in an embodiment of this application; Figure 3 This is the third flowchart illustrating the data processing method for data space provided in the embodiments of this application; Figure 4 This is the fourth flowchart of the data processing method for data space provided in the embodiments of this application; Figure 5 This is the fifth flowchart illustrating the data processing method for data space provided in the embodiments of this application; Figure 6 This is one of the structural schematic diagrams of a data processing device oriented towards data space provided in the embodiments of this application; Figure 7 This is a second schematic diagram of the structure of the data processing device for data space provided in the embodiments of this application; Figure 8 This is the third schematic diagram of the structure of the data processing device for data space provided in the embodiments of this application; Figure 9 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0028] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0029] The following description, in conjunction with the accompanying drawings, details the data processing method, data processing device, electronic device, and readable storage medium for data space provided in this application, through specific embodiments and application scenarios.

[0030] Among them, the data processing method oriented towards the data space can be applied to the terminal, specifically executed by the hardware or software in the terminal.

[0031] The data processing method for data space provided in this application embodiment can be executed by an electronic device or a functional module or entity in an electronic device that can implement the data processing method for data space. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras and wearable devices. The data processing method for data space provided in this application embodiment is described below using an electronic device as the execution subject.

[0032] It should be noted that the data processing method for data space in this application is applied to a federated clustering scenario.

[0033] The federated clustering scenario consists of a three-layer structure formed by multiple clients participating in the training, multiple intermediate terminals, and a global terminal; among them, the training data corresponding to each client may be different, and the training data corresponding to each intermediate terminal may be different.

[0034] In this application, references Figure 4 The client corresponds to the data storage and processing node at the institutional layer, and for ease of description, it will be referred to as an institutional node below. Institutional nodes can be application terminals, information systems, data service units, or edge computing nodes with local data processing capabilities, such as hospitals, enterprise information systems, or traffic sensing terminals. Institutional nodes are mainly used to perform data collection, preprocessing, and differential privacy clustering calculations locally, and provide the perturbed first clustering result to the upper-level nodes. It can be understood that one client can correspond to one institutional node, and multiple clients can correspond one-to-one with multiple institutional nodes; one client corresponds to one middleware, and the client connects to the middleware.

[0035] Continue to refer to Figure 4The intermediate nodes correspond to the district-level data aggregation and processing nodes, hereinafter referred to as regional nodes. These regional nodes can be regional data management platforms, such as regional health commissions or urban transportation bureaus, providing service nodes with regional data coordination capabilities. Regional nodes receive cluster centers and statistical summaries uploaded by clients within their jurisdiction and perform regional-level cluster fusion, reliability assessment, and other intermediate-level data processing tasks. It can be understood that one intermediate node corresponds to one regional node, and multiple intermediate nodes correspond one-to-one with multiple regional nodes; each intermediate node corresponds to at least one client and connects to the global endpoint.

[0036] Continue to refer to Figure 4 The global endpoint corresponds to the comprehensive aggregation processing node at the city level, hereinafter referred to as the global node. The global node can be an upper-level service node with full-domain model integration capabilities, such as a city-level data management platform or a city-level operation center. The global node receives cross-regional aggregation results from multiple regional nodes and performs global clustering set, global model optimization, and outputs city-level clustering results. It can be understood that the global endpoint corresponds to at least one intermediate endpoint.

[0037] This application focuses on the circulation of urban data elements as the core scenario. By introducing an intermediary between the client and the global end, a multi-layered federated clustering and differential privacy collaborative data security analysis framework is constructed. The entire framework consists of three layers: client (corresponding to the institutional layer), intermediary (corresponding to the district level), and global end (corresponding to the city level). Within a multi-level data governance system, it achieves the organic integration of privacy protection, cluster collaboration, and model feedback. By constructing a closed-loop structure of "hierarchical autonomy—collaborative aggregation—adaptive feedback," this framework is applicable to the multi-level data organization structures commonly found in practical data management systems (such as "institution—region—city"). It can effectively capture the hierarchical relationship information between multi-level data, breaking through the limitations of traditional single-layer federated clustering in hierarchical relationship modeling. It provides a scalable and dynamically adaptable implementation path for cross-domain secure intelligent analysis in the data space.

[0038] Understandably, since different levels have different impacts on the clustering results in terms of aggregation scope, data sensitivity, and upload noise intensity, the level weights and privacy budgets of different levels will be different in the subsequent clustering fusion process.

[0039] In some embodiments, conventional clustering algorithms can be used for training, such as DP-KMeans, DP-DBSCAN, etc., which are not limited herein.

[0040] In some embodiments, the hierarchical weights may include the hierarchical weights corresponding to the client, intermediate, and global ends, and can be represented as follows: , .

[0041] In this embodiment, For hierarchical weights, Corresponding to the client's hierarchical weight, Corresponding to the hierarchical weights of the intermediate end, This corresponds to the hierarchical weights of the global endpoint. Specifically, the hierarchical weight of the global endpoint is greater than that of the middleware endpoint, which in turn is greater than that of the client endpoint. The hierarchical weight of the middleware endpoint is greater than that of the client endpoint. > > .

[0042] In some embodiments, the privacy budget may include privacy budgets for the client, middleware, and global endpoints, and can be represented as follows: .

[0043] In this embodiment, Budget the local clustering perturbation for node i in the client, which is used for feature noise addition; Budget for perturbation of the upload cluster center in the intermediate stage; The global feedback budget is used for dynamic adjustments in subsequent iterations, allowing for training strategy adjustments such as model optimization or privacy budget reallocation.

[0044] In some embodiments, before performing steps 110-140, 210-230, and 310-340, refer to Figure 5 Node registration and privacy budget initialization need to be completed to ensure that the system has a trusted communication link and controllable privacy protection capabilities in a multi-level federated clustering mode.

[0045] In some embodiments, the nodes in the client, middleware, and global servers complete the registration in a hierarchical manner.

[0046] In this embodiment, the client generates a public-private key pair locally and submits a registration request to its superior node (i.e., the corresponding intermediate end). The registration request includes organizational metadata, data type summary, and privacy capability statement, etc. The middleware verifies the client's legitimacy through public key verification and digital credential registration forms (or smart contracts), and records the registration result in the trusted identity chain. Simultaneously, as a lower-level node, the middleware also needs to submit registration information to the global endpoint for identity verification. The global endpoint maintains the identity directory and trusted links of all nodes in the system to ensure that each node has a legitimate and verifiable identity during subsequent clustering processes.

[0047] like Figure 1As shown, the data processing method oriented towards the data space is applied to the middle end, which corresponds to at least one client and is connected to the global end. The data processing method oriented towards the data space includes steps 110, 120, 130 and 140.

[0048] Step 110: Receive the first information corresponding to the target wheel sent by each client; In this step, the first information includes at least the first clustering result corresponding to the target round; the target round is any round in the model training process, including the first round, the last round, and intermediate rounds.

[0049] The first clustering result is obtained by clustering local data using the client-side clustering model corresponding to the target round. The client-side clustering model is the model trained by each client.

[0050] In actual execution, in round (i-1), the intermediate end receives the first information obtained by each client through training. The first information includes the first clustering result. The intermediate end clustering model is used to train the first clustering result to obtain the second clustering result corresponding to round (i-1). The second information corresponding to round (i-1) is sent to each client and the global end for the global end to perform round i training and for the clients to adjust the model parameters in round i. Where i > 1 and i is an integer. When i is 2, round (i-1) is the initial training round, and the second information includes at least the second clustering result.

[0051] The global endpoint receives the second information obtained from the training of each intermediate endpoint, which includes the second clustering result. The global endpoint clustering model is used to train the second clustering result to obtain the third clustering result corresponding to the (i-1)th round. The global endpoint evaluates the third clustering result to obtain the evaluation result corresponding to the (i-1)th round. The global endpoint feeds back the evaluation result corresponding to the (i-1)th round to the intermediate endpoint and the client to guide the intermediate endpoint clustering model and the client clustering model to adjust the model parameters in the i-th round and to redistribute the privacy budget in the intermediate endpoint and the client in the i-th round.

[0052] In actual implementation, refer to Figure 4 The intermediate end, acting as the district-level layer, receives the cluster center set sent by the institutional layer corresponding to the client within its jurisdiction, which is the result of local clustering by multiple institutional nodes. The intermediate end calculates the weighted distance matrix of the cluster centers within the region, performs regional aggregation to obtain the regional-level cluster center result set, and sends it to the global end, which acts as the city layer.

[0053] In some embodiments, the result set of regional cluster centers can be obtained based on the institutional-level cluster center set, which can be expressed by the following formula:

[0054] in, This is the result set of regional cluster centers. The set of cluster centers at the institutional level contains i nodes. This indicates the aggregation operation in the middle.

[0055] In some embodiments, the first information may further include parameter information such as cluster centers, variance, or sample weights obtained by the client during training, which are used by the intermediate end to train based on the first clustering result.

[0056] Step 120: Use the intermediate clustering model corresponding to the target wheel to train the first clustering result and obtain the second clustering result corresponding to the target wheel; In this step, the intermediate end trains the first clustering result obtained in step 110 based on the intermediate end clustering model corresponding to the target wheel to obtain the second clustering result corresponding to the target wheel.

[0057] It should be noted that due to the different data distributions of each client, the first clustering results differ. During the training process, the intermediate end needs to perform weighted fusion of the first clustering results from different clients, resulting in different second clustering results.

[0058] After receiving the first information sent by each client, the intermediate terminal uses the intermediate terminal clustering model to train the first clustering result in the first information to obtain its own corresponding intermediate terminal clustering model and second clustering result.

[0059] Different intermediate clustering models obtained by training at different intermediate nodes will have different results.

[0060] The implementation of step 120 will be described below through specific embodiments.

[0061] In some embodiments, reference Figure 5 Step 120 may include: Based on the hierarchical weights corresponding to the intermediate nodes and the data sensitivity of the nodes in the first clustering result, the privacy budget of each node is determined. The credibility coefficient of each node is determined based on the clustering quality, noise perturbation intensity, and privacy budget. The nodes are weighted and fused based on the credibility coefficient to obtain the first clustering result after weighted fusion. Clustering is performed based on the first clustering result after weighted fusion to obtain the second clustering result corresponding to the target round.

[0062] The following uses the hierarchical weights corresponding to the middle end. and privacy budget For example, to determine the privacy budget of each node in the intermediate end. The process will be explained.

[0063] In some embodiments, the privacy budget for each node is determined based on the hierarchical weights corresponding to the intermediate nodes and the data sensitivity of the nodes in the first clustering result, which can be expressed by the formula:

[0064] in, For the first Privacy budget for each node. It is the hierarchical weight corresponding to the middle end; The first clustering result is the first Data sensitivity of each node; , These are weighting coefficients, which can be determined based on user-defined values ​​or experimental data.

[0065] In some embodiments, determining a node's data sensitivity based on its data value, content sensitivity, and re-identification risk can be expressed by the following formula:

[0066] in, For the first Data sensitivity of each node For the first The data value weight of each node For the first Content sensitivity coefficient of each node For the first Risk factors for re-identification of individual nodes. The function or model can be set according to the actual application requirements, such as a linear weighted model or a normalized weighted summation model, etc., and this application does not limit it.

[0067] It should be noted that the privacy budget of a node is negatively correlated with its data sensitivity. That is, the higher the data sensitivity, the smaller the privacy budget allocated to that node, in order to ensure that the sensitive data is subjected to greater perturbation during the federated clustering process, thereby enhancing the overall privacy protection capability.

[0068] In actual implementation, the privacy budget for each node is recorded in the privacy budget management table, and the budget consumption is checked regularly.

[0069] In some embodiments, if a node's privacy budget is lower than the privacy budget threshold, the system will automatically trigger a renegotiation mechanism, whereby its parent node will decide whether to increase the budget or decrease the clustering accuracy.

[0070] In some embodiments, the privacy budget for each node can also be determined based on factors such as the amount of data and the complexity of the data, without limitation.

[0071] In some embodiments, the credibility coefficient of each node is determined based on the clustering quality, noise perturbation intensity, and privacy budget, and can be expressed by the formula:

[0072] in, For the first The credibility coefficient of each node For the first The clustering quality of each node itself. For the first The node upload characteristics of each node or the noise standard deviation of the center. The maximum value among the noise standard deviations of the uploaded features or centers of each node. For the first Privacy budget deviation at each node This represents the average privacy budget for each node. and These are weighting coefficients, which can be determined based on user-defined criteria or experimental results.

[0073] In actual implementation, the system uses the node's credibility coefficient. A weighted average is applied to the cluster centers in the first clustering result to improve the stability and accuracy of the clustering results.

[0074] In some embodiments, clustering based on the first clustering result after weighted fusion to obtain the second clustering result corresponding to the target round can be achieved through the following steps: Step 1: Analyze the first clustering result after weighted fusion. Clustering, which yields local cluster centers, can be expressed by the formula:

[0075] Where i represents node i. Represents the sample set based on node i. The obtained cluster centers The set of cluster centers, k represents the number of cluster centers in the cluster center set; j is an integer, and .

[0076] Step 2, based on the Laplace mechanism, adds noise to each cluster center in the cluster center set, which can be expressed by the formula:

[0077] in, This represents the noisy clustering result (i.e., the second clustering result). Indicates the first Cluster centers of nodes, For Laplace function, For the first Data sensitivity of each node For the middle end of the first Privacy budget for each node.

[0078] In some embodiments, the second information sent by the intermediate end includes at least the second clustering result; or it may also include additional metadata such as timestamps, budget consumption rates, and clustering quality metrics.

[0079] In this embodiment, the intermediate end uploads the second information to the corresponding upper-layer node (i.e., the node in the global end) for aggregation. During the research and development process, the inventors discovered that related technologies typically employ a uniform differential privacy budget, applying the same noise level to all nodes. However, in reality, the amount and distribution of data across different nodes vary greatly (for example, small data nodes may cause severe distortion, while large data nodes, such as large hospitals, show significant differences in clustering results compared to small clinics). Using a fixed noise level cannot balance clustering accuracy and privacy protection strength, leading to excessive redundancy in privacy protection and consequently affecting clustering accuracy. Related technologies also lack an effective credibility assessment mechanism; when abnormal nodes upload erroneous clustering results or excessive noise, it can damage the clustering effect of the global model. Current federated frameworks fail to provide sufficient criteria to evaluate the contribution of nodes to the clustering results.

[0080] In this application, a privacy budget for each node is determined based on the intermediate hierarchical weights and the data sensitivity of nodes in the first clustering result, allowing the privacy budget to be dynamically adjusted according to data volume, sensitivity, and node importance. A node credibility coefficient is calculated by combining the node's clustering quality, noise perturbation intensity, and privacy budget. Based on the credibility coefficients, each node is weighted and fused to obtain the weighted first clustering result. The weighted fusion calculation effectively balances the influence of each node, enhances the contribution of nodes with higher credibility, and reduces interference from anomalous nodes. A second round of clustering is performed based on the weighted fusion result to obtain the final target round clustering result, improving the accuracy and stability of the clustering result while ensuring privacy. Weighted fusion reduces distortion caused by node data differences or anomalous nodes, enhancing the reliability of node contribution assessment.

[0081] According to the data processing method for data space provided in this application, the privacy budget of a node is determined by combining hierarchical weights and the data sensitivity of the node. A credibility coefficient is calculated and applied to weightedly fuse the clustering results using the clustering quality, noise perturbation intensity, and privacy budget of the node. This weighted fusion optimizes the contribution of nodes at different levels, ensuring a balance between privacy protection and clustering accuracy, and improving the robustness and accuracy of the clustering results. This credibility evaluation mechanism comprehensively considers the stability of clustering quality, noise perturbation, and privacy budget, improving the stability and accuracy of the clustering results, and better adapting to multi-layered data structures, thereby optimizing the effect of clustering analysis.

[0082] Step 130: Send the second information corresponding to the target wheel to the global terminal; In this step, the second information includes at least the second clustering result corresponding to the target round. The second clustering result is used for the global end to train the global end clustering model corresponding to the target round to obtain the third clustering result corresponding to the target round. The intermediate end sends the second information corresponding to the i-th round to each client and the global end, so that the client and the global end can perform the i-th round of training; where i>1 and i is an integer. When i is 2, the (i-1)-th round is the initial training round, and the second information includes at least the second clustering result.

[0083] The global endpoint receives the second information obtained from the training of each intermediate endpoint, which includes the second clustering result. The global endpoint clustering model is used to train the second clustering result to obtain the third clustering result corresponding to the (i-1)th round. The global endpoint evaluates the third clustering result to obtain the evaluation result corresponding to the (i-1)th round. The global endpoint feeds back the evaluation result corresponding to the (i-1)th round to the intermediate endpoint and the client to guide the intermediate endpoint clustering model and the client clustering model to adjust the model parameters in the i-th round and to redistribute the privacy budget in the intermediate endpoint and the client in the i-th round.

[0084] Different clients participating in the aggregation may produce different second clustering results.

[0085] In some embodiments, the intermediate terminal may also send the second information corresponding to the target wheel to the client.

[0086] In this embodiment, the second information sent by the intermediate end can optimize the client clustering model in the client in terms of both model accuracy and privacy budget, thereby achieving region-level clustering convergence.

[0087] Step 140: Receive the third information corresponding to the target wheel sent by the global terminal.

[0088] In this step, the third piece of information is the clustering feedback information obtained by the global end evaluating the global end clustering model, which is used by the client and the middle end to update the privacy budget and the corresponding model parameters, respectively.

[0089] In some embodiments, the third information may also include parameters such as cluster centers, variance, or sample weights obtained from global training.

[0090] In actual execution, each intermediate end receives the evaluation result corresponding to the (i-1)th round from the third information sent by the global end, repeats steps 120 to 130, and performs multiple rounds of training until the clustering quality indicators (such as clustering convergence or clustering accuracy) of the global end clustering model meet the requirements.

[0091] In some embodiments, after step 140, the method further includes: Based on the third information, determine the privacy budget of the intermediate node corresponding to the next round of the target round.

[0092] In this embodiment, the intermediate end receives the clustering quality feedback result sent by the global end, and adjusts the intermediate end clustering model parameters and privacy budget based on the clustering quality feedback result.

[0093] At the intermediate end, the system performs weighted fusion of the first clustering results sent by each client to form regional cluster centers and feature representations.

[0094] In some embodiments, techniques such as multi-party secure computation and privacy budget reallocation can also be used in the weighted fusion process to align the feature spaces of institutional models with large differences, so as to eliminate the problem of uneven distribution of data across institutions.

[0095] During their research, the inventors discovered that single-layer clustering methods in related technologies struggle to effectively represent the semantic relationships and spatial dependencies between data levels, leading to distorted clustering results. Real-world data organizations typically exhibit significant and complex hierarchical structures; for example, government, healthcare, and urban sensing networks all possess clearly defined hierarchical relationships. However, federated clustering methods in related technologies employ a single-layer structure, where all client nodes (such as hospitals, enterprises, and departments) perform clustering locally and upload the results to the server for global aggregation. This ignores the semantic relationships between data levels, resulting in aggregation results that fail to accurately reflect the hierarchical structure of the data.

[0096] In this application, by introducing an intermediary, the data is no longer flattened but processed hierarchically. After receiving the clustering results from the client, the intermediary performs region-level cluster fusion, effectively preserving the hierarchical relationships and structural information of the data. This method not only avoids neglecting hierarchical semantics in single-layer aggregation but also improves the accuracy and stability of the clustering results through region-level processing.

[0097] According to the data processing method for data space provided in the embodiments of this application, by introducing an intermediate layer, a multi-level federated clustering model is formed with the client and the global end, realizing the hierarchical fusion of cross-domain data. The intermediate end receives the clustering results uploaded by the client and performs regional-level clustering fusion, effectively combining the data characteristics of each level, ensuring that the correlation information between data levels is preserved, achieving a balance between privacy protection and clustering accuracy. At the same time, through hierarchical aggregation, the stability and accuracy of clustering results are improved, which can better adapt to multi-level data structures, thereby optimizing the effect of clustering analysis.

[0098] like Figure 2 As shown, this data processing method for data space is applied to a client, which corresponds to an intermediate terminal, and the client is connected to the intermediate terminal. The method includes steps 210, 220 and 230.

[0099] Step 210: Use the client clustering model corresponding to the target round, train it based on its own node data, and obtain the first clustering result corresponding to the target round; In this step, the client clustering model is the model trained by each client.

[0100] In the actual execution process, in the (i-1)th round, the intermediate end receives the first information obtained by each client through training. The first information includes the first clustering result. The intermediate end clustering model is used to train the first clustering result to obtain the second clustering result corresponding to the (i-1)th round. In some embodiments, step 210 includes: Based on the hierarchical weight of the client and the data sensitivity of its own node data, determine the privacy budget of each node. Cluster each node based on data sensitivity and privacy budget to obtain the first clustering result corresponding to the target round.

[0101] In this embodiment, the process of determining the privacy budget of each node and obtaining the first clustering result corresponding to the target round is consistent with the implementation process in step 120, which has been described in detail above, so it will not be repeated here.

[0102] Step 220: Send the first information corresponding to the target wheel to the intermediate end; In this practical step, the first information includes at least the first clustering result corresponding to the target round.

[0103] Step 230: Receive the second information corresponding to the target wheel sent by the intermediate end and the third information corresponding to the target wheel sent by the global end.

[0104] In this step, the client receives the clustering model evaluation results corresponding to the (i-1)th round from the third information of the global end and the second information of the intermediate end. Steps 220-230 are repeated to perform multiple rounds of training until the clustering quality indicators (such as clustering convergence or clustering accuracy) of the global end clustering model meet the requirements. Here, i > 1 and i is an integer. When i is 2, the (i-1)th round is the initial training round.

[0105] In some embodiments, after step 230, the method further includes: Based on the second and third information, determine the privacy budget for each node in the client corresponding to the next round of the target round.

[0106] In this embodiment, the client receives clustering quality feedback results sent by the intermediate end and the global end, and adjusts parameters such as the intermediate end clustering model parameters and privacy budget based on the clustering quality feedback results.

[0107] In actual implementation, refer to Figure 4 The client, acting as the institutional layer, performs local clustering based on each institutional node (such as medical institutions, enterprises, or sensor terminals). It also introduces an adaptive differential privacy mechanism, dynamically allocating a privacy budget based on data sensitivity and clustering complexity to ensure that its data is processed locally, preventing the leakage of original information. Furthermore, a perturbation mechanism effectively prevents reverse inference. The local clustering model parameters generated during this process are then encrypted and uploaded to the regional nodes.

[0108] like Figure 3 As shown, the data processing method oriented towards the data space is applied to the global end, the global end corresponds to at least one intermediate end, and the global end corresponds to the connection of the intermediate end. The data processing method oriented towards the data space includes: step 310, step 320 and step 330.

[0109] Step 310: Receive the second information corresponding to the target wheel sent by each intermediate end; In this step, the second information includes at least the second clustering result corresponding to the target wheel; the second clustering result is obtained by training the intermediate end clustering model corresponding to the target wheel based on the first clustering result. The first clustering result is obtained by clustering local data using the client clustering model corresponding to the target round.

[0110] Step 320: Use the global end clustering model corresponding to the target round to train the second clustering result and obtain the third clustering result corresponding to the target round; In actual execution, the global terminal, acting as the city level, receives the set of regional cluster centers contained in the second information sent by the intermediate terminal, performs global clustering and multi-level weighted clustering fusion, and obtains the third clustering result.

[0111] In some embodiments, the third clustering result obtained based on the set of regional cluster centers can be expressed by the following formula:

[0112] in, This is the result of the third clustering. , is the set of cluster centers at the institutional level, containing i nodes, and F represents the aggregation operation in the global end.

[0113] In some embodiments, step 320 includes: Based on the hierarchical weights corresponding to the global endpoint and the data sensitivity of nodes in the second clustering results, the privacy budget for each node is determined. The credibility coefficient of each node is determined based on the clustering quality, noise perturbation intensity, and privacy budget. The nodes are weighted and fused based on the credibility coefficient to obtain the second clustering result after weighted fusion. Clustering is performed based on the second clustering result after weighted fusion to obtain the third clustering result corresponding to the target round.

[0114] In this embodiment, the process of determining the privacy budget of each node and the credibility coefficient of each node is consistent with the implementation process in step 120, which has been described in detail above, so it will not be repeated here.

[0115] In actual execution, the second clustering result sent by the intermediate end is fused with global clustering and multi-level weighted clustering at the global end to obtain the third clustering result. Based on the global clustering model contained in the third clustering result, the stability and accuracy of the clustering result are evaluated according to the model convergence index (such as the inter-cluster distance convergence rate), and a privacy budget feedback signal is generated and sent to each node in the client and intermediate end to optimize the privacy budget allocation in the subsequent clustering process.

[0116] In some embodiments, reference Figure 5 The global endpoint can also connect to the intermediate endpoint and the client. After step 320, it may also include: Based on the clustering quality of the global end clustering model at the target round and the round before the target, the clustering accuracy gain corresponding to the target round is determined. Based on the privacy cost rate and clustering accuracy gain corresponding to the target round, the privacy budget of the node corresponding to the next round of the target round is determined.

[0117] In actual implementation, after each iteration, the system adjusts the privacy budget mechanical energy in the global, middleware, and client sides based on the clustering quality and privacy loss.

[0118] The following uses the t-th iteration as an example to illustrate the process of determining the clustering accuracy gain corresponding to the target round. Here, t > 1 and t is an integer; when t is 2, the (t-1)-th round is the initial training round.

[0119] In some embodiments, the privacy consumption rate corresponding to round t can be determined based on the following formula:

[0120] in, This represents the privacy consumption rate corresponding to round t. This represents the privacy budget already used in round t. This indicates the total privacy budget of the system.

[0121] In some embodiments, the privacy budget for the node corresponding to round (t+1) can be determined based on the following formula:

[0122] in, Let i represent the privacy budget for node i in round (t+1). This represents the privacy budget for node i in round t. This represents the clustering accuracy gain corresponding to round t; This represents the privacy consumption rate corresponding to round t, and α is an adjustable parameter that can be fine-tuned based on user customization or actual running results.

[0123] Let's continue with the t-th iteration as an example.

[0124] In some embodiments, the clustering accuracy gain corresponding to the target round is determined based on the clustering quality of the global end-clustering model at the target round and the round preceding the target, respectively, and can be expressed by the formula:

[0125] in, This represents the clustering accuracy gain corresponding to round t. and Representing the t-th round and the t-th round respectively The clustering accuracy corresponding to the round.

[0126] In some embodiments, clustering accuracy can be determined based on the following formula:

[0127] Where Q represents the clustering accuracy, a(i) is the average distance from node i to other samples in the same cluster, b(i) is the average distance from node i to the nearest other cluster, and n represents the total number of nodes.

[0128] In this embodiment, a higher clustering accuracy value indicates that the clustering result is closer to the true value. Clustering accuracy can be calculated from the silhouette coefficient, inter-cluster distance, intra-cluster variance, and clustering stability index in the local clustering metrics uploaded by each node.

[0129] In practice, the system dynamically adjusts the privacy budget by monitoring changes in clustering accuracy gain. In the early stages, before the model converges, a higher budget (i.e., adding less noise) is allowed to accelerate convergence. As the system enters the mid-stage and the model stabilizes, the budget is appropriately tightened to prevent privacy leaks. In the later stages, because the model is more sensitive to accuracy and has strong robustness to small perturbations, the budget can be slightly increased to balance clustering quality. This phased budget allocation mechanism not only ensures the constraint of the overall privacy budget (total ε remains constant) but also effectively achieves a balance between performance and privacy protection.

[0130] According to the data processing method for data space provided in the embodiments of this application, by introducing clustering accuracy gain as the basis for dynamic allocation of privacy budget, the clustering accuracy change in each iteration is monitored, the clustering accuracy gain is calculated, and the privacy budget of subsequent iterations is adjusted according to the gain and the privacy consumption rate. Under the premise of achieving privacy protection, the privacy budget is reasonably allocated, which improves the stability and accuracy of clustering results, can better adapt to multi-level data structures, and thus optimizes the effect of clustering analysis.

[0131] Step 330: Send the third information corresponding to the target wheel to the client and the middleware.

[0132] In this step, the third piece of information is used to determine the privacy budget of each node in the intermediate end and the client corresponding to the next round of the target round.

[0133] In actual implementation, refer to Figure 4 and Figure 5 The global endpoint, acting as the city layer, receives cluster centers from various regions sent by the intermediate endpoint, acting as the district layer. It then performs cross-regional global clustering to generate a unified clustering model. The clustering quality feedback module in the global endpoint evaluates the clustering profile coefficients, noise impact, and privacy loss function of the overall model. The feedback results are then transmitted to the intermediate endpoint (district layer) and the client (institutional layer) to guide the reallocation of privacy budgets and model adjustments in the intermediate endpoint and client.

[0134] The following sections use medical and urban transportation scenarios as examples to illustrate the implementation methods of using the client as the institutional layer, the middleware as the district-level layer, and the global layer as the city layer.

[0135] Taking the healthcare scenario as an example, the institutional layer consists of nodes such as hospitals and community health service centers. It is primarily responsible for conducting local cluster analysis based on local data (such as electronic medical records, laboratory indicators, and vital signs data), applied to patient disease classification and district-level public health risk assessment. This data is highly sensitive and requires strict privacy protection. The district-level layer consists of nodes such as district-level health commissions and district-level health platforms. It is responsible for receiving cluster centers and statistical summaries uploaded by hospitals, conducting district-level disease cluster analysis, and predicting epidemic transmission trends. The city-level layer consists of global nodes of the city-level health management platform. It is responsible for aggregating the clustering results from all district-level layers across the city, constructing a comprehensive data spatial model, and conducting city-wide cluster consistency analysis and privacy budget reallocation, providing support for city-wide public health decision-making and cross-regional data analysis.

[0136] Taking urban traffic scenarios as an example, the institutional layer consists of nodes such as bus companies and traffic police monitoring stations. It is responsible for analyzing road network congestion patterns and detecting traffic anomalies based on real-time traffic data (such as traffic flow, vehicle speed distribution, and travel heatmaps), requiring efficient real-time clustering processing capabilities. The district-level layer comprises multiple regional nodes from various branches of urban traffic management. It is responsible for aggregating traffic clustering results uploaded by traffic police stations and bus companies within its jurisdiction, performing multi-source fusion analysis, optimizing regional traffic conditions, and identifying cross-agency travel patterns and regional congestion. The city-level layer consists of global nodes from the city's traffic management platform. It is responsible for aggregating clustering results from all district-level layers, constructing a comprehensive city-wide data spatial model, performing global traffic analysis and privacy budget adjustments, and providing support for city-wide traffic flow optimization and traffic governance decisions.

[0137] The following section uses a medical scenario as an example to explain in detail how the data processing method oriented towards data space proposed in this application clusters medical and health data for the purpose of assisting diagnosis.

[0138] In the medical data space, different hospitals (institutional level) upload noisy cluster centers, regional health commissions (intermediate level) perform regional-level disease type clustering analysis, and the global model integrates disease patterns at the city level. Specifically, the institutional level outputs typical patient characteristic clusters (such as metabolic syndrome type, high-risk cardiovascular type, etc.) based on basic data and clustering models. When the district level finds that cluster centers from multiple hospitals exhibit similar characteristics, such as elevated BMI, blood glucose fluctuations, and increased inflammatory markers, the model can identify potential high-risk groups for diabetes and automatically provide corresponding "regional health intervention strategies" to the clinical decision-making system. Based on this, the city level aggregates the clustering model calculation results from multiple district levels and performs cross-regional cluster result alignment. This process further optimizes disease risk identification capabilities by merging, correcting, and recalibrating similar disease clusters. Based on the distribution density and evolution trend of clusters in each region, the global model can identify the growth or migration trends of disease groups and output city-level disease risk maps and trend curves. Furthermore, the privacy budget allocation at the district and institutional levels can be dynamically adjusted based on model stability and the contribution of regional data, and optimized parameters can be returned to the regional and institutional nodes for the next round of clustering calculations. This process continuously optimizes the results of urban healthcare trend analysis, thereby providing support for policy-making and the rational allocation of healthcare resources.

[0139] The data processing method for data space orientation provided in this application can be executed by a data processing device for data space orientation. This application uses an example of a data processing device for data space orientation executing the data processing method for data space orientation to illustrate the data processing device for data space orientation provided in this application.

[0140] This application also provides a data processing apparatus for data space.

[0141] like Figure 6 As shown, the data processing device for data space is applied to the middle end, which corresponds to at least one client and is connected to the global end. The device includes: a first processing module 610, a second processing module 620, a third processing module 630 and a fourth processing module 640.

[0142] The first processing module 610 is used to receive first information corresponding to the target round sent by each client; the first information includes at least the first clustering result corresponding to the target round; the first clustering result is obtained by the client using the client clustering model corresponding to the target round and clustering based on local data; The second processing module 620 is used to train the first clustering result using the intermediate clustering model corresponding to the target wheel, and obtain the second clustering result corresponding to the target wheel. The third processing module 630 is used to send the second information corresponding to the target wheel to the global end; the second information includes at least the second clustering result corresponding to the target wheel; the second clustering result is used by the global end to train the global end clustering model corresponding to the target wheel to obtain the third clustering result corresponding to the target wheel; The fourth processing module 640 is used to receive the third information corresponding to the target wheel sent by the global terminal.

[0143] According to the data processing apparatus for data space provided in the embodiments of this application, by introducing an intermediate layer, a multi-level federated clustering model is formed with the client and the global end, realizing the hierarchical fusion of cross-domain data. The intermediate end receives the clustering results uploaded by the client and performs regional-level clustering fusion, effectively combining the data characteristics of each level, ensuring that the correlation information between data levels is preserved, achieving a balance between privacy protection and clustering accuracy. At the same time, through hierarchical aggregation, the stability and accuracy of the clustering results are improved, which can better adapt to multi-level data structures, thereby optimizing the effect of clustering analysis.

[0144] In some embodiments, the second processing module 620 may also be used for: Based on the hierarchical weights corresponding to the intermediate nodes and the data sensitivity of the nodes in the first clustering result, the privacy budget of each node is determined. The credibility coefficient of each node is determined based on the clustering quality, noise perturbation intensity, and privacy budget. The nodes are weighted and fused based on the credibility coefficient to obtain the first clustering result after weighted fusion. Clustering is performed based on the first clustering result after weighted fusion to obtain the second clustering result corresponding to the target round.

[0145] In some embodiments, the apparatus may further include a twelfth processing module for: Based on the third information, determine the privacy budget of the intermediate node corresponding to the next round of the target round.

[0146] like Figure 7 As shown, the data processing device for data space is applied to the middle terminal, with each client corresponding to a middle terminal and the client corresponding to a connection to the middle terminal. The device includes: a fifth processing module 710, a sixth processing module 720, and a seventh processing module 730.

[0147] The fifth processing module 710 is used to train the client clustering model corresponding to the target round based on its own node data to obtain the first clustering result corresponding to the target round. The sixth processing module 720 is used to send the first information corresponding to the target wheel to the intermediate end; the first information includes at least the first clustering result corresponding to the target wheel; The seventh processing module 730 is used to receive the second information corresponding to the target wheel sent by the intermediate end and the third information corresponding to the target wheel sent by the global end.

[0148] According to the data processing apparatus for data space provided in the embodiments of this application, by introducing an intermediate layer, a multi-level federated clustering model is formed with the client and the global end, realizing the hierarchical fusion of cross-domain data. The intermediate end receives the clustering results uploaded by the client and performs regional-level clustering fusion, effectively combining the data characteristics of each level, ensuring that the correlation information between data levels is preserved, achieving a balance between privacy protection and clustering accuracy. At the same time, through hierarchical aggregation, the stability and accuracy of the clustering results are improved, which can better adapt to multi-level data structures, thereby optimizing the effect of clustering analysis.

[0149] In some embodiments, the fifth processing module 710 can also be used for: Based on the hierarchical weight of the client and the data sensitivity of its own node data, determine the privacy budget of each node. Cluster each node based on data sensitivity and privacy budget to obtain the first clustering result corresponding to the target round.

[0150] In some embodiments, the device may further include a thirteenth module for: Based on the second and third information, determine the privacy budget for each node in the client corresponding to the next round of the target round.

[0151] like Figure 8 As shown, the data processing device for data space is applied to the intermediate end, and the global end corresponds to at least one intermediate end, and the global end is connected to the intermediate end. The device includes: an eighth processing module 810, a ninth processing module 820 and a tenth processing module 830.

[0152] The eighth processing module 810 is used to receive the second information corresponding to the target wheel sent by each intermediate end; the second information includes at least the second clustering result corresponding to the target wheel; the second clustering result is obtained by the intermediate end using the intermediate end clustering model corresponding to the target wheel and training it based on the first clustering result; The ninth processing module 820 is used to train the second clustering result using the global end clustering model corresponding to the target round, and obtain the third clustering result corresponding to the target round. The tenth processing module 830 is used to send the third information corresponding to the target round to the client and the intermediate end; the third information is used to determine the privacy budget of each node in the intermediate end and the client corresponding to the next round of the target round.

[0153] According to the data processing apparatus for data space provided in the embodiments of this application, by introducing an intermediate layer, a multi-level federated clustering model is formed with the client and the global end, realizing the hierarchical fusion of cross-domain data. The intermediate end receives the clustering results uploaded by the client and performs regional-level clustering fusion, effectively combining the data characteristics of each level, ensuring that the correlation information between data levels is preserved, achieving a balance between privacy protection and clustering accuracy. At the same time, through hierarchical aggregation, the stability and accuracy of the clustering results are improved, which can better adapt to multi-level data structures, thereby optimizing the effect of clustering analysis.

[0154] In some embodiments, the tenth processing module 830 can also be used for: Based on the clustering quality of the global end clustering model at the target round and the round before the target, the clustering accuracy gain corresponding to the target round is determined. Based on the privacy consumption rate and clustering accuracy gain corresponding to the target round, obtain the third information corresponding to the target round.

[0155] The data processing device for data space in this application embodiment can be an electronic device or a component of an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM or self-service machine, etc. The embodiments of this application do not specifically limit it.

[0156] The data processing device for data space in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system.

[0157] The data processing apparatus oriented towards data space provided in the embodiments of this application can achieve... Figures 1 to 5The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0158] In some embodiments, such as Figure 9 As shown, this application embodiment also provides an electronic device 900, including a processor 901, a memory 902, and a computer program stored in the memory 902 and executable on the processor 901. When the program is executed by the processor 901, it implements the various processes of the above-described data processing method embodiment for data space and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0159] It should be noted that the electronic devices in the embodiments of this application include the aforementioned mobile electronic devices and non-mobile electronic devices.

[0160] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described data processing method embodiments for data space orientation and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0161] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0162] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described data processing method oriented towards data space.

[0163] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0164] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described data processing method embodiments for data space, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0165] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0166] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0167] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0168] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0169] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0170] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. A data processing method oriented towards data space, characterized in that, The method is applied to a middleware, which corresponds to at least one client and is connected to a global endpoint; the method includes: The system receives first information corresponding to a target wheel sent by each of the clients; the first information includes at least a first clustering result corresponding to the target wheel; the first clustering result is obtained by the client using the client clustering model corresponding to the target wheel and clustering based on local data. The intermediate clustering model corresponding to the target wheel is used to train the first clustering result to obtain the second clustering result corresponding to the target wheel; The second information corresponding to the target wheel is sent to the global terminal; the second information includes at least the second clustering result corresponding to the target wheel; the second clustering result is used by the global terminal to train the global terminal clustering model corresponding to the target wheel to obtain the third clustering result corresponding to the target wheel; Receive the third information corresponding to the target wheel sent by the global terminal.

2. The data processing method oriented towards data space according to claim 1, characterized in that, The step of using the intermediate clustering model corresponding to the target wheel to train the first clustering result and obtain the second clustering result corresponding to the target wheel includes: Based on the hierarchical weights corresponding to the intermediate nodes and the data sensitivity of the nodes in the first clustering result, the privacy budget of each node is determined. The credibility coefficient of each node is determined based on the clustering quality of the nodes, the noise perturbation intensity, and the privacy budget. Based on the credibility coefficient, each node is weighted and fused to obtain the first clustering result after weighted fusion; Based on the first clustering result after weighted fusion, clustering is performed to obtain the second clustering result corresponding to the target round.

3. The data processing method oriented towards data space according to claim 1, characterized in that, After receiving the third information corresponding to the target wheel sent by the global terminal, the method further includes: Based on the third information, the privacy budget of the intermediate node corresponding to the next round of the target round is determined.

4. A data processing method oriented towards data space, characterized in that, Applied to a client, wherein the client corresponds to an intermediate terminal, and the client is connected to the intermediate terminal, the method includes: The client clustering model corresponding to the target round is used and trained based on its own node data to obtain the first clustering result corresponding to the target round; The first information corresponding to the target wheel is sent to the intermediate terminal; the first information includes at least the first clustering result corresponding to the target wheel. Receive the second information corresponding to the target wheel sent by the intermediate terminal and the third information corresponding to the target wheel sent by the global terminal.

5. The data processing method oriented towards data space according to claim 4, characterized in that, The step of using the client-side clustering model corresponding to the target round, training it based on its own node data, and obtaining the first clustering result corresponding to the target round includes: Based on the hierarchical weight corresponding to the client and the data sensitivity of its own node data, the privacy budget of each of its own nodes is determined; Based on the data sensitivity and the privacy budget, each of its own nodes is clustered to obtain the first clustering result corresponding to the target round.

6. The data processing method oriented towards data space according to claim 4 or 5, characterized in that, After receiving the second information corresponding to the target wheel sent by the intermediate terminal and the third information corresponding to the target wheel sent by the global terminal, the method further includes: Based on the second information and the third information, the privacy budget of each node in the client corresponding to the next round of the target round is determined.

7. A data processing method oriented towards data space, characterized in that, Applied to a global endpoint, wherein the global endpoint corresponds to at least one intermediate endpoint, and the global endpoint is connected to the intermediate endpoint, the method includes: The intermediate terminal receives second information corresponding to the target wheel sent by each of the intermediate terminals; the second information includes at least the second clustering result corresponding to the target wheel; the second clustering result is obtained by the intermediate terminal using the intermediate terminal clustering model corresponding to the target wheel and training based on the first clustering result; The second clustering result is trained using the global end clustering model corresponding to the target wheel to obtain the third clustering result corresponding to the target wheel; The third information corresponding to the target round is sent to the client and the intermediate end; the third information is used to determine the privacy budget of each node in the intermediate end and the client corresponding to the next round of the target round.

8. The data processing method oriented towards data space according to claim 7, characterized in that, After training the second clustering result using the global clustering model corresponding to the target wheel to obtain the third clustering result corresponding to the target wheel, the method further includes: Based on the clustering quality of the target wheel and the previous wheel of the target corresponding to the global end clustering model, the clustering accuracy gain corresponding to the target wheel is determined. Based on the privacy consumption rate corresponding to the target round and the clustering accuracy gain, the privacy budget of the node corresponding to the next round of the target round is determined.

9. A data processing device oriented towards a data space, characterized in that, An application in a middleware, wherein the middleware corresponds to at least one client and is connected to a global endpoint, the device includes: A first processing module is configured to receive first information corresponding to a target wheel sent by each of the clients; the first information includes at least a first clustering result corresponding to the target wheel; the first clustering result is obtained by the client using the client clustering model corresponding to the target wheel and clustering based on local data; The second processing module is used to train the first clustering result using the intermediate clustering model corresponding to the target wheel, and obtain the second clustering result corresponding to the target wheel. The third processing module is used to send the second information corresponding to the target wheel to the global terminal; the second information includes at least the second clustering result corresponding to the target wheel; the second clustering result is used by the global terminal to train the global terminal clustering model corresponding to the target wheel to obtain the third clustering result corresponding to the target wheel; The fourth processing module is used to receive the third information corresponding to the target wheel sent by the global terminal.

10. A data processing device oriented towards a data space, characterized in that, Applied to a client, wherein the client corresponds to an intermediate terminal, and the client is connected to the intermediate terminal, the device includes: The fifth processing module is used to train the client clustering model corresponding to the target round based on its own node data to obtain the first clustering result corresponding to the target round; The sixth processing module is used to send the first information corresponding to the target wheel to the intermediate terminal; the first information includes at least the first clustering result corresponding to the target wheel; The seventh processing module is used to receive the second information corresponding to the target wheel sent by the intermediate terminal and the third information corresponding to the target wheel sent by the global terminal.

11. A data processing device oriented towards a data space, characterized in that, Applied to a global endpoint, wherein the global endpoint corresponds to at least one intermediate endpoint, and the intermediate endpoint is connected to the global endpoint, the device includes: The eighth processing module is used to receive second information corresponding to the target wheel sent by each of the intermediate terminals; the second information includes at least the second clustering result corresponding to the target wheel; the second clustering result is obtained by the intermediate terminal using the intermediate terminal clustering model corresponding to the target wheel and training based on the first clustering result; The ninth processing module is used to train the second clustering result using the global end clustering model corresponding to the target wheel, and obtain the third clustering result corresponding to the target wheel; The tenth processing module is used to send the third information corresponding to the target round to the client and the intermediate end; the third information is used to determine the privacy budget of each node in the intermediate end and the client corresponding to the next round of the target round.

12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the data processing method oriented towards data space as described in any one of claims 1-8.