Data space construction method and system based on dynamic ontology modeling and privacy computing

Through dynamic ontology modeling and privacy computing, a structured semantic network is generated and quantum compression is performed, which solves the problems of pattern rigidity and privacy leakage in data space construction and realizes efficient and secure data sharing and storage optimization.

CN120579218BActive Publication Date: 2025-09-30TROY INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511073952.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-09-30
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

Existing data space construction technologies have problems such as rigid models, privacy leakage, and high storage resource consumption. They are difficult to adapt to dynamically changing business entity relationships and are costly.

Method used

A method based on dynamic ontology modeling and privacy computing is adopted. Entities, attributes and relationships are defined through OWL ontology to generate a structured semantic network. The attention mechanism is used to align with the vector space, perform quantum compression and federated learning, and combine the differential privacy algorithm for model training. A low-dimensional index and lineage map are constructed to achieve dynamic adaptation of data patterns and privacy protection.

Benefits of technology

It realizes dynamic adaptive modeling of data space, reduces storage requirements, reduces the risk of privacy leakage, and improves the security and efficiency of data sharing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579218B_ABST
    Figure CN120579218B_ABST
Patent Text Reader

Abstract

The present invention provides a data space construction method and system based on dynamic ontology modeling and privacy computing, which belongs to the field of data processing technology. The method constructs a structured semantic network and a vector space of multimodal data based on multimodal data, and then uses the attention mechanism to align the nodes of the semantic network with the vector space to generate a joint knowledge representation. At the same time, the joint knowledge representation is compressed into a low-dimensional knowledge representation, and a low-dimensional index is constructed. The improved federated learning algorithm combined with differential privacy is used for modeling to obtain the original model; the initial model and initial model parameters are distributed to each participant for federated learning training, and the trained encrypted model parameters are used for model aggregation to obtain a global model; finally, the model parameters of the global model are distributed to each participant for model deployment to form a three-level collaborative data space. The present invention reduces the risk of privacy leakage and reduces the storage overhead of the data space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for constructing a data space based on dynamic ontology modeling and privacy computing. Background Art

[0002] Dataspace is a distributed, multi-label data storage framework for all objects throughout their lifecycle. It is a technical system that enables secure and efficient data connectivity. Building a dataspace is a complex process that integrates multiple technologies, aiming to achieve cross-domain data interconnection, secure sharing, and collaborative computing.

[0003] Current data space construction technologies include traditional data integration, data lake technology, federated learning frameworks, and blockchain data sharing. Transmission data integration solutions utilize a relational database federation architecture based on ETL tools, employing a centralized data warehouse model to build data spaces. Data lake technology builds data spaces based on the unstructured data storage architecture of the Hadoop ecosystem, providing Schema-on-Read capabilities. Data space modeling in both solutions relies on predefined schemas (collections of database objects), making it difficult to adapt to dynamically changing business entity relationships. The federated learning framework solution utilizes Google's horizontal federated learning architecture to build data spaces, which supports distributed model training. However, this solution cannot effectively defend against member inference attacks and poses the risk of model parameter leakage. Blockchain data sharing technology is used to build data spaces, but the resulting data spaces only record transaction hashes and lack fine-grained data lineage tracking capabilities. Furthermore, existing data space construction technologies consume significant storage resources and are costly. Summary of the Invention

[0004] In view of this, the present invention provides a data space construction method and system based on dynamic ontology modeling and privacy computing to solve the problems of model rigidity, privacy leakage and high cost in the current data space construction process.

[0005] The technical solution adopted in the present invention is:

[0006] In a first aspect, the present invention provides a data space construction method based on dynamic ontology modeling and privacy computing, which is applied to data coordination parties, including:

[0007] Obtain multimodal data that needs to be included in the data space, and use OWL ontology to define entities, attributes and relationships in the multimodal data to build a structured semantic network;

[0008] Vectorize multimodal data to generate context-dependent embedding representations and obtain the vector space of multimodal data;

[0009] Utilizing an attention mechanism to align nodes of the structured semantic network with the vector space of the multimodal data to generate a joint knowledge representation;

[0010] Perform online detection of data patterns of multimodal data through edge computing nodes, and expand the joint knowledge representation ontology when changes in data patterns are detected;

[0011] Perform quantum compression on the joint knowledge representation after ontology expansion to obtain a low-dimensional knowledge representation, and construct a low-dimensional index based on the low-dimensional knowledge representation;

[0012] Based on low-dimensional knowledge representation, an improved federated learning algorithm combined with differential privacy is used for modeling to obtain the initial model;

[0013] Distribute the initial model and initial model parameters to each participant, so that each participant can use local client data and initial model parameters to perform federated learning training on the initial model in a trusted execution environment. Perform gradient obfuscation on the local model parameters of the trained initial model and output encrypted model parameters.

[0014] Obtain the encryption model parameters of each participant, aggregate the encryption model parameters of each participant in the trusted execution environment, and obtain a global model;

[0015] The model parameters of the global model are distributed to each participant for local deployment of the global model, so that each participant can share data through the global model and form a three-level collaborative data space.

[0016] Furthermore, the online detection of the data pattern of the multimodal data by the edge computing node and the ontology expansion when the data pattern is detected to have changed include:

[0017] Deploy lightweight edge computing nodes and use online clustering in the edge computing nodes to detect changes in the data schema of multimodal data; wherein the change information includes newly added device fields and abnormal data distribution;

[0018] When a change in the data pattern of multimodal data is detected, the similarity between the current data pattern and the previous data pattern is calculated based on the change information. If the calculated similarity exceeds the preset pattern similarity threshold, the edge computing node generates an OWL extended description based on the change information based on the template engine, and adds the OWL extended description to the OWL ontology of the joint knowledge representation. At the same time, the iterative extension of the OWL ontology is managed through version control.

[0019] Furthermore, the initial model and initial model parameters are distributed to each participant, so that each participant uses local client data and initial model parameters to perform federated learning training on the initial model in a trusted execution environment, including:

[0020] The data coordinator sends the initial model and initial model parameters to the client of each participant. After receiving the initial model and initial model parameters, each participant uses the local client data as training data to input the initial model and conducts model training according to the initial model parameters.

[0021] During the model training process, each participant injects dynamic noise into the training data based on the changes in local client data distribution and the model convergence status, and performs adaptive norm clipping on the client gradients of each participant;

[0022] After the model training is completed, each participant outputs the trained local model and local model parameters.

[0023] Furthermore, performing gradient obfuscation on the local model parameters of the trained initial model and outputting encrypted model parameters includes:

[0024] Each participant calculates a local gradient based on the local model parameters on the local client and injects random noise into the local gradient to obtain the first gradient;

[0025] Each participant generates polynomial coefficients on the local client, performs gradient obfuscation on the first gradient according to the polynomial coefficients to obtain the second gradient, and calculates the local gradient share based on the second gradient through a secure computing protocol to obtain the encrypted model parameters.

[0026] Furthermore, the quantum compression of the joint knowledge representation after ontology expansion to obtain a low-dimensional knowledge representation and the construction of a low-dimensional index based on the low-dimensional knowledge representation include:

[0027] Convert the spatial data points corresponding to the joint knowledge representation after ontology expansion into quantum bit vector representation, and define the objective function to minimize the information loss after data projection;

[0028] The objective function is mapped to a quantum Hamiltonian, and a quantum annealing simulation is performed based on the quantum Hamiltonian. At the same time, the quantum gate parameters are iteratively adjusted to obtain the optimal dimensionality compression path.

[0029] Based on the optimal dimension compression path, the target feature dimensions to be retained are selected from the joint knowledge representation through the quantum random walk algorithm and principal component analysis enhanced analysis operation;

[0030] According to the optimal dimension compression path, the data space dimension corresponding to the joint knowledge representation is compressed into a low-dimensional space, and the data other than the target feature dimension that needs to be retained is projected into the low-dimensional space to obtain a low-dimensional knowledge representation;

[0031] According to the characteristics of the compressed data, R-Tree or KD-Tree is selected as the index structure, and a low-dimensional index is constructed on the low-dimensional knowledge representation based on the index structure.

[0032] Furthermore, the method further comprises:

[0033] Acquire sensitive data uploaded by edge computing nodes, encode the sensitive data into OWL semantic descriptions in a trusted execution environment, calculate the hash value of the OWL semantic description, write the calculated hash value into the blockchain smart contract, generate a hash value certificate, and associate the hash value certificate with the sensitive data uploaded by the edge computing node to form a complete lineage map;

[0034] Obtain a proof acquisition request issued by the edge computing node based on a lightweight verification protocol, generate a zk-SNARK proof based on the proof acquisition request, and feed the zk-SNARK proof back to the edge computing node for edge verification; the proof acquisition request includes the hash value of the sensitive data queried by the edge computing node in the blockchain and the edge computing node identity verification information.

[0035] In a second aspect, the present application provides a data space construction system based on dynamic ontology modeling and privacy computing, which is applied to data coordination parties and includes:

[0036] A semantic network construction module is used to obtain multimodal data that needs to be included in the data space, and use OWL ontology to define entities, attributes and relationships in the multimodal data to build a structured semantic network;

[0037] The vector embedding module is used to vectorize multimodal data, generate context-dependent embedding representations, and obtain the vector space of multimodal data;

[0038] a dynamic fusion module, configured to align the nodes of the structured semantic network with the vector space of the multimodal data using an attention mechanism to generate a joint knowledge representation;

[0039] The ontology expansion module is used to perform online detection of data patterns of multimodal data through edge computing nodes, and to expand the ontology of joint knowledge representation when changes in data patterns are detected;

[0040] The spatial compression module is used to perform quantum compression on the joint knowledge representation after ontology expansion to obtain a low-dimensional knowledge representation, and at the same time construct a low-dimensional index based on the low-dimensional knowledge representation;

[0041] The ontology modeling module is used to model based on low-dimensional knowledge representation and adopt an improved federated learning algorithm combined with differential privacy to obtain an initial model;

[0042] The collaborative training module is used to distribute the initial model and initial model parameters to each participant, allowing each participant to perform federated learning training on the initial model using local client data and initial model parameters in a trusted execution environment. The local model parameters of the trained initial model are then gradient-obfuscated and the encrypted model parameters are output.

[0043] The model aggregation module is used to obtain the encrypted model parameters of each participant and aggregate the encrypted model parameters of each participant in the trusted execution environment to obtain a global model;

[0044] The space construction module is used to distribute the model parameters of the global model to each participant for local deployment of the global model, so that each participant can share data through the global model and form a three-level collaborative data space.

[0045] Furthermore, the data space construction system of the present invention also includes a collaborative verification module, which is used to obtain sensitive data uploaded by the edge computing node, encode the sensitive data into an OWL semantic description in a trusted execution environment, calculate the hash value of the OWL semantic description, write the calculated hash value into the blockchain smart contract, generate a hash value certificate, and associate the hash value certificate with the sensitive data uploaded by the edge computing node to form a complete lineage map;

[0046] The collaborative verification module is also used to obtain a proof acquisition request issued by the edge computing node based on a lightweight verification protocol, generate a zk-SNARK proof based on the proof acquisition request, and feed the zk-SNARK proof back to the edge computing node for edge verification; the proof acquisition request includes the hash value evidence of the edge computing node querying sensitive data in the blockchain and the edge computing node identity authentication information.

[0047] In summary, the beneficial effects of the present invention are as follows:

[0048] The present invention provides a data space construction method based on dynamic ontology modeling and privacy computing. The method uses an attention mechanism to align the nodes of the structured semantic network with the vector space of the multimodal data, generates a joint knowledge representation, performs quantum compression on the joint knowledge representation, obtains a low-dimensional knowledge representation, and constructs a low-dimensional index based on the low-dimensional knowledge representation to store independent metadata in the data space, thereby reducing the storage requirements of the data space, thereby reducing the storage overhead of the data space, and only retaining the most critical features for modeling, reducing interference factors in model learning. At the same time, ontology modeling is performed through the joint knowledge representation after dimensionality reduction, realizing unified modeling of heterogeneous data, and at the same time, it can perceive data pattern changes in real time, support the dynamic evolution and expansion of the ontology, and adapt to a variety of data patterns, solving the problem of pattern rigidity in traditional methods. At the same time, during the federated learning training process, the method adopts a federated learning algorithm based on differential privacy for model training to ensure that each participant does not disclose ontology features during model training, thereby reducing the risk of privacy leakage. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work, and these are all within the scope of protection of the present invention.

[0050] Figure 1 This is a flow chart of a data space construction method based on dynamic ontology modeling and privacy computing of the present invention;

[0051] Figure 2 This is a schematic diagram of the data evidence storage process of the present invention;

[0052] Figure 3 This is a functional module block diagram of a data space construction system based on dynamic ontology modeling and privacy computing in the present invention. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. If there is no conflict, the various features of the present invention and the embodiments can be combined with each other and are all within the scope of protection of the present invention.

[0054] The detailed implementation process of the present invention is shown in the following examples.

[0055] Example 1: Reference Figure 1 As shown, Figure 1 This is a flow chart of a data space construction method based on dynamic ontology modeling and privacy computing in the present invention. Figure 1 As shown, the method of the embodiment of the present invention includes:

[0056] S1: Obtain multimodal data that needs to be included in the data space, and use OWL ontology to define entities, attributes and relationships in the multimodal data to build a structured semantic network;

[0057] S2: Vectorize the multimodal data to generate context-dependent embedding representations and obtain the vector space of the multimodal data;

[0058] S3: aligning the nodes of the structured semantic network with the vector space of the multimodal data using an attention mechanism to generate a joint knowledge representation;

[0059] S4: Perform online detection of data patterns of multimodal data through edge computing nodes, and expand the ontology of joint knowledge representation when changes in data patterns are detected;

[0060] S5: Perform quantum compression on the joint knowledge representation after ontology expansion to obtain a low-dimensional knowledge representation, and construct a low-dimensional index based on the low-dimensional knowledge representation;

[0061] S6: Based on low-dimensional knowledge representation, an improved federated learning algorithm combined with differential privacy is used for modeling to obtain the initial model;

[0062] S7: Distribute the initial model and initial model parameters to each participant, so that each participant can use local client data and initial model parameters to perform federated learning training on the initial model in a trusted execution environment. Gradient obfuscation is performed on the local model parameters of the trained initial model, and the encrypted model parameters are output.

[0063] S8: Obtain the encryption model parameters of each participant, aggregate the encryption model parameters of each participant in the trusted execution environment, and obtain a global model;

[0064] S9: Distribute the model parameters of the global model to each participant for local deployment of the global model, so that each participant can share data through the global model and form a three-level collaborative data space.

[0065] Specifically, in step S1 of this embodiment of the present invention, OWL (Web Ontology Language) technology is used to define the entities, attributes, and relationships in multimodal data. Based on these defined entities, attributes, and relationships, a structured semantic network, or knowledge graph, is constructed. The core concept of OWL is ontology, a model used to define concepts, entities, and the relationships between them in a domain. An ontology typically includes definitions of classes, attributes, instances, and relationships.

[0066] In step S2 of the embodiment of the present invention, when vectorizing multimodal data, the BERT model or GraphSAGE (Graph Samples and Aggregations) can be used to perform the vectorization operation. If the BERT model is used to vectorize unstructured data (such as text or sensor streams), the specific vectorization process is as follows:

[0067] Data preprocessing: Unstructured data is segmented and special tags, such as classification tags and separator tags, are added. The segmentation results are then digitized and converted into word IDs. Word order information is then represented using learnable positional embeddings (i.e., BERT's built-in positional encoding). Finally, segment encoding is performed based on word order information to distinguish multiple text segments in the input data and generate segment embeddings.

[0068] Vector generation: The processed word ID, position embedding, and segment embedding are fed into the BERT model for vectorization. The model then uses a multi-layer Transformer encoder to capture contextual semantics. Each layer includes a self-attention mechanism and a feed-forward neural network (FFN). The model then outputs a context-sensitive embedding representation.

[0069] Furthermore, if GraphSAGE (Graph Sampling and Aggregation) is used for vectorization, taking sensor data as an example, the specific vectorization process is as follows: data characteristics and graph construction. Data characteristics: Sensor stream data typically has spatiotemporal correlation (spatial node topology + temporal sequence characteristics).

[0070] Each sensor is a node in the graph, with the node's features being the sensor's measured values ​​(e.g., time series data such as temperature, pressure, and acceleration). The edges of the graph are defined based on spatial and temporal associations. Edge relationships are stored using an adjacency matrix or adjacency list.

[0071] Vectorized operations. Learns node embeddings by aggregating neighbor node features, supporting inductive learning (capable of handling unseen nodes). Inputs the adjacency matrix and adjacency table corresponding to node features into the GraphSAGE model, outputting context-sensitive embeddings.

[0072] The BERT model is suitable for text and sequence data, with its deep semantic modeling and pre-trained model generalization capabilities. The GraphSAGE model is suitable for graph-structured data (such as sensor networks), capturing spatial topological relationships and supporting inductive learning. In practical applications, it is important to select a model for vectorization based on data characteristics and task objectives, with emphasis on preprocessing (such as text segmentation and graph construction) and downstream task adaptation (such as fine-tuning and feature fusion).

[0073] Specifically, in the embodiment of the present invention, in step S3, the attention mechanism is used to align the nodes of the structured semantic network with the vector space of the multimodal data to generate a joint knowledge representation, which specifically includes:

[0074] First, a graph neural network is used to embed nodes and edges in a structured semantic network. Node embedding generates a node's semantic vector h by aggregating features from neighboring nodes (e.g., the attention mechanism of GAT), capturing structural context and semantic associations. Edge relationship embedding encodes the semantic relationship between edges into a vector r, enhancing the representation of semantic dependencies between nodes. Graph neural networks can employ neural networks such as GNN, GCn, GAT, and GraphSAGE. After embedding, nodes are subjected to hierarchical semantic modeling. A hierarchical attention mechanism is used to assign weights to nodes at different levels, highlighting key semantic nodes.

[0075] Next, a shared linear layer is used to map multiple feature vectors in the vector space to a common semantic space. The features in the common semantic space are then fused using either an early or late fusion strategy to obtain a multimodal joint representation. The early fusion strategy involves concatenating the features of each modality in the common semantic space and generating a joint feature using an MLP (Multi-Layer Perceptron). The late fusion strategy involves retaining the features of each modality separately and dynamically selecting modal information using an attention mechanism.

[0076] Finally, for each semantic node n (embedded as h) and multimodal features, attention weights are calculated, and a multimodal contextual representation corresponding to semantic node n is generated. Attention is calculated using an attention mechanism (such as the Transformer). The node's neighborhood structure and relationship type are taken into account when calculating attention. Attention is then allocated to global multimodal features for high-level semantic nodes, while attention is allocated to local features for low-level nodes. After attention allocation, for each semantic node n, its original semantic embedding h and the multimodal contextual representation are fused to obtain a node-modality joint embedding. Finally, a graph pooling operation is used to aggregate the joint embeddings of all nodes to obtain a joint knowledge representation.

[0077] Furthermore, in step S4 of the embodiment of the present invention, the data pattern of the multimodal data is detected online by the edge computing node, and when a change in the data pattern is detected, the ontology is expanded, specifically including:

[0078] Lightweight edge computing nodes are deployed, and online clustering is used in the edge computing nodes to detect changes in the data patterns of multimodal data; wherein the change information includes newly added device fields and abnormal data distribution.

[0079] When a change in the data schema of multimodal data is detected, the similarity between the current data schema and the previous data schema is calculated based on the change information. If the calculated similarity exceeds a preset schema similarity threshold, the edge computing node generates an OWL extension description based on the change information based on the template engine. This OWL extension description is then added to the OWL ontology of the federated knowledge representation, and the iterative extension of the OWL ontology is managed through version control. This embodiment of the present invention achieves dynamic ontology evolution by sensing changes in data schemas in real time, performing incremental updates based on schema change information, and automatically extending the ontology.

[0080] Specifically, the online clustering method in embodiments of the present invention is implemented using an online clustering algorithm, such as the Streaming K-means algorithm. Streaming K-means is an online version of the K-means algorithm. Its core concept is to process samples in the data stream one by one, dynamically updating the cluster centers to avoid reprocessing historical data. Using online clustering, changes in the data patterns of multimodal data are detected. The specific process is as follows: Monitoring parameters are first initialized and set. These monitoring parameters include the initial number of clusters K, the distance threshold (the minimum distance for determining whether a sample is an outlier), the center drift threshold (the displacement amplitude for determining whether a cluster has significantly changed), and the attenuation factor (controlling the weight of the influence of historical samples on the cluster centers). After the parameters are set, K samples are randomly selected from the initial data stream as initial cluster centers. Online clustering is performed based on the monitoring parameters and the initial cluster centers. The algorithm cyclically monitors for abnormal signals, center drift signals, and new cluster signals to determine whether the data patterns of the multimodal data have changed. If a sample's distance from all cluster centers exceeds the distance threshold, an outlier signal is detected, and a newly added device field is reported. When a center drift signal is detected in which the cluster center displacement is greater than the center drift threshold, or a new cluster signal in which the dynamic cluster number K1 exceeds the initial cluster number K is detected, abnormal data distribution is fed back.

[0081] Furthermore, in step S5 of the embodiment of the present invention, quantum compression is performed on the joint knowledge representation after ontology expansion to obtain a low-dimensional knowledge representation, and a low-dimensional index is constructed based on the low-dimensional knowledge representation, including:

[0082] Convert the spatial data points corresponding to the joint knowledge representation after ontology expansion into quantum bit vector representation, and define the objective function to minimize the information loss after data projection;

[0083] The objective function is mapped to a quantum Hamiltonian, and a quantum annealing simulation is performed based on the quantum Hamiltonian. At the same time, the quantum gate parameters are iteratively adjusted to obtain the optimal dimensionality compression path.

[0084] Based on the optimal dimension compression path, the target feature dimensions to be retained are selected from the joint knowledge representation through the quantum random walk algorithm and principal component analysis enhanced analysis operation;

[0085] According to the optimal dimension compression path, the data space dimension corresponding to the joint knowledge representation is compressed into a low-dimensional space, and the data other than the target feature dimension that needs to be retained is projected into the low-dimensional space to obtain a low-dimensional knowledge representation;

[0086] According to the characteristics of the compressed data, R-Tree or KD-Tree is selected as the index structure, and a low-dimensional index is constructed on the low-dimensional knowledge representation based on the index structure.

[0087] In this embodiment, quantum compression uses the principle of quantum state superposition to capture high-probability feature directions of joint knowledge representation (similar to the principal components of classical PCA) and filter out redundant dimensions (such as background noise dimensions in sensor data). The compressed data only retains the most critical features for modeling (such as edges and texture features in images), reducing interference factors in model learning.

[0088] Specifically, constructing a low-dimensional index of the data space in a low-dimensional space is crucial for addressing efficiency bottlenecks, storage costs, and computational complexity in high-dimensional data processing. It also deeply integrates with technical scenarios such as data privacy protection and federated learning. Traditional indexes only associate data locations, while low-dimensional indexes preserve semantic information through quantum dimensionality reduction (e.g., reducing image pixel vectors to feature vectors for the "cat / dog" classification), enabling "semantic retrieval" (e.g., searching for "images containing cars" directly matches vector clusters corresponding to the semantic meaning of "car" in the low-dimensional space).

[0089] The pseudo code example of dimension compression is:

[0090] def Q_DimReduce(data, target_dim):

[0091] qubit_states = encode_to_qubits(data) # Data quantum state encoding

[0092] energy_fn = define_energy_function(qubit_states) # Build energy function

[0093] optimized_params = quantum_annealing(energy_fn) # Quantum annealing optimization

[0094] compressed_data = project_dimensions(optimized_params, target_dim) # Dimension compression

[0095] return build_index(compressed_data) # Build a low-dimensional index

[0096] Specifically, R-Tree is a data structure obtained by extending B-tree to multidimensional conditions. It is applicable to geometric objects in multidimensional space (such as points, rectangles, polygons, etc.), and is particularly suitable for unstructured or semi-structured data (such as GIS geographic data, image area features). KD-Tree is a tree-shaped data structure that stores instance points in k-dimensional space for rapid retrieval. It is mainly used for point data in multidimensional space and is suitable for structured numerical data (such as coordinate points, feature vectors). Therefore, in an embodiment of the present invention, the corresponding index structure can be selected according to the actual application scenario requirements. For example, if the data is a non-point object or the data needs to be dynamically updated, R-Tree is selected as the index structure. If the data is pure point data and has a low dimension, KD-Tree can be selected as the index structure.

[0097] Specifically, in step S6 of this embodiment of the present invention, differential privacy noise (DP-FedAvg+ algorithm) is introduced into federated learning. Based on low-dimensional knowledge representation, an improved federated learning algorithm incorporating differential privacy is used for modeling to obtain an initial model. This ensures that model training does not leak ontology features. By embedding privacy protection into the ontology modeling process, the risk of privacy leakage is addressed.

[0098] The training process of the classic federated learning framework can be summarized as follows: 1. The coordinator establishes an initial model; 2. The model's initial structure and parameters are sent to each participant; 3. Each participant trains the model using local data; 4. Each participant sends the resulting model parameters to the coordinator; 5. The coordinator aggregates the participants' models to construct a more accurate global model. Each participant trains the model on their local client and establishes a shared model mechanism in the cloud to update the model. Through federated learning, all training data remains on each participant's device, and the resulting trained model achieves the desired results.

[0099] Furthermore, in step S7 of the embodiment of the present invention, the initial model and initial model parameters are distributed to each participant, so that each participant uses local client data and initial model parameters to perform federated learning training on the initial model in a trusted execution environment, specifically including:

[0100] The data coordinator sends the initial model and initial model parameters to the client of each participant. After receiving the initial model and initial model parameters, each participant uses the local client data as training data to input into the initial model and performs model training according to the initial model parameters.

[0101] The data coordinator uses Shamir's secret sharing to enable secure multi-party collaborative gradient computation, allowing multiple parties to collaboratively compute gradients in an encrypted state. When distributing the initial model parameters to each participant's client, the data coordinator uses Shamir's secret sharing to decompose the initial model parameters into n shares and distribute them to the n participants. Each participant can then use their local shares and data to compute gradients, secretly sharing the gradients to generate new shares.

[0102] During the model training process, each participant injects dynamic noise into the training data based on the changes in local client data distribution and the model convergence status, and performs adaptive norm clipping on the client gradients of each participant to prevent outliers from affecting the noise utility.

[0103] Finally, after the model training is completed, each participant outputs the trained local model and local model parameters.

[0104] During dynamic noise injection, the noise level is set inversely proportional to the number of training rounds, and data is weighted based on sensitivity, with medical data receiving a higher weight, for example. The client-side gradient refers to the direction of model parameter updates calculated by the client on the local dataset. Compared to traditional differential privacy approaches that use a fixed amount of noise, which can lead to reduced model accuracy, this method reduces accuracy loss by dynamically adjusting noise injection (experiments have shown a 12% improvement in model accuracy).

[0105] Furthermore, in step S7 of the embodiment of the present invention, gradient obfuscation is performed on the local model parameters of the trained initial model to output the encrypted model parameters, including:

[0106] Each participant calculates a local gradient based on the local model parameters on the local client and injects random noise into the local gradient to obtain a first gradient. In this embodiment of the present invention, the local model parameters trained by each participant are equivalent to the new shares calculated based on the distributed shares.

[0107] Each participant generates polynomial coefficients on the local client, performs gradient obfuscation on the first gradient according to the polynomial coefficients to obtain the second gradient, and calculates the local gradient share based on the second gradient through a secure computing protocol to obtain the encrypted model parameters.

[0108] In the embodiment of the present invention, participants can hide individual data features and resist reasoning attacks between members by dynamically injecting random noise locally on the client.

[0109] Specifically, in step S8 of the embodiment of the present invention, the encrypted model parameters of each participant are obtained, and the encrypted model parameters of each participant are aggregated in a trusted execution environment to obtain a global model. Specifically, the data coordinator uses Shamir's reconstruction method to sum the shares of the encrypted model parameters uploaded by each participant, aggregates the local gradient shares into a complete global gradient through Lagrange interpolation, uses the aggregated global gradient to update the model parameters, and repeats the model parameter update process until the model converges to obtain a global model.

[0110] Specifically, the three-level collaborative data space constructed in step S9 of this embodiment of the present invention is divided into three layers based on the characteristics of the data space: an edge computing layer, a fog computing coordination layer, and a cloud core layer. Its core function is to dynamically bind data governance granularity (device / domain / cross-domain) with the computing layer (edge / fog / cloud). The edge computing layer is deployed on IoT gateways / edge servers and is primarily responsible for protocol adaptation, streaming cleansing, local decision-making, data pattern monitoring, and data desensitization. The fog computing coordination layer is deployed in regional data centers / 5G MEC nodes and is primarily responsible for intra-domain federated aggregation, dynamic routing, and TEE privacy computing. The cloud core layer is deployed in cloud server clusters and is primarily responsible for providing a trusted execution environment (TEE), global model updates, blockchain evidence storage, and cross-domain policy arbitration.

[0111] After the global model is deployed globally, the edge computing layer, fog computing coordination layer, and cloud core layer all possess the global model. The edge computing layer extracts features from local data on edge computing nodes, ensuring that raw data never leaves the device. The fog computing coordination layer then fuses these extracted features to achieve intra-domain privacy computing. The cloud core layer makes global decisions and integrates knowledge across domains.

[0112] Example 2: In an embodiment of the present invention, based on the above-mentioned Example 1, a dynamic GeoHash partitioning strategy for the index is also provided to achieve grid granularity adjustment of the data space. Among them, the dynamic GeoHash partitioning strategy mainly monitors the changes in data distribution density through edge computing nodes through real-time data perception, such as the movement of IoT device locations. High-density areas and low-density areas are divided according to the data distribution density, and the grid granularity is adapted. Among them, the high-density area adopts the method of refining the network to adjust the granularity, for example, increasing the GeoHash accuracy from 6 bits to 10 bits. The low-density area adjusts the grid granularity by merging grids, for example, reducing the accuracy from 6 bits to 4 bits. The code for adjusting the grid granularity of the data space is as follows:

[0113] def adjust_geohash_precision(region_density):

[0114] if region_density>1000: # points / square kilometer

[0115] return 10 # high precision

[0116] elif region_density<100:

[0117] return 4 # low precision

[0118] else:

[0119] return 6 # default precision

[0120] This embodiment of the present invention also provides a cross-layer cache consistency protocol (CL-CacheSync), which ensures consistent cache versions across the three layers of data in mid-flight and uses incremental updates (Delta Sync) to reduce the amount of synchronized data. This cross-layer cache consistency protocol is based on an improved and optimized Paxos algorithm. Its main improvements and optimizations include:

[0121] (1) Set up three layers of cache: edge cache, fog cache, and cloud index. Edge cache is used to cache hot data at edge computer nodes, such as frequently queried local area indexes. Fog cache caches regional index metadata. Cloud index stores global indexes.

[0122] (2) Set up incremental synchronization to transfer only the changed area data.

[0123] The specific protocol optimization code is as follows:

[0124] class CL_CacheSync:

[0125] def __init__(self):

[0126] self.edge_cache = {} # Edge cache

[0127] self.fog_cache = {} # Fog cache

[0128] self.cloud_index = None # Cloud index

[0129] def sync(self, updated_region):

[0130] # Incremental synchronization: only transfer changed area data

[0131] delta = compute_delta(self.cloud_index, updated_region)

[0132] self.edge_cache.update(delta)

[0133] self.fog_cache.update(delta)

[0134] self.cloud_index.merge(delta)

[0135] Furthermore, in an embodiment of the present invention, the data space construction method based on dynamic ontology modeling and privacy computing further includes step S10:

[0136] Obtain sensitive data uploaded by the edge computing node, encode the sensitive data into OWL semantic description in a trusted execution environment, calculate the hash value of the OWL semantic description, write the calculated hash value into the blockchain smart contract, generate a hash value certificate, and associate the hash value certificate with the sensitive data uploaded by the edge computing node to form a complete lineage map.

[0137] The edge computing node receives a proof request based on a lightweight verification protocol, generates a zk-SNARK proof based on the proof request, and feeds the zk-SNARK proof back to the edge computing node for edge verification. The proof request includes the hash value of the sensitive data queried by the edge computing node in the blockchain and the edge computing node's identity verification information.

[0138] Specifically, this embodiment of the present invention uses the zk-SNARKs lightweight protocol to verify the integrity of cloud-based computations. Edge computing nodes use zk-SNARKs proofs to verify the correctness of computations in the cloud-based executable environment (TEE) without accessing the original data D. The zk-SNARKs lightweight protocol allows verifiers to process only a minimal amount of data (e.g., a few KB of proof), ensuring computational integrity without having to re-run the entire calculation. Furthermore, the proof process does not require the disclosure of original data or intermediate computational steps, thus ensuring data security.

[0139] In traditional blockchain evidence storage solutions, a common misunderstanding in data traceability is that "on-chain means traceability". In fact, simply storing hashes can only solve the problem of data existence proof, but cannot trace data lineage. Users often only store transaction hashes on the blockchain (such as the Bitcoin model), which can only prove that "a certain transaction has occurred", but cannot know the transaction content and data evolution process. Therefore, this embodiment adopts a structured summary hash chain storage method, where each hash corresponds to a semantic description (OWL semantics) of a data operation event (such as cleaning, conversion), and then calculates its hash value and stores it on the chain. Figure 2 As shown, this embodiment encodes sensitive data (or key operational events) on edge computing nodes into standardized OWL semantic descriptions through edge data operations. This OWL semantic description is then structured and digested to calculate a digest hash. This digest hash is then written into a blockchain smart contract, generating a hash value certificate. Simultaneously, the hash value certificate is linked to sensitive data in the edge computing node's local database (i.e., the off-chain metadata repository) to create a complete lineage map of the sensitive data, facilitating subsequent data lineage tracking.

[0140] The embodiment of the present invention utilizes the attention mechanism to align the nodes of the structured semantic network with the vector space of the multimodal data, generates a joint knowledge representation, and performs ontology modeling through the joint knowledge representation, thereby realizing unified modeling of heterogeneous data. At the same time, it can also perceive data pattern changes in real time, support the dynamic evolution and expansion of the ontology, and adapt to a variety of data patterns, thus solving the problem of pattern rigidity in traditional methods. At the same time, during the federated learning training process, the method uses a federated learning algorithm based on differential privacy for model training, ensuring that each participant does not disclose ontology features during model training, thereby reducing the risk of privacy leakage. In addition, the method also performs quantum compression on the joint knowledge representation after ontology expansion to obtain a low-dimensional knowledge representation, and constructs a low-dimensional index based on the low-dimensional knowledge representation to store independent metadata in the data space, thereby reducing the storage requirements of the data space and thereby reducing the storage overhead of the data space.

[0141] Specifically, the method of the embodiment of the present invention has the following technical advantages in terms of privacy computing architecture, algorithm integration depth, trusted execution mechanism, and attack defense capabilities:

[0142] (1) The privacy computing architecture adopts a three-level collaborative architecture, namely the edge layer (data desensitization), fog computing layer (federated learning) and cloud layer (TEE+blockchain), which work together to improve the privacy protection of the data space;

[0143] (2) A dynamic privacy budget allocation mechanism based on data sensitivity and adaptive noise adjustment during model training (DPFedAvg+) is adopted to deeply integrate differential privacy with the federated learning modeling process to solve the problem of "privacy leakage risk";

[0144] (3) The trusted execution environment (TEE) is coordinated with the blockchain to complete sensitive calculations within the TEE, and the result hash is stored on the chain. Sensitivity weighted calculation solves the problem of privacy protection during the training phase, and blockchain verification solves the problem of trusted proof of data operations. Together, they form a closed loop of the privacy lifecycle.

[0145] (4) A gradient obfuscation mechanism was designed to inject obfuscated noise into federated learning, which can resist member inference attacks (leakage probability <0.01%) and solve the problem of model parameter leakage in the federated learning process.

[0146] Example 3: Reference Figure 2 As shown, based on Example 1, this embodiment of the present invention further provides a data space construction system based on dynamic ontology modeling and privacy computing, which is applied to data coordination parties and includes:

[0147] A semantic network construction module is used to obtain multimodal data that needs to be included in the data space, and use OWL ontology to define entities, attributes and relationships in the multimodal data to build a structured semantic network;

[0148] The vector embedding module is used to vectorize multimodal data, generate context-dependent embedding representations, and obtain the vector space of multimodal data;

[0149] a dynamic fusion module, configured to align the nodes of the structured semantic network with the vector space of the multimodal data using an attention mechanism to generate a joint knowledge representation;

[0150] The ontology expansion module is used to perform online detection of data patterns of multimodal data through edge computing nodes, and to expand the ontology of joint knowledge representation when changes in data patterns are detected;

[0151] The spatial compression module is used to perform quantum compression on the joint knowledge representation after ontology expansion to obtain a low-dimensional knowledge representation, and at the same time construct a low-dimensional index based on the low-dimensional knowledge representation;

[0152] The ontology modeling module is used to model based on low-dimensional knowledge representation and adopt an improved federated learning algorithm combined with differential privacy to obtain an initial model;

[0153] The collaborative training module is used to distribute the initial model and initial model parameters to each participant, allowing each participant to perform federated learning training on the initial model using local client data and initial model parameters in a trusted execution environment. The local model parameters of the trained initial model are then gradient-obfuscated and the encrypted model parameters are output.

[0154] The model aggregation module is used to obtain the encrypted model parameters of each participant and aggregate the encrypted model parameters of each participant in the trusted execution environment to obtain a global model;

[0155] The space construction module is used to distribute the model parameters of the global model to each participant for local deployment of the global model, so that each participant can share data through the global model and form a three-level collaborative data space.

[0156] Furthermore, the data space construction system of an embodiment of the present invention also includes a collaborative verification module, which is used to obtain sensitive data uploaded by the edge computing node, and encode the sensitive data into an OWL semantic description in a trusted execution environment, calculate the hash value of the OWL semantic description, and write the calculated hash value into the blockchain smart contract, generate a hash value evidence, and associate the hash value evidence with the sensitive data uploaded by the edge computing node to form a complete bloodline map.

[0157] The collaborative verification module also receives proof requests from edge computing nodes based on a lightweight verification protocol, generates zk-SNARK proofs based on these proof requests, and feeds the zk-SNARK proofs back to the edge computing nodes for edge verification. The proof requests contain the hash value of the sensitive data queried by the edge computing node in the blockchain and the edge computing node's identity verification information.

[0158] The embodiment of the present invention uses the attention mechanism through the dynamic fusion module to align the nodes of the structured semantic network with the vector space of the multimodal data to generate a joint knowledge representation, and performs ontology modeling based on the joint knowledge representation through the ontology modeling module to achieve unified modeling of heterogeneous data. At the same time, it can also perceive the changes in data patterns in real time through the ontology expansion module, support the dynamic evolution and expansion of the ontology, and adapt to a variety of data patterns, solving the problem of pattern rigidity in traditional methods. At the same time, during the federated learning training process, the collaborative training module uses a federated learning algorithm based on differential privacy to train the model, ensuring that each participant does not disclose the ontology features during model training, reducing the risk of privacy leakage. In addition, the system's space compression module also performs quantum compression on the joint knowledge representation after ontology expansion to obtain a low-dimensional knowledge representation, and at the same time constructs a low-dimensional index based on the low-dimensional knowledge representation to store independent metadata in the data space, reducing the storage requirements of the data space, and thereby reducing the storage overhead of the data space.

[0159] In the embodiment of the present invention, a comparative analysis of the technical innovation advantages is conducted with traditional technologies, with the comparison dimensions including mode adaptability, privacy protection strength, computing efficiency and storage overhead. The specific advantages of the embodiment of the present invention are shown in Table 1 below:

[0160] Table 1 Comparison of technological innovation advantages

[0161]

[0162] Specifically, the embodiments of the present invention can be applied in smart city scenarios to achieve secure data integration across multiple departments. Actual measurements show a 17-fold increase in cross-domain query efficiency. The latency of genetic data sharing in healthcare scenarios has been reduced to 200ms. Applied in industrial IoT device data storage scenarios, the cost of building a data space can be reduced by 42%. The embodiments of the present invention effectively address the core pain points of existing data space technologies, creating significant technical advantages in terms of dynamic adaptability, privacy protection, and computational efficiency.

[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data space construction method based on dynamic ontology modeling and privacy computing, applied to data coordination, characterized by: include: Obtain multimodal data that needs to be included in the data space, and use OWL ontology to define entities, attributes and relationships in the multimodal data to build a structured semantic network; Vectorize multimodal data to generate context-dependent embedding representations and obtain the vector space of multimodal data; Utilizing an attention mechanism to align nodes of the structured semantic network with the vector space of the multimodal data to generate a joint knowledge representation; Perform online detection of data patterns of multimodal data through edge computing nodes, and expand the joint knowledge representation ontology when changes in data patterns are detected; Perform quantum compression on the joint knowledge representation after ontology expansion to obtain a low-dimensional knowledge representation, and construct a low-dimensional index based on the low-dimensional knowledge representation; Based on low-dimensional knowledge representation, an improved federated learning algorithm combined with differential privacy is used for modeling to obtain the initial model; Distribute the initial model and initial model parameters to each participant, so that each participant can use local client data and initial model parameters to perform federated learning training on the initial model in a trusted execution environment, perform gradient obfuscation on the local model parameters of the trained initial model, and output encrypted model parameters; Obtain the encryption model parameters of each participant, aggregate the encryption model parameters of each participant in the trusted execution environment, and obtain a global model; The model parameters of the global model are distributed to each participant for local deployment of the global model, so that each participant can share data through the global model and form a three-level collaborative data space.

2. The data space construction method according to claim 1, characterized in that: The online detection of the data pattern of the multimodal data by the edge computing node and the ontology expansion when the data pattern is detected to have changed include: Deploy lightweight edge computing nodes and use online clustering in the edge computing nodes to detect changes in the data schema of multimodal data; wherein the change information includes newly added device fields and abnormal data distribution; When a change in the data pattern of multimodal data is detected, the similarity between the current data pattern and the previous data pattern is calculated based on the change information. If the calculated similarity exceeds the preset pattern similarity threshold, the edge computing node generates an OWL extended description based on the change information based on the template engine, and adds the OWL extended description to the OWL ontology of the joint knowledge representation. At the same time, the iterative extension of the OWL ontology is managed through version control.

3. The data space construction method according to claim 1, characterized in that: The initial model and initial model parameters are distributed to each participant, so that each participant uses local client data and initial model parameters to perform federated learning training on the initial model in a trusted execution environment, including: The data coordinator sends the initial model and initial model parameters to the client of each participant. After receiving the initial model and initial model parameters, each participant uses the local client data as training data to input the initial model and conducts model training according to the initial model parameters. During the model training process, each participant injects dynamic noise into the training data based on the changes in local client data distribution and the model convergence status, and performs adaptive norm clipping on the client gradients of each participant; After the model training is completed, each participant outputs the trained local model and local model parameters.

4. The data space construction method according to claim 1, characterized in that: The step of performing gradient obfuscation on the local model parameters of the trained initial model and outputting encrypted model parameters includes: Each participant calculates a local gradient based on the local model parameters on the local client and injects random noise into the local gradient to obtain the first gradient; Each participant generates polynomial coefficients on the local client, performs gradient obfuscation on the first gradient according to the polynomial coefficients to obtain the second gradient, and calculates the local gradient share based on the second gradient through a secure computing protocol to obtain the encrypted model parameters.

5. The data space construction method according to claim 1, characterized in that: The method of performing quantum compression on the joint knowledge representation after ontology expansion to obtain a low-dimensional knowledge representation and constructing a low-dimensional index based on the low-dimensional knowledge representation includes: Convert the spatial data points corresponding to the joint knowledge representation after ontology expansion into quantum bit vector representation, and define the objective function to minimize the information loss after data projection; The objective function is mapped to a quantum Hamiltonian, and a quantum annealing simulation is performed based on the quantum Hamiltonian. At the same time, the quantum gate parameters are iteratively adjusted to obtain the optimal dimension compression path. Based on the optimal dimension compression path, the target feature dimensions to be retained are selected from the joint knowledge representation through the quantum random walk algorithm and principal component analysis enhanced analysis operation; According to the optimal dimension compression path, the data space dimension corresponding to the joint knowledge representation is compressed into a low-dimensional space, and the data other than the target feature dimension that needs to be retained is projected into the low-dimensional space to obtain a low-dimensional knowledge representation; According to the characteristics of the compressed data, R-Tree or KD-Tree is selected as the index structure, and a low-dimensional index is constructed on the low-dimensional knowledge representation based on the index structure.

6. The data space construction method according to claim 1, characterized in that: Also includes: Acquire sensitive data uploaded by edge computing nodes, encode the sensitive data into OWL semantic descriptions in a trusted execution environment, calculate the hash value of the OWL semantic description, write the calculated hash value into the blockchain smart contract, generate a hash value certificate, and associate the hash value certificate with the sensitive data to form a complete lineage map; Obtain a proof acquisition request issued by the edge computing node based on a lightweight verification protocol, generate a zk-SNARK proof based on the proof acquisition request, and feed the zk-SNARK proof back to the edge computing node for edge verification; the proof acquisition request includes the hash value of the sensitive data queried by the edge computing node in the blockchain and the edge computing node identity verification information.

7. A data space construction system based on dynamic ontology modeling and privacy computing, applied to data coordination, characterized by: include: A semantic network construction module is used to obtain multimodal data that needs to be included in the data space, and use OWL ontology to define entities, attributes and relationships in the multimodal data to build a structured semantic network; The vector embedding module is used to vectorize multimodal data, generate context-dependent embedding representations, and obtain the vector space of multimodal data; a dynamic fusion module, configured to align the nodes of the structured semantic network with the vector space of the multimodal data using an attention mechanism to generate a joint knowledge representation; The ontology expansion module is used to perform online detection of data patterns of multimodal data through edge computing nodes, and to expand the ontology of joint knowledge representation when changes in data patterns are detected; The spatial compression module is used to perform quantum compression on the joint knowledge representation after ontology expansion to obtain a low-dimensional knowledge representation, and at the same time construct a low-dimensional index based on the low-dimensional knowledge representation; The ontology modeling module is used to model based on low-dimensional knowledge representation and adopt an improved federated learning algorithm combined with differential privacy to obtain an initial model; The collaborative training module is used to distribute the initial model and initial model parameters to each participant, allowing each participant to perform federated learning training on the initial model using local client data and initial model parameters in a trusted execution environment. The local model parameters of the trained initial model are then gradient-obfuscated and the encrypted model parameters are output. The model aggregation module is used to obtain the encrypted model parameters of each participant and aggregate the encrypted model parameters of each participant in the trusted execution environment to obtain a global model; The space construction module is used to distribute the model parameters of the global model to each participant for local deployment of the global model, so that each participant can share data through the global model and form a three-level collaborative data space.

8. The data space construction system according to claim 7, characterized in that: The system also includes a collaborative verification module, which is used to obtain sensitive data uploaded by the edge computing node, encode the sensitive data into an OWL semantic description in a trusted execution environment, calculate a hash value of the OWL semantic description, write the calculated hash value into a blockchain smart contract, generate a hash value certificate, and associate the hash value certificate with the sensitive data uploaded by the edge computing node to form a complete lineage map; The collaborative verification module is also used to obtain a proof acquisition request issued by the edge computing node based on a lightweight verification protocol, generate a zk-SNARK proof based on the proof acquisition request, and feed the zk-SNARK proof back to the edge computing node for edge verification; the proof acquisition request includes the hash value evidence of the edge computing node querying sensitive data in the blockchain and the edge computing node identity authentication information.

Citation Information

Patent Citations

  • Multimodal content processing method, apparatus, device and storage medium

    US20210192142A1

  • KR20210037619A